Story 006 - The Parallelism That Made Things Slower
Published:
Back in 2025, I added a ThreadPoolExecutor to the retrieval pipeline of the RAG chatbot I was building. The idea of concurrency came into the retrieval picture because, for every query, this pipeline had three independent FAISS sources (.index files) to retrieve similar texts from. And I intuitively thought this could obviously be searched concurrently.
Although my original intuition was to implement true multi-core processing via ProcessPoolExecutor, after a vibe brainstorming, I eventually went with ThreadPoolExecutor as I learned that FAISS releases the GIL when performing vector calculations in C++ under the hood. And I figured that while heavy calculations for one source were going on in C++, the GIL would be released, allowing other Python threads to make progress, including operations around the search such as fetching required metadata from RAM and collecting the top search results along with their corresponding texts, while the C++ calculations for another source were going on.
And it worked. I had three independent searches. They could run concurrently. Problem solved.
That was pretty much where I stopped thinking.
I didn’t benchmark it. I didn’t profile it. I didn’t ask whether the retrieval workload was actually large enough for the overhead of external concurrency to be worthwhile.
I had essentially done what I would later learn is one of the cardinal sins of performance engineering: I optimized something I hadn’t measured.
Fast forward to 2026.
As I started diving deeper into HPC and performance engineering, I began looking at some of my old engineering decisions with a slightly different mindset. One of them suddenly bothered me.
“Wait… how much did that concurrency actually help?”
So I measured it, and the measurements shattered my reality.
0.3 ms was recorded by Sequential execution (Search A then B then C).
~1.4 ms was recorded by the ThreadPoolExecutor implementation.
And then came the one that made me stare at the terminal for a while. ~69 ms for the ProcessPoolExecutor.
Oh.
This led me down the cProfile profiling journey, which eventually became the motivation to document all of this and write the technical report paper: Analyzing Internal and External Parallelism in Multi-Index FAISS Retrieval: Threads, Processes, OpenMP and the Cost of Concurrency.
Lesson Learned
Measure before you optimize
Something being parallelizable doesn’t mean making it parallel will make it faster. The workload, the granularity of the work, and the overhead introduced by the parallel execution model all matter.
I also learned that parallelism exists at multiple layers. FAISS was already doing its own internal parallel work through OpenMP, while I was adding another layer of concurrency around it. Understanding how those layers interact turned out to be far more interesting than simply making three searches run at the same time.
Most importantly, I learned that performance engineering is less about having the right intuition and more about being willing to challenge your intuition with measurements.
Turns out, I had some measuring to do.
