Story 004 - The MCP Detour That Taught Me Microservices

Published:

I was asked to migrate two chatbots to Model Context Protocol (MCP). One was newer and had been built around a tool-calling architecture from the beginning. Let’s call it chatbot B. The other was the first chatbot I built. Let’s call it chatbot A. It had started as a RAG-based chatbot built on top of a knowledge base containing three FAISS sources, but it was now growing beyond retrieval as a new ticketing-automation use case was being added through tool calling. Its knowledge base and retrieval pipeline, including the embedder and reranker, were still part of the chatbot repository and application code, technically making it a monolith.

Fortunately, I am a hardcore functional programmer. This made moving chatbot B to an MCP architecture surprisingly easy. Instead of providing the planner LLM with a small set of statically defined tool schemas, the planner started discovering tools from MCP servers, and the local function executor was replaced by an MCP client. Moving the actual tool functions to the MCP servers was also easy because everything was already modularized by functionality.

Although chatbot A was a monolithic deployment, it had also been built using functional programming. FAISS index loading, vector search, embedding and reranking calls, and context construction were already independent functions with explicit inputs and outputs. This made the migration much easier than it initially looked. The difficult part was not rewriting the retrieval pipeline. It was deciding the service boundaries.

To separate chatbot A from its knowledge base and retrieval pipeline, I had to decide which modules represented knowledge and retrieval and which belonged to the chatbot experience. Everything involving the knowledge files, FAISS indexes, embeddings, reranking, and context construction moved into a new repository. Conversation state, tool planning, response generation, streaming, and presentation stayed with the chatbot. This became my first proper polyrepo split.

The new repository exposed the existing retrieval pipeline through a small FastAPI service. All the functions moved without changing a single line inside them. They were only renamed and rearranged into new files according to their responsibilities. The new retrieval service would run on the same server where the inference stack was already serving the required models for the retrieval pipeline. The compute placement stayed the same. Only the software boundary changed. Even the vLLM endpoints for the embedder and reranker remained unchanged. The final flow became: Chatbot -> RAG MCP -> New Retrieval API -> Embedding + FAISS + Reranking -> Retrieved Documents.

Chatbot A also became a lightweight tool-calling agent similar to chatbot B. Instead of RAG being executed inside chatbot A’s app.py, the RAG MCP called the exposed retrieval service’s BASE_URL/api/kb/search endpoint. It passed the same inputs used by the monolithic version and returned the outputs in the same format. That separation felt ridiculously clean.

This also solved an ownership problem that had been annoying me during every UAT cycle. Instead of testing the existing functionality and giving sign-off, users would keep sending more documents to add to the knowledge base. My future plan for this is basically an UNO reverse card: build them a portal so the application-support team can manage approved knowledge themselves, while engineering continues to own the ingestion and retrieval pipeline.

Of course, I also learned that microservices have their own problems. A Python function call becomes a network call. One application becomes multiple processes, with the chatbot, MCP servers, and retrieval API listening on different ports. Starting everything requires process supervision. Errors now need to cross service boundaries cleanly. Repositories need compatible API contracts. A function signature can no longer be changed casually because another independently deployed service may depend on it. But the current scale is not large enough for these problems to become painful yet, so I will let them be problems for the future.

The migration taught me that microservices do not remove complexity. They move complexity away from implementation and into interfaces, deployment, networking, and operations. But when the boundary matches actual ownership, that trade can be worth it.

Looking back, MCP was only the detour that forced me to examine the architecture. Registering Python functions as MCP tools was the easy part. The valuable work was deciding what the chatbot should own, what the retrieval system should own, and where the knowledge base should live.

The monolith had worked because its functional modules were already clean. The microservice split worked because those functions could be lifted into another repository without rewriting their internals. I still prefer functions, explicit data flow, and boring modules over elaborate object hierarchies. This migration only reinforced that preference. Good functional boundaries made it possible to introduce service boundaries later.

Chatbot A’s system has more repositories, more ports, more processes, more network calls, and more ways to fail. But it also feels much cleaner.