Multi-Agent Architecture: Why Only One Agent Should Write
Splitting an agent into specialists breaks four very specific things in production. The principle that makes them impossible: one writes, the others think.
Splitting a big AI agent into small specialized agents is the natural reflex once an assistant grows. Too many tools, too many responsibilities, a prompt that overflows: you decompose, the way we decomposed monoliths into microservices.
I argued for that decomposition here in March, and the arguments still hold. But production taught me a nuance that changes everything, and it is the subject of this article: what you decompose matters. You can distribute intelligence. You do not distribute the right to write. That lesson came from building agents and workflows powered by language models, not from theory.
To keep everyone on board, we will follow a deliberately simple example: an AI assistant that manages a cooking blog over chat. Its user writes "add my plum tart recipe", "create a Summer Menus page", "reply to this comment". The assistant reads and writes directly into the site. Everything below applies to any agent that modifies a real system: a shop, a CRM, a workspace.
Three words before we start
Three terms come up everywhere, with an image for each:
- a tool is a button the model can press: create a recipe, read a page, publish a reply;
- an agent is the model, plus its written instructions, plus the list of buttons it is allowed to press;
- a sub-agent is a colleague the agent hands the file to when a request falls outside its scope. The one distributing the files is the supervisor.
The supervisor pattern, and why teams adopt it
Once the assistant grows, the classic decomposition looks like this:
user
|
supervisor no business tools: it routes, then writes the reply
|
+-- recipes create, edit, organize recipes writes
+-- pages compose the blog's pages writes
+-- readers reply to comments writesA supervisor with no business tools at all, whose only job is to understand the request, hand it to the right specialist, then write the final answer. Three sub-agents behind it, each with its own prompt, its own tools, its own scope.
The reasons to adopt this pattern are solid. A single agent carrying every tool sees its context window fill up with irrelevant descriptions. Its permissions become impossible to scope. A prompt change for one use case breaks another. Specialization fixes all of that on paper.
On paper only. Because this diagram contains one detail that will decide everything: all three specialists write. Each of them can modify the blog. And every modification crosses a boundary.
Where it breaks: the boundary
Follow one ordinary request: "Create a Summer Menus page with my three latest recipes".
The supervisor receives the message. It writes a handover note for the pages sub-agent: free text summarizing what it understood. The pages sub-agent wakes up with that note as its only luggage, works, writes into the site, then writes a report. The supervisor reads the report and answers the user.
Four handoffs. Each one can fail silently.
The report gets lost in transit. A sub-agent works under a step ceiling, a cap on how many actions it can chain before it is cut off. If it hits that cap right after creating the page but before writing its report, the supervisor gets a blank sheet. A blank sheet looks like a failure. So it announces a failure and offers to retry, while the page exists. The user retries: now there are two pages.
Consent does not cross the boundary. The user said "yes, publish". That yes stays on the supervisor's desk, while the sub-agent is the one who executes. Between them sits nothing but the handover note. If the note rephrases badly, the sub-agent asks again for a confirmation already given. From the outside, the assistant looks like it is not listening.
Neither does memory. The sub-agent arrives without having read the conversation: every delegation starts on a blank sheet. "My three latest recipes", "like last time", "same style as the home page": anything the supervisor forgets to copy into its note does not exist for the specialist.
And state fragments. The recipes sub-agent creates a recipe and confirms: it does exist in the database. But creating a recipe displays it nowhere: it has to be added to a page, and pages belong to another sub-agent. Neither of them sees the whole board. The assistant then confuses "it exists in the database" with "a visitor can see it", and keeps insisting everything is fine while the user stares at an empty page.
These four failures share one trait: the model did nothing wrong. Each agent, taken alone, did its job. What fails is the handoff. The boundary between agents is made of free text, and free text loses information at every crossing.
The telltale sign is the day you catch yourself writing a prompt rule like "at most one modification per delegation, even if the user asked for several". That is not a business constraint. It is architectural scar tissue dressed up as an operating policy.
What the state of the art says
This diagnosis is not one person's opinion. Three sources, three angles, same conclusion.
Mastra says it in its documentation: keep a single agent when the task is short, tightly coupled, and supported by a manageable tool set. For anything one good agent already handles, extra agents only add cost and a bigger surface to debug.
LangChain supplies the missing criterion: read actions are inherently more parallelizable than write actions. Multi-agent shines on breadth-first queries, the ones where you go fetch information in ten directions at once. It fails in domains that require all agents to share the same context. An assistant that modifies a system through conversation is precisely that second case.
Cognition had published the sharpest position, "Don't Build Multi-Agents". Their follow-up, "Multi-Agents: What's Actually Working", is more useful still because it states what works rather than what fails:
Multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions.
And an observation most readers skim past: most multi-agent setups in production are limited to read-only sub-agents, which mostly resemble tool calls rather than true multi-agent collaboration. That is not a criticism. That is the shape that holds.
The principle: one writes, several think
Writes are single-threaded. Intelligence is distributed.
One agent writes. It holds the conversation memory, the user's consent, the tone, the voice. It is the only thing with authority over system state.
The specialists stop being colleagues you hand the file to. They become consultants: you call them for an opinion, they think it through on their own with their own instructions, they return a structured answer, and they never touch the site. Technically they are no longer agents but model-powered tools.
THE WRITER · one agent, one voice, one authority
write tools consolidated by family
read tools deterministic, free, testable
intelligence tools the consultants, no write rights
compose_page intent + blog state -> a plan
write_copy brief + constraints -> the text
review_page state after writing -> what is offReplay the earlier request. "Create a Summer Menus page with my three latest recipes": the writer reads the site, calls the compose_page consultant which returns a plan, asks for confirmation exactly once, writes the blocks, calls review_page to check the result, and answers. The memory, the yes and the site state never leave the same context. The four failures from the previous chapter are not fixed one by one: they become impossible to express.
The important nuance: specialization does not disappear. Each consultant keeps its isolated context, its dedicated prompt, a different model if needed. What disappears is its right to write.
And a short message stays short. "Feature the plum tart" triggers no delegation at all: the writer answers on its own. In the supervisor pattern, every message, even a trivial one, crossed a boundary.
The monolith objection, and how to defuse it
Immediate objection, and a fair one: if one agent carries every tool, you land back on the monolith you were escaping. The research on tool selection is unambiguous, accuracy degrades past ten to fifteen tools.
The answer fits in one word: measure. Every tool has to be described to the model, with its name, what it does and the shape of its parameters. That manual ships in full on every message. An assistant with a few dozen tools can easily send the equivalent of twenty pages of text in descriptions on every call, before the user has said a word. Measuring takes an hour, with a script that compiles each description exactly the way your framework does right before calling the model. No model calls, no tokens spent.
And the measurement almost always reveals the same thing: the weight comes from tool families. Ten tools named add_menu_block, add_heading_block, add_grid_block are ten ways of saying "add something to the page". A family consolidates into a single tool where one parameter carries the type. Ten buttons become one button and a dropdown.
Two measurement lessons along the way:
Measure the total, not the worst case. Descriptions ship on every call: it is the sum you pay for in tokens, not the largest tool.
Date your constraints. If a family was split into ten tools, it is often because of a vendor limit, such as the caps on structured outputs, the mechanism that forces the model to answer in exactly the shape you asked for. Those limits move regularly, and usually upward. An architectural decision made under a vendor limit carries an expiry date nobody writes down. Check yours: the constraint that justified your split may no longer exist.
Consolidate first, merge second. In that order. Consolidation is not a bonus, it is the precondition: without it, the merge produces an obese agent that picks its tools badly; with it, the single writer often carries less than one of your former sub-agents did.
What single-threaded writes delete
The most convincing argument is not what the principle brings, it is what it lets you never build. All the machinery that manages the boundary disappears with it:
- The reports between agents, and the guards that detect when they arrive empty or truncated
- The per-delegation cost accounting, to know which sub-agent spent what
- The defensive rules like "one modification per delegation"
- The prompt blocks explaining to each specialist what it cannot see
- The suspend and resume workflows, those state machines that pause an operation while a human answers, and that only exist because a sub-agent cannot wait
That last one deserves a sentence: with a single writer, an operation awaiting confirmation is no longer a suspended workflow to persist and resume. It is a state, inside a conversation that keeps going.
What you decompose, and what you do not
None of this contradicts the March article. The monolithic agent does break, decomposition is still the right answer, MCP and A2A are still the right plumbing: the former to wire agents to tools, the latter to let them talk to each other.
The nuance is what you decompose.
You decompose intelligence: contexts, prompts, models, specialties. That is where separation pays, because it isolates reasoning, and reasoning parallelizes well.
You do not decompose write authority. It stays in one place, with the memory and the consent. A write spread across several agents is a consensus to rebuild on every turn, and nobody builds a reliable consensus out of free text passed between two models.
The microservices parallel says the same thing, pushed one notch further. We split the services. We never split data ownership: each piece of data belongs to exactly one service, and the others read it. That is precisely the rule missing from most multi-agent architectures.
One agent that writes. Several that think.
The rest is plumbing.