# Multi-agent systems: parallelise the reading, keep the writing in one thread

> Anthropic and Cognition published opposite conclusions a day apart. Read the two posts together, though, and the real dividing line is fairly clear: don't ask how many agents you need. Ask which agent is allowed to write.

Original: https://fdetimes.net/en/analysis/multi-agent-parallel-reads-single-writer/

On 12 June 2025, Walden Yan of Cognition published a post with a blunt title: "Don't Build Multi-Agents". The next day, Anthropic unveiled its Research system, in which a lead agent running Claude Opus 4 orchestrates Sonnet 4 subagents. On Anthropic's own internal research eval, the system outperformed a single-agent Opus 4 by 90.2%.

Two companies building agents, two opposite conclusions, 24 hours apart. If you work as an FDE, a client will put the question to you soon enough: "Should we be using multi-agent?"

Getting it wrong in either direction is expensive. Say no when the problem genuinely needs it, and the client misses a significant gain in quality. Say yes when it doesn't, and you leave behind a system that costs several times more, fails unpredictably, and may well need debugging at 2am by you.

## The two camps are not really arguing

Read closely, Anthropic never claimed multi-agent is good for everything. Its post says the architecture is strongest on high-value tasks that involve heavy parallelisation. In the same post, it concedes that most coding tasks have fewer truly parallelisable parts than research does.

Cognition, which builds a coding agent, looks at the problem from the other side. Walden Yan's argument is that when many agents collaborate, context and decisions get scattered, and the result is a fragile system.

His second principle explains why: every action carries implicit decisions, and conflicting implicit decisions produce bad results.

A LangChain post, published three days after Anthropic's, reconciles the two views. Its observation: reading is inherently easier to parallelise than writing. Seen through that split, the other two posts make sense: research is reading, writing code is writing.

**Key point:** The right question is not "one agent or many" but "how many agents are allowed to write".

Consider an example. One subagent reads documentation on vendor A, another reads about vendor B. The two jobs do not tread on each other, because reading does not change the world.

But if the first subagent writes a date-handling function in UTC and the second writes interface code that implicitly assumes local time, each is "correct" on its own and wrong once combined. Those are exactly the implicit decisions Cognition is talking about.

## The cost is not just the bill

The figure Anthropic published, roughly 15 times the tokens of chat, is easier to grasp with a hypothetical: an exchange costing 10,000 tokens becomes about 150,000. For a market analysis report that takes an analyst a full day, that is cheap; for a chatbot answering questions about leave policy, it is absurd.

The token bill is the visible part. Since late 2024, Anthropic has warned that agent autonomy brings higher costs and the risk of compounding errors. A second hypothetical: if each step of an agent is 95% correct, a chain of 10 dependent steps is correct only about 60% of the time (0.95 to the power of 10).

More agents means more steps, more handoffs, more places for errors to multiply.

## Why debugging gets much harder

With ordinary software, you rerun the same input and see the same bug. Agents do not work that way. Anthropic writes that agents make dynamic decisions and are non-deterministic between runs, even with identical prompts.

Multiply that across many agents talking to each other, and you have a system where yesterday's failure may not reproduce today. Anthropic also admits that LLM agents are not yet good at coordinating and delegating to other agents in real time. In other words, the hardest part of the architecture is the part where the models are weakest.

Cognition also talks about traces, but from a different angle. Its first principle is to share context: pass other agents the full agent trace, not just individual messages. For Anthropic, the trace is a tool for engineers to find bugs; for Cognition, it is what an agent needs so it does not make decisions with missing information.

Two different purposes, but they point to the same thing: what needs preserving is the whole chain of steps, not just the final answer.

## The pattern that works: many eyes reading, one hand writing

In April 2026, Cognition returned to the subject with an update that was softer than the previous year's headline. Its conclusion: multi-agent currently works best when writing stays in a single thread. That matches LangChain's read/write split.

The most useful concrete pattern from that post is the clean-context reviewer. According to Cognition, such a reviewer catches bugs that the coding agent itself cannot see. The reason is simple: the agent that wrote the code carries all of its own assumptions, while the reviewer only reads the result, is not steered by the earlier line of thinking, and, crucially, writes nothing.

Putting the sources together yields a decision table that none of them sets out in full:

| Client scenario | Read or write | Sensible architecture | Why |
|---|---|---|---|
| Synthesising information from many independent sources | Mostly read | Lead agent + parallel subagents | Splittable, and valuable enough to justify the token cost |
| Editing code across many files | Write | One linear agent | Implicit decisions from multiple agents will conflict |
| Checking code an agent has just written | Read | Add a clean-context reviewer | Catches bugs the writing agent misses, without touching the write path |
| Simple Q&A or lookup | Light reading | One model call or one agent | The simplest option is enough; the cost isn't justified |

The last row is where many projects go wrong. Anthropic's advice from "Building effective agents" still holds: find the simplest solution possible, and add complexity only when it is genuinely needed.

## On a client site, where do you start?

Suppose an insurance company wants an agent that reads claim files, checks them against policy terms, and updates claim status in the core system. Your first step is not to draw a five-agent diagram. List every action and label it: reading the claim, reading the policy terms and looking up the customer's history are reads; updating the status is a write.

From there, the architecture almost suggests itself. Independent read steps can go to subagents running in parallel if the volume is large enough. The write to the core system is handled by exactly one agent, with the full context in hand. If mistakes are costly, add a read-only reviewer that checks before the write.

And before writing the first line of agent code, set up tracing. When the client calls to say "yesterday it approved a claim it shouldn't have", what you need is the full trace of that exact run, not a promise to try to reproduce it.

## Talk about decisions, not agent counts

If you see "multi-agent experience" in a job description, don't read it as a demand to build systems with as many agents as possible. An interviewer who knows the field may well turn the question around: why didn't you use a single agent?

So when you write your CV or describe a project, talk about decisions, not the number of agents. A line such as "split document lookup into parallel subagents, kept data updates in a single agent, added a reviewer and tracing" shows you understand both the benefits and the costs.

If you have ever collapsed a multi-agent system back into a single agent and it became more stable, that is an even better interview story.

The skill worth practising alongside this is reading traces. If you have followed a chain of tool calls and pinpointed where two implicit decisions collided, that is a concrete example to bring to an interview, and far more convincing than a list of orchestration frameworks.

More than a year after the dueling headlines of June 2025, the answer does not lie with either camp. It lies in a simple classification you can make at your first discovery session: is this job reading or writing?

**Try this week:**

- Take an agent you are working on, list every tool call and mark each as a read or a write; only the all-read stretches are worth trying to split into subagents.
- Turn on full tracing for one agent run (prompts, tool calls, the result of each step), then try to find the cause of a failure from the trace alone.
- Add a reviewer step: call a fresh model with clean context, give it only the diff and the original request, then compare the number of bugs it catches with your own self-review.

## Sources

- [How we built our multi-agent research system (Anthropic Engineering)](https://www.anthropic.com/engineering/multi-agent-research-system)

- [Don't Build Multi-Agents (Cognition)](https://cognition.com/blog/dont-build-multi-agents)

- [Multi-Agents: What's Actually Working (Cognition)](https://cognition.com/blog/multi-agents-working)

- [How and when to build multi-agent systems (LangChain)](https://www.langchain.com/blog/how-and-when-to-build-multi-agent-systems)

- [Building effective agents (Anthropic)](https://www.anthropic.com/research/building-effective-agents)
