Agents and MCP · 6 MIN
Are multi-agent systems worth it? When to use one agent or many
Multi-agent systems cost more and add coordination risk. Here is when splitting work across agents pays off, and when a single agent is the better call.
Multi-agent systems are worth it when the work is broad, can be split into independent parts, and is valuable enough to pay for extra compute. They are not worth it when the parts depend on each other or all need the same context. Anthropic's own research system, which uses a lead agent directing parallel subagents, reports roughly 15 times the token use of a chat interaction, so the decision is an economic one as much as a technical one.
- Start with a single agent or a workflow. Split into multiple agents only when one context window or one sequence of steps is the bottleneck.
- Anthropic reports that agents use about four times the tokens of chat and multi-agent systems about fifteen times, so the task value must justify the cost.
- Multi-agent designs fit breadth-first work with independent directions. They fit poorly when agents must share context or depend on each other's output.
- Most coding tasks have fewer truly parallel parts than research tasks, which limits the benefit.
- Subagents with their own context windows are also useful for isolation and role separation, even without parallel speedups.
- Nactore designs multi-agent systems only where evals show a single agent falls short.
What is a multi-agent system?
A multi-agent system uses more than one model-driven agent to complete a task. The most common shape is orchestrator and workers. A lead agent analyzes the request, plans, and spawns subagents that each handle one part. The subagents work in parallel and return findings, and the lead combines them.
Anthropic's write-up of its research system describes this pattern. Subagents act as intelligent filters, searching and returning only what matters, and a separate component handles source attribution. It is a useful reference because the authors also publish the costs and the limits.
If the term agent is new, read AI agents vs workflows first.
What do the published numbers say about cost?
Anthropic reports that agents typically use about 4 times more tokens than chat interactions, and multi-agent systems about 15 times more. It also reports that, on its internal research evaluation, a multi-agent setup with a stronger lead model and smaller subagents beat a single stronger agent by 90.2 percent. Those are Anthropic's figures for its own research task, not a general promise, and your results will depend on your task and your measurement.
The takeaway is not "more agents is better". The authors state that the approach needs tasks where the value is high enough to pay for the added performance. They also report that token usage alone explained most of the performance variance on a browsing evaluation, which hints that part of the gain comes from spending more compute on the problem.
| Setup | Relative cost | Where it tends to fit |
|---|---|---|
| Single LLM call or fixed workflow | Lowest | Known steps, high volume |
| Single agent with tools | Higher | Open-ended tasks with one thread of work |
| Multi-agent, orchestrator and workers | Highest | Broad tasks with independent parallel directions |
When does splitting work across agents help?
The same source lists the conditions where it works well. Use them as a checklist.
- The task is breadth-first. Many independent directions can be explored at once, such as researching ten competitors or scanning many data sources.
- The work exceeds one context window. Each subagent holds its own context and passes back a summary, which keeps the lead's context manageable.
- Parallelism is real. The pieces do not wait on each other.
- The tool set is complex. Different agents can specialize in different interfaces.
- The value per task is high. A report, an analysis, or a migration that a person would take hours to do.
A second, quieter benefit is separation of concerns. A subagent with a narrow role, a short system prompt, and a limited tool allowlist is easier to reason about and test than one agent that does everything. That holds even when nothing runs in parallel.
When should you stay with a single agent?
The research write-up is candid about the poor fits. Domains where all agents need to share the same context, or where there are many dependencies between agents, are not a good match yet. It also says most coding tasks involve fewer truly parallelizable parts than research does, and that real-time coordination between agents remains hard.
Practical red flags in our experience:
- Agents hand off vague summaries. The second agent lacks the detail it needs, and quality drops.
- Agents duplicate work. Without clear boundaries, two subagents search the same thing.
- The coordinator becomes the bottleneck. If the lead must read everything anyway, you have added cost without saving context.
- Debugging gets harder. A failure might sit in the plan, a handoff, or a worker, and you need traces to tell which.
See also coding agents in engineering teams, where the shared-context problem is most visible.
How do you design one that works?
Anthropic's lessons translate into a short design guide.
- Write explicit delegation instructions. Each subagent needs an objective, an output format, guidance on tools, and clear boundaries.
- Embed scaling rules. Tell the lead how much effort a task deserves. A simple lookup needs one agent with a few calls, not ten.
- Design the tools carefully. Tool descriptions steer behavior. See tool calling design for agents.
- Start broad, then narrow. Let agents explore before they focus.
- Evaluate early and small. Anthropic suggests starting with around 20 representative test cases because early changes tend to show large effects. Add an automated grader and keep human review for what it misses.
- Trace everything. You need to see the plan, each subagent's actions, and the handoffs.
Before adding a second agent, run the same task with one agent given more time or better tools, and score both on the same eval. If the single agent is close, keep it. You will save cost and debugging effort.
How do you decide for a business process?
Ask three questions. Is the work naturally parallel? Does each part need only a slice of the context? Is the output valuable enough to justify several times the compute? If you answer yes to all three, prototype a multi-agent design and compare it with a single agent on a fixed set of real cases. If any answer is no, a workflow or single agent is probably the better engineering choice.
Frequently asked questions
Are multi-agent systems more accurate?
Sometimes. Anthropic reports a gain on its internal research eval, but it ties the approach to tasks where extra compute is justified. Test on your own task before assuming a gain.
Is a subagent the same as a second model?
No. A subagent is a separate agent instance with its own context window, instructions, and tool access. It can use the same underlying model or a different one.
Can we use multiple agents for coding?
Carefully. Parallel sessions on separate branches or worktrees can work well for independent tasks. Tightly coupled changes are harder to split.
What is the biggest risk?
Cost and coordination errors. Set token budgets, limit turns, and trace handoffs so you can find failures.
Takeaway
Use more agents only when the task is broad, parallel, and valuable, and prove it with an eval against a single-agent baseline. Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.