---
title: "The Case for Multi-Agent Systems"
url: "https://www.morphllm.com/blog/multi-agent-systems"
description: "Anthropic's multi-agent system outperformed single-agent Opus by 90%. The reason isn't better models. It's that intelligence degrades when you ask one agent to do everything, and improves when you let specialists work in isolation."
date: "2026-04-05"
author: "Tejas Bhakta"
---
# The Case for Multi-Agent Systems

Anthropic built a research system where Opus 4 hands work down to Sonnet 4 sub-agents. It [beat single-agent Opus 4 by 90.2%](https://www.anthropic.com/engineering/multi-agent-research-system).

The setup that used a weaker model for most of the work beat the stronger model working on its own by 90 percent.

It's not a fluke either. Once you know what actually caps agent performance, it's the result you'd predict. And the cap isn't how smart the model is. It's context.

## What breaks when one agent does everything

Every LLM has a context window, and every token sitting in that window is pulling on the model's attention. Watch an agent research something. It opens a couple dozen documents, chases a few dead ends, doubles back, and finally lands on the thing it needed. The part that mattered might be 2,000 tokens. Getting there cost it 50,000.

In a single-agent setup all 50,000 of those stick around. The model drags them into every reasoning step that follows, and the signal-to-noise ratio falls a little further with each tool call.

None of this is hand-waving. [Chroma ran 18 LLMs](https://research.trychroma.com/context-rot) across 8 input lengths, and every one of them got worse as the context grew. Opus 4.6 sheds about 14 percentage points over a 750K-token span. Models that clear 90% accuracy on short prompts fall off a cliff by 32K tokens. Years earlier Liu et al. saw the same thing in [Lost-in-the-Middle](https://arxiv.org/abs/2307.03172): bury the relevant fact in the middle of a long context and accuracy drops more than 30%.

The cause is baked into the architecture. Ten thousand tokens means the transformer is juggling 100 million pairwise relationships. A hundred thousand tokens makes it 10 billion. Attention doesn't scale linearly, so extra context does more than water down the relevant part. It makes the model physically worse at attending to anything.

Manus put a number on the imbalance. The [input-to-output token ratio](https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus) for an agent runs around 100 to 1. Nearly everything an agent processes is input: tool results, file reads, search output. Your frontier model spends the bulk of its capacity reading rather than thinking. Which makes controlling what it reads the single highest-leverage thing you can do.

## How multi-agent systems fix it

In Anthropic's design, each sub-agent works a question inside its own context window. It reads 10,000-plus tokens of documents, weighs them, and sends back 1,000 to 2,000 tokens of condensed findings. The lead agent never lays eyes on the dead ends. It only ever gets the distillate.

The mechanism is context isolation. Every agent starts on a clean window, and the dead ends die inside the sub-agent that hit them instead of bleeding into the lead.

The interesting thing is how many teams landed on this independently. Anthropic pushes research sub-tasks down to Sonnet agents so Opus can stay on synthesis. Claude Code fans out [Task agents](https://docs.anthropic.com/en/docs/agents-and-tools/claude-code/overview) into parallel context windows to explore. Cognition shipped [SWE-grep](https://cognition.ai/blog/swe-grep), a search agent that hands back nothing but the files that matter. Sourcegraph's Amp describes its sub-agents as ["fundamentally changing your relationship with the context window by giving you a multiplication of context windows"](https://ampcode.com/notes/how-to-build-an-agent). Cursor runs its background search isolated from whatever you're editing.

When five separate teams reach for the same architecture for the same reason, the architecture is probably correct.

## Multi-agent is also cheaper

Splitting the work across agents costs less, not more.

In Anthropic's system a Sonnet sub-agent is a fraction of an Opus token. The heavy lifting, the research and the exploration and the retrieval, runs on the cheap model, and Opus only ever touches the distilled result.

We watched this happen ourselves with [WarpGrep on SWE-Bench Pro](/blog/warpgrep-v2). Bolting a dedicated search sub-agent onto Opus 4.6 made the whole system **15.6% cheaper**, $2.51 a task against $3.06, and **28% faster**, 445 seconds against 618. We added a model to the pipeline and both the cost and the latency went down.

The mechanism is simple. The expensive model burns fewer tokens. Left on its own Opus opens dozens of files during a search, carries all of them forward, and reasons over a bloated window. Give it a sub-agent and it receives only the files that survived the filter. Fewer input tokens, fewer output tokens, done sooner.

Anthropic found that [token usage explains 80% of the performance variance](https://www.anthropic.com/engineering/multi-agent-research-system) on BrowseComp. Not the choice of model. Not the prompt. Token efficiency. Multi-agent systems come out ahead here for free, because a sub-agent throws away its exploration before any of it reaches the lead.

There's an obvious analogy. Spend the expensive intelligence on synthesis and the decisions; spend the cheap intelligence on the digging. It's the same split that lets a company function. The CEO doesn't read every email in the building. Analysts filter and summarize and escalate.

## Why a bigger context window won't save you

The natural pushback: if long context is the disease, just make the window bigger and sturdier. Models handle 1M and 2M tokens now. Doesn't that end the conversation?

It doesn't. Chroma's study already included the newest long-context models, and all 18 of them fell off anyway. This rot lives in attention, not in some ceiling on token count. Hand a 2M window 500K tokens of noise and it does worse than a 32K window holding 32K tokens of pure signal.

[Augment Code watched](https://www.augmentcode.com/tools/context-window-wars-200k-vs-1m-token-strategies) accuracy slide from 89% at 8K tokens to 25% at 1M. So the million-token window is mostly a spec-sheet figure. What you can actually use sits far below what's advertised.

Spotify's engineers ran into the same thing [in production](https://engineering.atspotify.com/2025/11/context-engineering-background-coding-agents-part-2), where their agents "tended to get lost when context window filled up, forgetting the original task after a few turns."

A bigger window hands you more rope. A multi-agent system means you need less rope in the first place.

## This isn't only about code

Coding agents are where the measurements are cleanest, which is why this post leans on them. But nothing about the argument is specific to code.

It's easy to forget that Anthropic's 90% came from a *research* system. The job there was answering hard questions by pulling together information scattered across many sources. Whatever isolation buys a coding agent, it buys any agent grinding through multi-step information work just the same.

You can spot the domains where it applies by their shape. There's a lot of sources to sift before you find what's relevant. Most of what you open is a dead end. And the last reasoning step, the one that actually produces the answer, needs a clean window or the output falls apart. Research fits that. So does analysis, customer support sitting on a knowledge base, legal review, financial diligence, and most enterprise work that touches unstructured data.

The plainer fact under all of it: intelligence organizes into hierarchies whenever resources are tight. The moment a single agent can't keep everything in working memory, the job gets split. Each specialist runs at full attention on its narrow slice while a coordinator stitches the outputs together. Nobody designed that as a workaround. It's just what a system does once its cognitive capacity is finite and the task isn't.

## When it's worth doing

A lot of tasks get nothing out of going multi-agent. A chatbot answering simple questions with no tools has no use for a sub-agent. Orchestration is pure overhead until context pollution becomes the thing holding you back, and then it starts paying for itself.

Watch the token split. Once an agent is spending more than half its tokens on retrieval and exploration, multi-agent will help it.

[Cognition measured](https://cognition.ai/blog/swe-grep) coding agents burning 60% of their time on search, which is well past that line. Anything with a comparable retrieval-to-reasoning ratio is worth trying it on.

A few other signals worth watching. Success rate that falls off as tasks run longer, dropping after ten or more tool calls, usually means context pollution. An agent that keeps re-reading things it already found has pushed its earlier reads out of effective attention. And results that get *worse* when you feed in more context are the textbook sign of attention dilution.

Start small. Take the highest-volume retrieval task, put one sub-agent on it, and measure whether the lead agent's output improves. In our experience, and in Anthropic's numbers, the jump is big and it shows up right away.

## The evidence just keeps stacking

On SWE-Bench Pro the same model can land [17 problems apart](https://www.swebench.com/) depending on the scaffold around it. On SWE-Bench Lite, GPT-4 scored [2.7% under one scaffold and 28.3% under another](https://arxiv.org/pdf/2509.16941). Identical model, different harness, and the harness that handles context well is the one that wins.

Anthropic found that [going from Sonnet 3.7 to Sonnet 4](https://www.anthropic.com/engineering/multi-agent-research-system) bought a bigger gain than doubling Sonnet 3.7's token budget. The right model in the right context beats more tokens in a noisy one.

[SWE-Search](https://arxiv.org/abs/2410.20285) got a 23% relative improvement across five models by running Monte Carlo Tree Search over the agent's exploration, no bigger model and no extra training. [LocAgent](https://aclanthology.org/2025.acl-long.426/) lifted downstream code resolution by 12% on nothing but better file localization. Same coding model. Sharper search. Better result.

Every one of these points the same way. The bottleneck was never how smart the model is. It's what the model is forced to attend to. Multi-agent systems handle that by giving each agent a clean window and one focused job, and the gains are large, they're consistent, and once the mechanism clicks they stop being surprising at all.

---

If you're building agents that keep slamming into context limits, we make two tools for exactly this: [WarpGrep](/products/warpgrep), an RL-trained search sub-agent that lifts every major coding model to #1 on SWE-Bench Pro, and [Morph Fast Apply](/products/fast-apply), a specialized model that merges code edits at 10,500 tok/s without cluttering the lead agent's context. Both are built on the same idea: keep the frontier model's window clean by handing the specialized work to specialized models.
