---
title: "The Bitter Lesson Applied: Why Coding Agents Need More Compute, Not More Cleverness"
url: "https://www.morphllm.com/blog/bitter-lesson"
description: "Rich Sutton's bitter lesson predicted that scaling compute beats hand-engineering. Seven years later, coding agents are proving him right. The fix isn't smarter models. It's specialized infrastructure."
date: "2026-03-16"
author: "Tejas Bhakta"
---
# The Bitter Lesson Applied: Why Coding Agents Need More Compute, Not More Cleverness

Back in 2019, Rich Sutton, one of the people who built reinforcement learning into a field, wrote a [short essay](http://www.incompleteideas.net/IncIdeas/BitterLesson.html) that ended up assigned reading at OpenAI, DeepMind, and pretty much every lab that takes itself seriously. The whole argument is one sentence:

> The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin.

He named it "The Bitter Lesson" because nobody in research wants it to be true. What you want to do, as a smart person, is pour your knowledge into the machine. Chess heuristics. Grammar rules. Features you hand-designed over months. And for a while that wins. Then compute catches up, some cruder method with a lot more of it walks in, and your careful work loses.

Chess is the one people always reach for. All that grandmaster intuition got encoded so carefully into the engines over the years, and then [brute-force search](https://en.wikipedia.org/wiki/Deep_Blue_(chess_computer)) that understood nothing about the game walked in and beat it anyway. If it were only chess you could call it a fluke. But the same thing happened to speech recognition when statistical models on raw audio made the hand-tuned linguistic rules look antique, and it happened to computer vision the year ImageNet-trained convolutional nets rendered a shelf of carefully designed features obsolete. That's the discomfort in Sutton's essay. Seventy years of looking, and the exception never turns up.

## The lesson, applied to LLMs

If you want the loudest confirmation of the bitter lesson in the history of the field, it's the language models. The jump from GPT-2 to GPT-4 wasn't some new understanding of how language works. The compute budget behind those models went from millions of dollars to billions, and that alone was enough to get us from one to the other.

There's a wrinkle Sutton's essay didn't get into, though. Once you've actually got a powerful general model, the interesting question isn't how to make it smarter. It's an economics question. How do you run it at scale without burning an absurd amount of money doing so.

The years from 2020 to 2024 were about scaling training. What we're in now is about scaling inference. [Deloitte reckons](https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/compute-power-ai.html) inference was already half of all AI compute in 2025 and gets to two-thirds in 2026, with inference-optimized chips alone becoming a $50 billion market this year. The rule Sutton wrote down hasn't budged through any of that. All that's really changed is where it catches you.

## Coding agents feel it first

Of everything running in production, coding agents are the cleanest test of the bitter lesson I know of. Their job is unforgiving in a very particular way. They work over codebases that run to millions of lines and have to land edits that are exact and usually spread across several files, all inside a loop tight enough that any latency at all just becomes a developer somewhere waiting on a spinner.

And here's the number that reframed the whole thing for me. Agents spend most of their first turn not coding but searching. Cognition, the Devin and Windsurf team, [measured](https://cognition.ai/blog/swe-grep) it at over 60% of that first turn going purely to retrieving context, before a single edit or line of reasoning. [Cerebras hit the same pattern](https://x.com/CerebrasSystems/status/1978874694825840679) on their own, separately.

And the obvious fix backfires. You'd assume more context helps the model find what it needs. It doesn't. Chroma ran [18 frontier models](https://research.trychroma.com/context-rot) through the test, GPT-4.1 and Claude Opus 4 and Gemini 2.5 in the mix, and every last one of them degraded as the input grew. The [Stanford "Lost in the Middle" paper](https://arxiv.org/abs/2307.03172) has the number for the worst case, where accuracy drops by more than 30% once the relevant fact is buried in the middle of the window rather than at either end.

So if search is where the time goes and stuffing more into context backfires, it stands to reason that search quality is doing a lot of the work on code quality. The papers bear that out. [SWE-Search](https://arxiv.org/abs/2410.20285) (ICLR 2025) got a 23% lift across five models by improving the search alone, with no bigger model and no extra training data. [LocAgent](https://aclanthology.org/2025.acl-long.426/) (ACL 2025) got 12% more issues resolved just by getting better at finding the right file.

The tempting reaction to all of this is to reach for a smarter model, one with a bigger window or a cleverer attention mechanism or a few more billion parameters. Sutton's whole essay is a warning against exactly that instinct, and the instinct does pay off for a while before it runs into the wall it always runs into. The move the lesson actually recommends is the opposite. Keep spending more compute, but pay attention to the shape you spend it in.

## Intelligence organizes into hierarchies

This is where the lesson stops being a training-run observation and turns into a design principle. The answer was never going to be one smarter model. It's several models, each tuned to a different compute profile and a different job.

Anthropic's own multi-agent setup [beat single-agent Opus by 90%](https://www.anthropic.com/engineering/claude-code-best-practices), and the reason isn't that the subagents were smarter than Opus. It's that the lead agent's context stayed clean. All the search noise, the dead-ends, the files that turned out not to matter, happened off in separate windows and never got a chance to pollute the reasoning model's working memory.

What convinced me this was real is how fast everyone landed on it at once. February 2026 got almost comical. Inside a single month Grok Build shipped 8 parallel agents and Windsurf shipped 5, Claude Code launched Agent Teams, Codex CLI wired in the Agents SDK, and Devin added parallel sessions. None of these teams was in a room together, and they all walked out with the same verdict: one model doing everything is the wrong shape for the problem.

That's the bitter lesson at the level of systems. The elegant answer, one brilliant model handling search and reasoning and editing inside a single context, loses to the graceless one: a stack of specialized models, each one burning compute on a narrow slice.

## Where the compute should actually go

The nice thing about the bitter lesson is that it tells you where to put your money. Not into making one model do all the jobs. Into making each layer of the stack as fast as it can be at the one job it has.

Take search. A reasoning model firing off grep calls one after another is the 2026 version of a chess engine leaning on grandmaster heuristics: it works, and it's leaving a lot on the floor. [WarpGrep](https://www.morphllm.com/blog/fast-context-rl-retrieval) fires 8 tool calls in parallel per turn, lands on the relevant code in 3.8 steps, and finishes a median codebase search in 5 seconds against 75 for the sequential way. On SWE-Bench Pro, bolting WarpGrep v2 onto frontier models lifted scores by 2.1 to 3.7 points while spending 17% fewer input tokens and costing 15.6% less.

Or code merging. When a frontier model rewrites an entire file to change three lines, every one of those wasted tokens is also degrading the reasoning that comes after through context rot. [Fast Apply](https://www.morphllm.com/fast-apply-model) is a 7B model that does nothing but merge code edits, served on custom CUDA kernels with speculative decoding, running at 10,500 tokens a second and pushing a 500-line file through in 0.8 seconds. On a scoped job, the purpose-built model just beats the general one.

Or compression. Twenty turns into a session, the window is thick with stale search results, edits that got superseded, exploration nobody needs anymore. [Flash Compact](/products/compact) cuts that by 50 to 70% at north of 33,000 tokens a second, keeping 98% of surviving text verbatim. It doesn't rewrite or summarize or invent, so there's no hallucination to worry about; it just mechanically strips the window back down to what the reasoning model should be looking at.

Underneath the search layer and the merge layer and the compression layer, it's the same bet each time. You stop trying to prompt-engineer one model into juggling all of it, and you build a dedicated piece of machinery for the job in front of you, one that keeps getting faster as the compute underneath it gets cheaper.

## The second bitter lesson

Sequoia's AI newsletter [pulled out](https://inferencebysequoia.substack.com/p/richard-suttons-second-bitter-lesson) what they called a second bitter lesson from Sutton, which is that the winners won't only scale compute. They'll build systems that keep learning and adapting as the world underneath them shifts.

For coding agents that turns the subagent architecture from a nice performance win into the only shape that can actually improve piece by piece. A better search model ships, you swap the search layer and leave everything else alone. Inference hardware gets twice as fast and every layer inherits it. A new code representation lands and the embedding layer picks it up without the reasoning model ever noticing. A single monolithic agent can't do any of that. A hierarchy can. That's Sutton applied to how you draw the boxes.

## Where this is heading

The trajectory isn't subtle. AI data-center capex is pegged at [$400 to $450 billion globally in 2026](https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/compute-power-ai.html) and pointed at a trillion by 2028, and most of that spend is inference, not training. The people building inference infrastructure, rather than cleverer prompts, are the ones standing on the right side of the lesson.

Sutton's own line was that we want AI that can discover the way we do, not AI that merely contains what we've already discovered.

For coding agents I read that as building the stack so each layer is free to find the best approach to its own job, with as much compute as that job needs, without getting boxed in by what happens to fit in one context window.

The lesson earns the word bitter because it's humbling. The thing that unlocks coding agents isn't some breakthrough in reasoning. It's plumbing. Faster search, faster apply, faster compression. More compute, put at the right layer, at the right moment.

That's the thing we're building at Morph.

---

<details>
<summary>References</summary>

- [Rich Sutton, "The Bitter Lesson" (2019)](http://www.incompleteideas.net/IncIdeas/BitterLesson.html)
- [Deloitte, "More compute for AI, not less" (2026)](https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/compute-power-ai.html)
- [Cognition, "SWE-grep" (2025)](https://cognition.ai/blog/swe-grep)
- [Chroma, "Context Rot" (2025)](https://research.trychroma.com/context-rot)
- [Liu et al., "Lost in the Middle" (Stanford/TACL 2024)](https://arxiv.org/abs/2307.03172)
- [SWE-Search, ICLR 2025](https://arxiv.org/abs/2410.20285)
- [LocAgent, ACL 2025](https://aclanthology.org/2025.acl-long.426/)
- [Anthropic, "Claude Code Best Practices" (2025)](https://www.anthropic.com/engineering/claude-code-best-practices)
- [Sequoia, "Richard Sutton's Second Bitter Lesson"](https://inferencebysequoia.substack.com/p/richard-suttons-second-bitter-lesson)
- [CES 2026: AI compute shift from training to inference](https://www.computerworld.com/article/4114579/ces-2026-ai-compute-sees-a-shift-from-training-to-inference.html)
- [Majgaonkar et al., ICSE 2026](https://arxiv.org/abs/2511.00197)
- [Caumartin et al., Query Reformulation (2025)](https://arxiv.org/abs/2512.07022)
- [Xia et al., Agentless (2024)](https://arxiv.org/abs/2407.01489)
- [Weller et al., Google DeepMind Embedding Limits (2025)](https://arxiv.org/abs/2508.21038)

</details>
