---
title: "Coding Agents Fail at Search, Not Coding: 15 Papers Prove It"
url: "https://www.morphllm.com/blog/code-search-bottleneck"
description: "60% of coding agent time is spent searching, not coding. Bigger context windows make it worse. 15 papers from Anthropic, DeepMind, and Cognition explain why."
date: "2026-02-18"
author: "Tejas Bhakta"
---
# Coding Agents Fail at Search, Not Coding: 15 Papers Prove It

The usual instinct is that a better coding agent means a better model under it, one with more context and more reasoning behind a bigger frontier. Buy up the stack and the agent rides along.

The gains aren't really there, though, and the literature keeps saying so. I read most of what came out from 2024 into early 2026, from Anthropic and Google DeepMind and Cognition and Stanford and a scatter of independent labs, and it all keeps circling back to the same point. These agents can write code. What they can't do dependably is locate the code they're meant to be writing against. The rest of this post is the paper trail for that claim.

## Agents spend 60% of their time searching, not coding

Cognition, the team behind Devin and Windsurf, has the cleanest number on it. They went back through agent trajectories across both products and found [agents spending upward of 60% of their first turn just pulling in context](https://cognition.ai/blog/swe-grep) before they wrote a single edit. And that wasn't some pathological tail. It was the median across their production workloads.

The reason shows up the moment you watch agentic search run. The model issues grep and file-read calls one at a time, and it takes something like 10 to 20 serial turns before it even has enough context to start the actual task. Each turn is a round trip: read the results, pick the next search, read those results. Each turn also drops fresh tokens into the window, and the model then lugs all of them around for the rest of the session.

Cerebras described [the same wall](https://x.com/CerebrasSystems/status/1978874694825840679). "Context retrieval has been one of the biggest bottlenecks in agentic coding. When you ask an agent to work on a large codebase, it can spend 60% of its time just searching for relevant files."

One [OpenReview study](https://openreview.net/forum?id=1bUeVB3fov) followed the money and found input tokens dominating the bill even with caching on. Comparable tasks could differ by 10x in tokens consumed, and almost the entire gap traced to search quality. The authors couldn't even predict a run's total token count going in, Pearson's r under 0.15, because search efficiency is that erratic from one run to the next. So the bulk of what you pay to run a coding agent goes to search rather than to generation.

## More context makes the model worse

The tempting fix is to make the window bigger and pour everything in. The research is blunt about why that fails.

Start with Liu et al.'s [Stanford paper](https://arxiv.org/abs/2307.03172), published in TACL 2024 and cited everywhere since. LLM performance drops by more than 30% when the relevant information moves from the start or end of the context into the middle. Accuracy follows a U-shaped curve. The model attends hard to the first and last tokens and poorly to everything sitting in between.

For a coding agent that's brutal. Grep for a function name, read 8 files, hit the relevant code in file number four, and you've effectively buried that code in the model's blind spot.

Chroma pushed on this at scale. They [ran 18 LLMs](https://research.trychroma.com/context-rot), GPT-4.1, Claude 4, Gemini 2.5, Qwen 3 and more, across 8 input lengths and 11 needle positions. Models don't use their context evenly; performance gets less reliable as the input grows, even on tasks that are otherwise trivial. Needle-question similarity, distractors, the structure of the haystack itself, all of it drags accuracy down. Claude Opus 4 even started refusing outright at longer lengths, a 2.89% refusal rate. Their line for it: "What matters more than whether relevant information is *present* is how that information is *presented*."

And the cause of it is baked into the architecture. A transformer at 10,000 tokens is tracking 100 million pairwise relationships between them; at 100,000 tokens that figure is 10 billion. Attention is quadratic. So piling on context doesn't merely dilute the part that matters, it actively degrades the model's ability to attend to anything at all.

The rule that falls out is blunt. Search quality caps reasoning quality. Give the agent the right 50 lines and you get correct code. Give it 500 lines of might-be-relevant results and you get a hallucination.

## Where agents actually break on SWE-Bench

SWE-bench is how people benchmark coding agents on real GitHub issues, and several recent papers stopped to ask *where* in the pipeline the agents actually fall over. Retrieval, more or less every time.

Localization is a good place to start, because agents are better at it than you'd think. Majgaonkar et al.'s [study of agent trajectories](https://arxiv.org/abs/2511.00197), accepted at ICSE 2026, found agents naming the right problematic file in 72-81% of the attempts that *failed*. So they usually knew the neighborhood. Neighborhood isn't address, though. The agent would nail the file but miss the lines, or find one of the three files that all needed changing. Their failed runs also came out consistently longer and more variable than the wins. These weren't agents that couldn't write the fix. They were agents that spent too long searching, filled their own context with wrong results, and then edited off that bad context.

Flip that around and better localization buys better resolution outright. LocAgent ([ACL 2025](https://aclanthology.org/2025.acl-long.426/)) showed it cleanly, using graph-guided search on a fine-tuned Qwen-2.5-Coder-32B to hit 92.7% file-level localization accuracy and lift downstream GitHub issue resolution by 12%, without touching the code generation model. Caumartin et al.'s [query reformulation paper](https://arxiv.org/abs/2512.07022) from December 2025 went further and just rewrote the search query, no change to the coding model at all, and got 35% better first-file retrieval and 22% better file retrieval than SWE-agent. Same coding model. Better search. Better outcome.

The [SWE-Search paper](https://arxiv.org/abs/2410.20285) ([ICLR 2025](https://arxiv.org/abs/2410.20285)) wrapped Monte Carlo Tree Search around the solution-space exploration and got 23% relative improvement across five models on SWE-bench. Their phrasing is that you can move a coding agent 23% "without requiring larger models or additional training data." Better search, nothing else.

Agentless is the accidental version of the same point. Xia et al.'s [pipeline](https://arxiv.org/abs/2407.01489) scored 32% on SWE-Bench Lite for $0.70 an issue. Three stages: localize, repair, validate. Two of those stages are nothing. Repair is a diff generator you could write in an afternoon; validate just runs the tests. The whole system's intelligence lives in localization, which walks carefully down from file to class or function to the exact edit location. So the pipeline that invested everything in finding, and almost nothing in writing, is the one that set the bar.

## Why RAG doesn't work for code

Search being the bottleneck, retrieval-augmented generation looks like the answer. Embed the codebase, embed the query, take the nearest neighbors. On code it breaks in more than one place, though, and the first place is the ugliest.

Take the LIMIT benchmark. It has 50K documents in it, which is small. The best embedding models still couldn't push recall@100 past 20% on it. And BM25, which is just keyword matching and older than anything on the leaderboard, beat all of them anyway. Tuning won't fix this. Last August a [DeepMind team](https://arxiv.org/abs/2508.21038) worked out the reason. Every fixed embedding size has a saturation point, and once your corpus crosses it the vector can't fit any more of the query-document relationships you need it to. They connect the ceiling to something called sign-rank, from communication complexity theory.

And the ceiling is low. And the ceiling shows up early. A 512-dimensional embedding starts breaking down somewhere around half a million documents. Double the dimension to 1024 and you buy your way to maybe four million before it breaks the same way. Point four million at a monorepo, though, and it stops sounding like plenty. That's the reason embeddings look wonderful on a demo repo and fall over once they're in production, and why a stronger model does exactly nothing about it.

Code makes all of this worse, because a code query is not really a text query. Ask "where does the auth middleware check JWT expiration?" and you're actually asking about call graphs, import chains, where the middleware got registered, and the conventions of whatever framework this is. That's a chain of hops, and one vector has nowhere to put it. RAGFlow's [year-end writeup](https://ragflow.io/blog/rag-review-2025-from-rag-to-context) named the other half of the bind. Context wants big chunks, 1024 tokens and up. Precise matching wants small ones, 100 to 256. Code won't grant you both. A function body needs to come back whole, yet it has to match on one identifier buried inside it, so every setting is a compromise between fragmented-but-precise and whole-but-fuzzy.

Then there's staleness. Production codebases move constantly, and embeddings you computed yesterday may not describe today's code. Stale embeddings [cost up to 20%](https://medium.com/@yashtripathi.nits/when-embeddings-go-stale-detecting-fixing-retrieval-drift-in-production-778a89481a57) on downstream LLM tasks. For an active repo with a dozen people pushing daily, re-indexing is a real maintenance load and skipping it is a real accuracy hit. An [exploratory study of code retrieval](https://www.preprints.org/manuscript/202510.0924) from October 2025 put it plainly: "indexing an entire codebase with embeddings is seen as not only potentially unnecessary but also a security risk, leading some of the most prominent agent development teams to abandon RAG in favor of more direct, exploratory methods."

And look at what Anthropic actually ships. Claude Code, their flagship coding agent, runs no RAG whatsoever, just grep across repositories line by line. When the team with arguably the best models on the planet built their own agent, they reached for grep over embeddings.

## The architecture everyone is converging on

Search is the bottleneck, RAG doesn't fit code, so what does? Across the papers the answer keeps coming out the same. Dedicated search sub-agents, each in an isolated context window.

Anthropic has the biggest number of the bunch. In their [multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system) from June 2025, an Opus 4 lead delegating to Sonnet 4 sub-agents beat plain single-agent Opus 4 by 90.2% on their internal research eval. Mechanically it's not complicated. The lead spins up a handful of sub-agents (3-5 of them, running at once), each gets a clean window, and each one does its own searching and filtering and reports back only the surviving material for the lead to reason over. Anthropic sums up why it works in a line: "Multi-agent systems work mainly because they help spend enough tokens to solve the problem." The word doing the work there is *separate*. The tokens get spent in the sub-agents' windows, never the lead's, so the lead's reasoning context stays clean.

Notice what Cognition did with their own 60% number. A bigger coding model wasn't the move they made. They built [SWE-grep](https://cognition.ai/blog/swe-grep), a sub-agent whose only job is retrieving code, and it's fast in a way that matters: about 2,800 tokens a second, on the order of 20x Haiku's pace, while still matching frontier-model retrieval accuracy. It gets there in 4 turns, firing 8 tool calls in parallel on each. Cognition's framing of the whole thing is that "context retrieval sub-agents are the perfect hand-off point between a smart model and a fast model." One model works out what to look for. The other one goes and finds it.

[WarpGrep v2](/products/warpgrep) takes that further still. On code retrieval it beats SWE-grep, Haiku, and Sonnet 4.6. Pair it with a frontier coding model, whether that's Opus, Minimax, or Kimi, and the two together top SWE-bench Pro. The benchmark is the part I'd underline. SWE-bench Pro throws agents at production-scale codebases rather than the toy repos most benchmarks use.

You don't even need the full sub-agent to watch this play out. GrepRAG ([ISSTA 2026](https://arxiv.org/abs/2601.23254)) just bolted some cheap post-processing onto agentic grep, identifier-weighted re-ranking and structure-aware dedup, and that alone beat state-of-the-art by 7-15% in code exact match on CrossCodeEval. Same retrieval model, smarter pipeline. And Augment Code [got to the top of SWE-Bench Verified](https://jxnl.co/writing/2025/09/11/why-grep-beat-embeddings-in-our-swe-bench-agent-lessons-from-augment/) with grep and find and zero embeddings, letting the agent's stubbornness across many search turns paper over the simpler tools. Their own caveat is worth keeping in mind: SWE-bench repos are small, and an enterprise codebase is another animal entirely.

## What this means for agent infrastructure

Stack the findings up and they lean one direction.

Search takes north of 60% of an agent's resources. Cognition measured it across production workloads and outside researchers reproduced it, so it's not a guess. When agents fail, the cause is context rot and not some shortfall in raw ability. The same model that handles a problem cleanly falls apart once its window is packed with irrelevant search results. RAG won't save you at scale, since DeepMind's proof puts a hard mathematical ceiling on how much code-query complexity an embedding can hold. Sub-agent isolation is what actually works, and the receipts pile up quickly: Anthropic's 90% off multi-agent architecture, Cognition's SWE-grep, WarpGrep v2 at #1 on SWE-bench Pro, SWE-Search's 23% from better search and nothing else. Those wins carry downstream into the code too. That's the story behind LocAgent's 12% lift, query reformulation's 35% first-file gain, and Agentless keeping pace on a bare localize-then-repair pipeline.

The frontier of AI coding was never bigger models. It's better search.

---

This is where [WarpGrep](/products/warpgrep) comes from. It's a search sub-agent living in its own context that uses RL-trained parallel search to turn up the relevant code in 3.8 steps and returns only the precise file spans the coding model needs. Nothing embedded, nothing indexed, no context rot to inherit.

[WarpGrep v2](/products/warpgrep) shipped on February 23, 2026. It beats SWE-grep, Haiku, and Sonnet 4.6 at code retrieval, and it lifts model performance on [SWE-bench Pro](https://www.swebench.com/) (real production-scale codebases) across the board. Run it alongside Opus, Minimax, or Kimi and the pairing lands at #1 on SWE-bench Pro.

The models can already write the code. First they have to find it.

---

<details>
<summary>References (15 papers)</summary>

The measurements on how agents spend their turns and their tokens: Cognition AI, "Introducing SWE-grep and SWE-grep-mini" ([cognition.ai/blog/swe-grep](https://cognition.ai/blog/swe-grep), 2025); "How Do Coding Agents Spend Your Money?" on OpenReview ([openreview.net/forum?id=1bUeVB3fov](https://openreview.net/forum?id=1bUeVB3fov), 2025); and Hrubec, "Reducing Token Usage of Software Engineering Agents" (TU Wien, 2025).

On context degradation: Liu et al., "Lost in the Middle: How Language Models Use Long Contexts" (TACL, 2024, [arxiv.org/abs/2307.03172](https://arxiv.org/abs/2307.03172)); and Hong, Troynikov, and Huber at Chroma, "Context Rot: How Increasing Input Tokens Impacts LLM Performance" ([research.trychroma.com/context-rot](https://research.trychroma.com/context-rot), 2025).

On where agents fail and what fixes it: Majgaonkar et al., "Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories" (ICSE 2026, [arxiv.org/abs/2511.00197](https://arxiv.org/abs/2511.00197)); Chen et al., "LocAgent: Graph-Guided LLM Agents for Code Localization" (ACL 2025, [aclanthology.org/2025.acl-long.426](https://aclanthology.org/2025.acl-long.426/)); Caumartin et al., "Reformulate, Retrieve, Localize: Agents for Repository-Level Bug Localization" ([arxiv.org/abs/2512.07022](https://arxiv.org/abs/2512.07022), 2025); "SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search" (ICLR 2025, [arxiv.org/abs/2410.20285](https://arxiv.org/abs/2410.20285)); and Xia et al., "Agentless: Demystifying LLM-based Software Engineering Agents" ([arxiv.org/abs/2407.01489](https://arxiv.org/abs/2407.01489), 2024).

On why RAG struggles with code: Weller, Boratko, Naim, and Lee at Google DeepMind, "On the Theoretical Limitations of Embedding-Based Retrieval" ([arxiv.org/abs/2508.21038](https://arxiv.org/abs/2508.21038), 2025); and "An Exploratory Study of Code Retrieval Techniques in Coding Agents" ([preprints.org/manuscript/202510.0924](https://www.preprints.org/manuscript/202510.0924), 2025).

On grep and sub-agent search: Anthropic, "How We Built Our Multi-Agent Research System" ([anthropic.com/engineering/multi-agent-research-system](https://www.anthropic.com/engineering/multi-agent-research-system), 2025); Wang et al., "GrepRAG: An Empirical Study and Optimization of Grep-Like Retrieval for Code Completion" (ISSTA 2026, [arxiv.org/abs/2601.23254](https://arxiv.org/abs/2601.23254)); and Flaherty and Liu at Augment Code, "Why Grep Beat Embeddings in Our SWE-Bench Agent" ([jxnl.co](https://jxnl.co/writing/2025/09/11/why-grep-beat-embeddings-in-our-swe-bench-agent-lessons-from-augment/), 2025).

</details>
