---
title: "The Long-Running Agent Era: Why Code Search and PR Review Are All That Matter"
url: "https://www.morphllm.com/blog/long-running-agents"
description: "As coding agents run for hours and days, two things change: agents need real code search to navigate, and human oversight moves from the IDE to the pull request."
date: "2026-02-10"
author: "Tejas Bhakta"
---
# The Long-Running Agent Era: Why Code Search and PR Review Are All That Matter

Something changed about coding work in the last year, which is that a decent chunk of it now runs for hours instead of minutes. People leave an agent porting a whole codebase while they sleep, or building a browser from scratch over a week, or grinding through a refactor that touches a thousand files and nobody wants to do by hand.

Cursor wrote up their [self-driving codebases](https://www.cursor.com/blog/towards-self-driving-codebases) work and the number that stuck with me was a system peaking at 1,000 commits per hour, across 10 million tool calls, over a single week. Around the same time Rakuten pointed Claude Code at vLLM, which is roughly 12.5 million lines, and the agent came back seven hours later with a complex feature done. This isn't a roadmap slide anyone's promising you. It already runs today.

It took me a while to internalize that a 30-minute agent and a 30-hour agent are not the same product with the timer turned up. They break in different places. Watch enough people run these things past the point they're comfortable with and you keep hitting the same two walls: the agent can't find the code, and you can't review what it wrote.

## Search is where the context rots

People assume the thing that kills a long run is the token limit, and it mostly isn't. What kills it is context rot, and context rot almost always starts the moment the agent goes looking for a file and can't find it cleanly.

You've probably felt this yourself. The window fills with junk and the agent loses the thread it was holding. The way [Nate's Newsletter](https://natesnewsletter.substack.com/p/i-read-everything-google-anthropic) tells it, some constraint from an early step ends up buried under everything the agent piled on later, which sounds right to me but only moves the question around. The junk still had to come from somewhere.

It came from searching. An agent on a big repo spends most of its time hunting around, and much of that hunting is dead loss, files it opened on a bad guess. There's a nice bit of bookkeeping in one of [Sankalp](https://sankalp.bearblog.dev/my-experience-with-claude-code-20-and-how-to-get-better-at-using-coding-agents/)'s writeups where he tracks a 50-tool-call session and finds the searching outweighing the code the agent actually wrote.

The arithmetic there is brutal. Sixty percent of your window goes to finding code, you reason with what's left, and after eight hours of thinking on forty percent of a brain, things start to slip. I remember one agent in particular that dropped a constraint it had clearly been holding onto around hour two, and my honest first reaction was that the model must be degrading somehow. It wasn't. The model was doing exactly what it always did. Its window had just quietly filled up with old search results that nobody, human or machine, was ever going to read a second time.

I don't really think the tools are at fault here either. Code search was built for people, and people tend to show up already half-knowing the answer. When I grep `handleSubmit`, I'm going on a hunch that it lives in a form handler off in some corner of the tree.

The agent has no hunch to run on. Picture the strong new hire on day one who hasn't opened a file. Give them `grep -r "validate"` and they're buried in matches with no way to tell which one matters. What they need instead is to ask in plain words where form validation happens and get the call graph back, and they need it quick, since on an overnight job every slow search is just window leaking away while the agent waits.

[WarpGrep](/products/warpgrep) is our answer to that. It's a search sub-agent that reads what the calling agent meant, ranks by relevance, and returns the code that matters instead of whole files, in under 6 seconds.

In an interactive session 6 seconds barely registers. Over a long run the savings compound. Run for 8 hours, search 200 times, and you've skipped thousands of lines you'd otherwise have read into the window. That difference is what keeps the agent lucid at hour seven instead of drifting off into rot by hour two.

[Shrivu Shankar](https://blog.sshh.io/p/how-i-use-every-claude-code-feature) found that letting the model read files directly, rather than lean on lossy summaries from explore agents, gives better reasoning because it "enables better pair-wise relationships and attention." Search quality sets reasoning quality. Hand the agent the right 50 lines and it writes clean code. Hand it 500 lines of maybe and it starts making things up.

The people getting real work out of long runs, Cursor's planner-worker setup, Anthropic's [harness patterns](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents), land on the same rule: show the agent only what matters. For code search, WarpGrep is that filter.

## So where does the human go?

Long-running agents force an awkward question. If the thing codes on its own for 8 hours, what are you doing during those 8 hours?

You don't leave the loop. You move to a different spot in it.

[Addy Osmani](https://addyosmani.com/blog/ai-coding-workflow/) treats the model as a powerful pair programmer that needs clear direction and oversight, which I think is right for the interactive case and just doesn't survive contact with an agent grinding away at 3am while you're asleep. There's nobody there to pair with. What you actually do is set it going before bed and read whatever came out in the morning.

Some people have taken this a good deal further. [Jesse Vincent](https://blog.fsck.com/2025/10/05/how-im-using-coding-agents-in-september-2025/) runs two separate sessions, one acting as architect that checks the design and one acting as implementer that writes the code, and the upshot is that his IDE isn't really where the work happens now. It's more like the table where those two sessions sit down together.

[Bored Hacking](https://boredhacking.com/coding-with-llms-2026/) has a warning post about the multi-thousand-line PR that nobody can review and nobody can safely roll back, and I regret to report it describes the completely normal output of a long agent run. A teammate's PR is a couple hundred lines that fit in my head, so I skim it, approve it, and go get lunch. The overnight agent handed me 2,000 lines across 47 files, and somewhere in the middle of counting hunks on the first one of those I admitted the old way was over. There was no holding it in my head and no clean undo if something had gone sideways forty files deep, so most of my attention has moved from writing code to checking the agent's, and the checking had to change shape too. I don't read these changes line by line anymore. I watch them run.

That watching part is what [Glance](/products/glance) does. When the PR lands, Glance reads the diff to figure out which corners of the app the agent disturbed, sends a browser agent to click through those corners the way a user would, and films the session. The recording is sitting in the PR by the time I open it, which means review starts with thirty seconds of the checkout form actually submitting rather than with me squinting at hunk seventeen of forty, trying to simulate the browser in my head.

I held out on trusting this longer than I should have, mostly out of some idea that a serious engineer reads every line. Then an agent spent a night porting a module from JavaScript to TypeScript across thirty files, and at 7am the PR had recordings of the migrated components rendering, the forms submitting, the error states firing. I signed off in ten minutes on a change that would otherwise have cost me two hours of diff archaeology. The question I answer in review has quietly changed, from whether the agent wrote correct code to whether the app still works, and with video in front of me the second question is just faster.

## The bugs a diff can't show you

Diff review has one blind spot it can't fix: code that reads right and behaves wrong. Our [RL-trained agent](/blog/browser-verification) is trained to go find those.

- Z-index bugs, where a component renders but something else sits on top of it.
- Dead handlers, where the button is right there in the diff but `onClick` never fires.
- Scroll traps, where the component exists and you can't get it on screen.
- Layouts that work on desktop and fall apart on mobile.
- Races between a user action and an async state update.

None of these show up in a diff. All of them are obvious in a 15-second clip. And they're the exact failures a long run generates: quiet integration bugs spread across files, invisible on any single diff line.

## What the workflow actually looks like now

The people getting the most out of long runs have settled into a pattern that has almost nothing to do with the old IDE-first way of working.

**Say what you want, not how to build it.** Cursor found [constraints beat instructions](https://www.cursor.com/blog/towards-self-driving-codebases). "No TODOs, no partial implementations" outperforms "remember to finish implementations." Write a spec, draw the boundaries, let the agent decide the rest.

**Give it real search.** [Craig Motlin](https://motlin.com/blog/claude-code-running-for-hours) kept an agent going past two hours by pushing verbose output into sub-agents. The deeper point is that an agent with good search spends less time lost and more time building. WarpGrep turns "find the auth middleware" from a 30-second multi-file grep hunt into one sub-6-second lookup.

**Let it run. Review the PR, not the process.** [Sankalp](https://sankalp.bearblog.dev/my-experience-with-claude-code-20-and-how-to-get-better-at-using-coding-agents/) says don't kick off a hard task mid-conversation. The flip side: don't hover over the agent either. Let it work, produce a PR, read the output.

**Review with your eyes, not only your head.** A 2,000-line diff needs deep focus and real expertise to review. A Glance video of the same change takes two minutes. Both help. Run them together and you keep quality at agent speed.

**Expect some mess.** Cursor found that demanding 100% correctness before every commit ground the system to a stop. "Workers would go outside their scope and start fixing irrelevant things." Better to accept a small error rate on the working branch and keep a green branch you fix up on a pass.

## The stack under a long run

Every long-running setup that works has three layers.

The inner loop is how fast the agent can search, edit, and check its work. This is where [Morph Fast Apply](/products/fast-apply) (200ms edits instead of 2-second edits, across 500 operations) and [WarpGrep](/products/warpgrep) (one sub-6-second search instead of dozens of grep calls) add up to hours saved.

The execution layer is the planner-worker split, the sub-agents, the isolated contexts. Cursor's research and people like [Shrivu Shankar](https://blog.sshh.io/p/how-i-use-every-claude-code-feature) and [Craig Motlin](https://motlin.com/blog/claude-code-running-for-hours) all land in the same place: isolate the workers, pass summaries up, keep the planner's context clean.

The review layer is where you come back in. As the agent takes over more of the execution, the PR becomes the one place you touch the work, and [Glance](/products/glance) makes that touch worth something by showing you what changed instead of just telling you.

## The IDE becomes optional

That's the direction. IDEs don't vanish. They're still good for a quick edit, a debug session, poking around. But the center of gravity for real software work is moving.

When the agent codes for 8 hours and you review for 20 minutes, the IDE isn't where the value gets made. The value is in the spec you wrote, the search that kept the agent on track, and the review tools that let you trust the output without reading all of it.

Agents keep getting better at running longer. The open question is whether the stuff around them, the search and the review and the fast apply, keeps up.

We think it will.
