---
title: "Flash Compact: 33,000 tok/sec Context Compaction"
url: "https://www.morphllm.com/blog/compact-sdk"
description: "Flash Compact drops 50-70% of an agent's context at 33,000+ tokens/second while keeping every surviving line verbatim. Two modes: objective compaction strips filler with no guidance, query-based compaction weights keep/drop decisions against what the agent needs next."
date: "2026-03-07"
author: "Tejas Bhakta"
---
# Flash Compact: 33,000 tok/sec Context Compaction

By turn 200, an agent is dragging around roughly 800K tokens. Some of that is the load you'd expect, like the system prompt and the files it read and the history of the conversation itself. But a lot of it is just exhaust. Grep output nobody ended up needing. A test run that passed. A retry that worked on the second try. When we sat down and measured it, the filler came out north of 70% on average, and that held across all four of our benchmarks. Then we tried cutting it and hit something we didn't expect. You save the tokens, obviously, but the resolve rate on SWE-Bench goes up 2 points on top of that.

So the hard part was never deciding to cut. It's deciding what goes.

## Two modes of compaction

### Objective compaction

Objective compaction runs with no query at all. The model keeps or drops each line purely on how structurally load-bearing it looks. An import or some boilerplate or a test suite that passed all score low. A function signature scores high, and so does an error message or the point where a real decision got made. You never tell it anything about where you're headed next.

This is the mode for when there isn't a next step to aim at yet. Maybe the session's done and you're archiving it. Maybe you're squeezing a sub-agent's output before the coordinator has even decided what to ask. Maybe it's a general assistant and the user could reasonably go anywhere from here.

```typescript
const result = await compact.compact({
  input: agentContext,
  compressionRatio: 0.5,
  preserveRecent: 2,
});
```

Run it at `compressionRatio: 0.5` and what stays is the skeleton of the conversation, meaning the signatures and error messages and decisions plus whatever tool output is still fresh. What falls away is the repetitive bulk sitting under all that. The test suites that passed. The files you read twice. The imports and the retry loops that finally caught.

### Query-based compaction

Give it a `query` and the behavior shifts. The query is a short line saying what the agent needs to do next, and the model grades every line against it. Whatever bears on the query lives at a higher rate, and whatever doesn't gets cut harder than it would have.

```typescript
const result = await compact.compact({
  input: agentContext,
  query: 'fix the JWT validation bug in auth middleware',
  compressionRatio: 0.5,
  preserveRecent: 2,
});
```

You get the same verbatim guarantee here: every line that survives is character-for-character what you put in. What changes is the precision. Take a 100-line file read. Objective compaction might keep 30 of those lines. If the query happens to match what's in the file, that could jump to 50. And if the query is about something else entirely, it might drop to 10.

It reasons better downstream for a plain reason. The context reaching the frontier model has already been narrowed to the step it's about to take, so it works over signal rather than digging through noise to find it.

You don't need much of a query. "Fix auth middleware bug" is plenty, and so is "refactor database connection pooling." It's a signal, not a prompt. All it does is settle ties, telling the model which of two lines to keep when both look important in the abstract but only one of them matters for what you're doing right now.

## When to use which

| Scenario | Mode | Why |
|----------|------|-----|
| Pre-call compression (agent about to act) | Query-based | You know the next task. Weight context toward it. |
| Session archival | Objective | No next task. Preserve general structure. |
| Sub-agent output for a coordinator with a known objective | Query-based | Coordinator's objective filters each sub-agent's noise. |
| Sub-agent output before the coordinator decides what to do | Objective | No objective yet. Keep the structural signal. |
| General-purpose assistant between user messages | Objective | User might ask anything. Don't bias toward a specific topic. |
| Tool output after a known next step | Query-based | Drop the 90% of grep/test output irrelevant to the fix. |

## The API

It's a single endpoint and a single call, and what comes back is the compacted text plus the real usage stats.

```typescript
import { CompactClient } from '@morphllm/morphsdk/tools/compact';

const compact = new CompactClient({ morphApiKey: 'sk-...' });

const result = await compact.compact({
  input: agentContext,               // string or message array
  query: 'fix the JWT validation bug in auth middleware',  // omit for objective compaction
  compressionRatio: 0.5,            // keep ~50% of content
  preserveRecent: 2,                // last 2 messages untouched
});

// result.output        → compacted text (verbatim lines only)
// result.usage         → { input_tokens, output_tokens, compression_ratio, processing_time_ms }
// result.messages      → per-message results with compacted_line_ranges
```

`compressionRatio` sets how hard the cut is. At 0.5 you keep about half, at 0.3 about a third. The model reads this as a target rather than a hard rule, so it won't do something dumb like split a function signature off from its body just to land on the number.

`preserveRecent` walls off the last N messages so they never get compacted, since the most recent turns are nearly always the ones that matter. It defaults to 2.

`query` is optional. Pass it and you get query-based compaction; leave it off and you get objective. If you hand in a message array with no query, the model reads the last user message and figures the query out for itself. Hand in a raw string with no query and it just falls back to pure objective mode.

Anything you wrap in `<keepContext></keepContext>` survives no matter the mode or the ratio. That kept content still counts against your `compressionRatio` budget, though, which means everything around it gets squeezed a little harder to make the target.

## Production patterns

### 1. Compress before every LLM call (query-based)

Right before you send context to your main model, compact it against the next task. The main model ends up seeing fewer tokens, costing less, and reasoning over cleaner signal.

```typescript
import { CompactClient } from '@morphllm/morphsdk/tools/compact';
import Anthropic from '@anthropic-ai/sdk';

const compact = new CompactClient();
const anthropic = new Anthropic();

async function agentStep(messages: Array<{ role: string; content: string }>, task: string) {
  // Query-based: compact toward the next task
  const compacted = await compact.compact({
    messages,
    query: task,
    compressionRatio: 0.5,
    preserveRecent: 3,
  });

  const response = await anthropic.messages.create({
    model: 'claude-sonnet-4-20250514',
    max_tokens: 8192,
    messages: compacted.messages.map(m => ({
      role: m.role as 'user' | 'assistant',
      content: m.content,
    })),
  });

  return response;
}
```

The main model has no idea any of this happened. It just sees a shorter conversation, and it reads naturally because nothing got rewritten, only removed. The `(filtered N lines)` markers are there to tell it something was cut, so it can ask for a re-read if it later needs one.

On cost: say your agent runs 500K input tokens a call at $3/M. Compact that down to 250K and you've saved $0.75 on the call, which is $37.50 across a 50-call session. And the compact call barely registers next to that, since it runs at 33,000+ tokens/second for a tiny fraction of what the main model costs.

### 2. Multi-session memory (objective)

Keep compacted transcripts around between sessions. No query this time. The session's finished and you have no idea what the next one will need.

```typescript
async function endSession(
  sessionMessages: Array<{ role: string; content: string }>,
  sessionId: string,
) {
  // Objective: no query, preserve general structure
  const compacted = await compact.compact({
    messages: sessionMessages,
    compressionRatio: 0.3,     // aggressive for storage
    preserveRecent: 0,         // session is over, compact everything
  });

  await db.sessions.update({
    where: { id: sessionId },
    data: {
      compactedTranscript: compacted.output,
      originalTokens: compacted.usage.input_tokens,
      storedTokens: compacted.usage.output_tokens,
    },
  });
}

async function startSession(priorSessionIds: string[]) {
  const priorSessions = await db.sessions.findMany({
    where: { id: { in: priorSessionIds } },
    select: { compactedTranscript: true },
  });

  const priorContext = priorSessions
    .map(s => s.compactedTranscript)
    .join('\n---\n');

  return [
    { role: 'user', content: `Prior session context:\n${priorContext}` },
    { role: 'assistant', content: 'I have context from prior sessions. Ready to continue.' },
  ];
}
```

A 500K-token session drops to about 150K at `compressionRatio: 0.3`. That's tight enough that three past sessions fit inside a 128K window and still leave room for the conversation you're actually having.

### 3. Sub-agent output compression (query-based)

A coordinator spins up sub-agents, and each one comes back with 10-50K tokens of findings. But the coordinator has a specific objective in mind, so you run query-based compaction to filter every sub-agent's output against that objective.

```typescript
async function coordinatorStep(
  subAgentResults: Array<{ agent: string; output: string }>,
  objective: string,
) {
  // Query-based: compact each sub-agent's output toward the coordinator's objective
  const compactedResults = await Promise.all(
    subAgentResults.map(async ({ agent, output }) => {
      const result = await compact.compact({
        input: output,
        query: objective,
        compressionRatio: 0.4,
      });
      return { agent, output: result.output, ratio: result.usage.compression_ratio };
    }),
  );

  const coordinatorContext = compactedResults
    .map(r => `## ${r.agent} (${Math.round((1 - r.ratio) * 100)}% compressed)\n${r.output}`)
    .join('\n\n');

  return coordinatorContext;
}
```

Do the math and it adds up fast. Five sub-agents at 30K tokens each dumps 150K on the coordinator, enough to blow past the context limit somewhere around the second one. Run each through compaction at 0.4 and the whole pile drops to 60K, which the coordinator can take in on a single call.

The `query` is doing the heavy lifting in that case. A search sub-agent will happily hand back 40K tokens of code matches even when barely 5K of them touch the bug the coordinator actually cares about.

## Line ranges

Every message in the response comes back with a `compacted_line_ranges` field marking which lines were dropped. I mostly use it to debug. When an agent starts making mistakes right after a compaction, that field is the first place I check to see what it lost. The same field is what lets the agent go grab a specific range back from the source when it later needs one, instead of re-reading an entire file to recover a handful of lines.

```typescript
const result = await compact.compact({
  messages: conversationHistory,
  query: 'database migration',       // or omit for objective compaction
  includeLineRanges: true,           // default: true
  includeMarkers: true,              // adds "(filtered N lines)" markers, default: true
});

for (const msg of result.messages) {
  if (msg.compacted_line_ranges.length > 0) {
    console.log(`${msg.role}: removed lines`, msg.compacted_line_ranges);
    // [{ start: 15, end: 42 }, { start: 78, end: 95 }]
  }
}
```

## Benchmarks

We ran context compression head to head with RAG and LLM summarization on four coding benchmarks. Every other method forced a trade, buying you fewer tokens at the cost of worse answers. Compression was the one that refused the trade. It cut the tokens and the accuracy climbed anyway.

### SWE-Bench Verified

Fifty real GitHub issues. For each one the agent has to track down the bug, get its head around the surrounding codebase, and produce a patch that actually holds.

| Method | Resolve Rate | Total Tokens |
|---|---|---|
| Baseline (no compression) | 62.0% | 972K |
| RAG (4K chunks) | 50.0% | 771K |
| LLM Summarization | 56.0% | 794K |
| Token-level pruning (LLMLingua2) | 56.0% | 699K |
| **Context compression** | **64.0%** | **670K** |

### SWE-Bench Pro

The problems get harder here, the trajectories longer, the tool calls more numerous. It's the point where sloppy context management stops being survivable and starts deciding whether an agent system works at all.

| Method | Resolve Rate | Total Tokens |
|---|---|---|
| Baseline (no compression) | 40.0% | 1.4M |
| RAG (4K chunks) | 30.0% | 1.1M |
| LLM Summarization | 35.0% | 1.1M |
| Token-level pruning (LLMLingua2) | 34.0% | 980K |
| **Context compression** | **42.0%** | **950K** |

### Long Code Completion

Hand the model a code file 8x its training context and ask it to predict the next block. Scored on edit similarity (ES), where higher wins.

| Method | Edit Similarity | Compression Ratio |
|---|---|---|
| Baseline | 56.56 | 1.0x |
| RAG | 55.82 | 6.60x |
| LLM Summarization | 52.80 | 9.68x |
| Token-level pruning (LLMLingua2) | 55.96 | 8.47x |
| **Context compression** | **57.58** | **10.92x** |

### Long Code QA

Point it at a long codebase and ask questions about how the code behaves, how it's put together, and why.

| Method | Accuracy | Compression Ratio |
|---|---|---|
| Baseline | 55.89% | 1.0x |
| RAG | 55.86% | 5.87x |
| LLM Summarization | 56.37% | 6.53x |
| Token-level pruning (LLMLingua2) | 55.38% | 9.02x |
| **Context compression** | **58.71%** | **14.84x** |

Whichever benchmark you look at, compression lands in the upper-left of the chart, the highest accuracy at the lowest token count. And on SWE-Bench Pro, where the contexts run 40% longer and the tasks are harder, the gap only opens up wider. RAG sheds 10 points there and hardly saves you anything, while compression picks up 2 points and still cuts a third of the context.

## Numbers

- 33,000+ tokens/second processing speed
- 100K tokens compressed in under 2 seconds
- 50-70% compression at `compressionRatio: 0.5`
- 98% verbatim accuracy (surviving lines are character-identical)
- 0% hallucination risk (the model never generates new content)
- +2 points on SWE-Bench Verified resolve rate vs. uncompressed baseline
- +2 points on SWE-Bench Pro resolve rate vs. uncompressed baseline

It's fast enough to sit inline in front of every LLM call. A 500K-token context compacts in under 3 seconds, and the main model call behind it takes anywhere from 10 to 60. So you're adding less than 10% to the round trip and cutting the input tokens in half.

## Try it

There's a [compact playground](/dashboard/playground/compact) where you can paste in some text, set a query (or leave it empty to run objective compaction), and compact in one click. The metrics it shows you are the real API stats: input tokens, output tokens, compression ratio, processing time, throughput.

The SDK is [`@morphllm/morphsdk`](https://docs.morphllm.com/sdk/components/compact), and you can grab an API key at [morphllm.com/dashboard/api-keys](/dashboard/api-keys).

```bash
npm install @morphllm/morphsdk
```

---

Compaction is just one layer of the stack. [WarpGrep](/products/warpgrep) does the code search, [Fast Apply](/products/fast-apply) does the merging, and each one is a small specialized model that's good at exactly one job, which is what frees the frontier model up to spend its attention on reasoning. The case for building it this way is laid out in [Everything Is Models](/blog/everything-is-models).
