---
title: "Compact: Context compaction at 33,000 tok/s"
url: "https://www.morphllm.com/products/compact"
canonical_url: "https://www.morphllm.com/products/compact"
docs_url: "https://docs.morphllm.com/sdk/components/compact"
description: "Custom inference engine that shrinks agent context 50-70% while keeping every surviving sentence byte-identical to the input. It deletes filler instead of summarizing. Enables 24+ hour agent sessions. Model: morph-compactor. 33,000 tok/s, under 3 seconds typical."
---

# Compact

Compresses agent context 50-70% by deleting filler. Every surviving sentence is byte-identical to the input, with no paraphrasing. Use when a long-running coding agent approaches its context limit, when web search results need shrinking before being passed to a downstream model, or when you want to run compaction inline before every LLM call instead of waiting for the 95% capacity cliff. Model: `morph-compactor`. 33,000 tok/s.

- Canonical: https://www.morphllm.com/products/compact
- Docs: https://docs.morphllm.com/sdk/components/compact
- Model spec: https://docs.morphllm.com/models/compact
- API reference: https://docs.morphllm.com/api-reference/endpoint/compact

## Problem

Agents hit a quality cliff when compaction triggers at 95% context capacity: they contradict earlier decisions, loop on solved problems, lose file paths. Summarization-based compaction (the default in most agent frameworks) paraphrases aggressively and loses precision; Factory's evaluation scored it 3.4–3.7/5 on accuracy. Compact deletes filler, keeps the rest verbatim.

## Why it's fast

33,000 tok/s comes from a custom inference engine and hand-written GPU kernels built for the compaction workload specifically: long-context attention with code-shaped sparsity, streaming batched decode. Not a fine-tune on off-the-shelf serving infrastructure.

## Quickstart

```ts
import { MorphClient } from "@morphllm/morphsdk";

const morph = new MorphClient({ apiKey: process.env.MORPH_API_KEY });

const compacted = await morph.compact.execute({
  messages: chatHistory,           // OpenAI message array
  objective: "implement retry logic in billing.ts", // optional, prompt-aware filtering
  targetReduction: 0.6,            // shrink to ~40% of original
});

console.log(compacted.messages);   // compacted chat history
console.log(compacted.stats);      // { inputTokens, outputTokens, ratio }
```

Also callable via OpenAI-compatible chat completions with `model: "morph-compactor"`. The compression is byte-identical, so the output passes to GPT-4/Claude/Gemini without re-serialization drift.

## Numbers

| Metric | Value |
|--------|-------|
| Speed | 33,000 tok/s |
| Latency (typical) | under 3 seconds |
| Context reduction | 50–70% |
| Claude Code native compact | ~90 seconds |
| Morph Compact | ~2.5 seconds |
| Input / Output $/1M | 0.20 / 0.50 |

## When to use

- Long-running agent sessions (4+ hours) where context grows unbounded
- Inline compaction before every LLM call to cap token spend
- Web search tool results: agents pull 10k+ tokens per page; shrink to the signal in <300ms
- Multi-session memory: store compacted transcripts and rehydrate on next session
- Pre-processing before sending to an expensive frontier model

## When not to use

- Short conversations (<2k tokens): no benefit
- Content that must preserve every token (legal, medical, exact quoting): Compact deletes what it judges filler
- Output that needs reformatting or translation (Compact does neither)

## Prompt-aware filtering

Pass the next objective. Compact keeps what's relevant to that objective and drops the rest. Omit the objective for general-purpose compaction.

## Patterns

- **Proactive, not reactive**: run Compact at 40–60% context, not at the 95% cliff
- **Before web search results**: compact tool output inline before appending to history
- **Cross-session memory**: persist compacted transcripts to disk, restore on session start
- **Multi-agent handoffs**: compact before sending context to a sub-agent

## See also

- [TypeScript SDK](https://docs.morphllm.com/sdk/quickstart)
- [API reference: compact endpoint](https://docs.morphllm.com/api-reference/endpoint/compact)
- [Authentication](https://docs.morphllm.com/auth)
- Blog: [Compact SDK](https://www.morphllm.com/blog/compact-sdk)
- Blog: [Long Running Agents](https://www.morphllm.com/blog/long-running-agents)
- Blog: [Coding Agent Harness Lessons](https://www.morphllm.com/blog/coding-agent-harness-lessons)
- Pricing: https://www.morphllm.com/pricing
