Claude Context Window Size (2026): 1M Tokens on Opus 5.5, Opus 5, Opus 4.8, and Sonnet 5

Claude context window size in 2026: Opus 5.5, Opus 5, Opus 4.8, Sonnet 5, and Fable 5.1 have a 1M-token context window at standard pricing with 128K max output; Haiku 4.5 has 200K. The claude.ai apps cap some models at 500K. Exact specs for every Claude model, Claude Code and claude.ai limits, prompt caching, and the research on effective context.

July 13, 2026 ยท 2 min read
Claude Context Window Size (2026): 1M Tokens on Opus 5.5, Opus 5, Opus 4.8, and Sonnet 5

Last updated September 22, 2026. Specs verified against Anthropic's models overview, pricing, and context-window docs on that date.

Claude's context window size is 1 million tokens on Opus 5.5, Opus 5, Opus 4.8, Sonnet 5, and Fable 5.1, and 200,000 tokens on Haiku 4.5. The 1M window needs no beta header and bills at standard per-token rates. Every 1M model outputs up to 128K tokens per request; Haiku 4.5 outputs 64K. The claude.ai chat app caps older models at 500K.

1M
Context window: Opus 5.5, Opus 5, Opus 4.8, Sonnet 5, Fable 5.1 (tokens)
200K
Context window: Haiku 4.5 (tokens)
128K
Max output per request on 1M models (300K on Batch API)
500K
claude.ai chat cap for Opus 4.8, Sonnet 4.6, Fable 5

Every Claude Model's Context Window (Opus 5.5, Opus 5, Sonnet 5, Fable 5.1, Haiku 4.5)

Anthropic's current lineup is Fable 5.1, Opus 5.5, Sonnet 5, and Haiku 4.5. Three of the four have 1M tokens. Opus 5.5 is the newest: it launched September 22, 2026, and replaced Opus 5 (July 24, 2026) as the default Opus. Fable 5.1 launched September 1, 2026. Everything from Claude 4.6 on, except Haiku, carries a 1M window with a 128K synchronous output cap. Current API IDs are claude-opus-5-5, claude-fable-5-1, claude-sonnet-5, and claude-haiku-4-5. Anthropic's Opus 5.5 post says Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks.

Claude Model Context Windows (September 2026)
ModelContext WindowMax OutputInput $/MTokOutput $/MTokStatus
Claude Opus 5.51M128K$4$20Latest, Sep 22 2026
Claude Fable 5.11M128K$10$50Latest, Sep 1 2026
Claude Sonnet 51M128K$2$10Current
Claude Haiku 4.5200K64K$1$5Current, retires no sooner than Oct 15 2026
Claude Mythos 5.11M128K$10$50Project Glasswing only
Claude Opus 51M128K$5$25Legacy, Jul 24 2026
Claude Fable 51M128K$10$50Legacy
Claude Opus 4.81M128K$5$25Legacy
Claude Opus 4.71M128K$5$25Legacy
Claude Opus 4.61M128K$5$25Legacy
Claude Sonnet 4.61M128K$3$15Legacy
Claude Opus 4.5200K64K$5$25Legacy
Claude Sonnet 4.5200K64K$3$15Legacy, retires no sooner than Sep 29 2026

For every 1M model, 1M is the default. You do not send a beta header, and a 900K-token request bills at the same per-token rate as a 9K one. On the Message Batches API, Opus 5.5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, and Sonnet 4.6 can return up to 300K output tokens with the output-300k-2026-03-24 beta header. Sonnet 5's $2/$10 launched as introductory pricing and is now the standard price; the scheduled September 1 move to $3/$15 was cancelled. Claude Opus 4.1 retired from the Claude API on August 5, 2026, and Opus 4 and Sonnet 4 retired on June 15. Sonnet 4.5 has no formal deprecation notice yet, but its earliest retirement date is September 29, 2026. Sources: Anthropic models overview, pricing, and context windows docs, September 22, 2026.

The same 1M window holds less text on newer models

Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text. Anthropic's docs put 1M tokens at about 555,000 words (2.5M Unicode characters) on the new tokenizer, versus about 750,000 words on models before Opus 4.7. 200K tokens is about 150,000 words. If you migrate a pipeline from Sonnet 4.5 or Opus 4.5 to Opus 5.5, re-run the token counting API on your largest prompts. A document that fit before may not fit now, and your per-request cost moves with the token count as well as the rate.

Two smaller limits sit beside the token limit. A single request can carry up to 600 images or PDF pages on a 1M model, but only 100 on a 200K model. And the Models API returns max_input_tokens and max_tokens for every model, so you can read these limits programmatically instead of hardcoding them.

Claude.ai and Claude Code Context Window Size

The API window is not the window you get in the Claude apps. Anthropic's support article on paid-plan context windows lists three separate tables: claude.ai chat, Claude Code, and Claude Cowork. The chat app halves the window for every model older than the current generation.

Claude Context Window by Surface (paid plans, September 2026)
ModelClaude APIclaude.ai chatClaude CodeClaude Cowork
Opus 5.51M1M1M1M
Fable 5.11M1M1M1M
Sonnet 51M1M1M1M (compacts at 500K)
Opus 51M1M1M1M
Fable 51M500K1M1M
Opus 4.81M500K1M1M
Opus 4.71M500K1M1M
Opus 4.61M500K1M ([1m] variant)200K
Sonnet 4.61M500K1M ([1m] variant)200K
Haiku 4.5200K200K200K200K

In Claude Code, Opus 4.6 and Sonnet 4.6 only get 1M if you pick the [1m] model with /model (for example claude-sonnet-4-6[1m]) and have usage credits enabled (Opus 4.6 needs them on Pro; Sonnet 4.6 on every plan except usage-based Enterprise). Opus 5 had a similar gotcha at launch. On July 25, 2026, Wade Tregaskis reported that Opus 5 defaulted to a 0.2M window on paid plans and that enabling usage credits restored 1M, with no spend required. That report is where the "Opus 5 context window 200K" searches come from. Anthropic's current table lists Opus 5 at 1M on every surface.

A Claude Code session is never empty. Before you type, the system prompt, your CLAUDE.md files, auto memory (the first 200 lines or 25KB of MEMORY.md), skill descriptions, and MCP tool names are already loaded, according to Anthropic's Claude Code context-window guide. After /compact, Claude Code re-reads up to five recently modified files and re-injects each invoked skill, capped at 5,000 tokens per skill. The skill listing itself does not reload. Anthropic's session-management post gives the simplest rule: start a new session for a new task.

Long sessions cost more per turn

One Claude user on Hacker News measured 144 sessions and 14,640 turns. Past turn 141, a turn cost 2.1x what it cost in the same session's first 20 turns. Past turn 180 it was 2.86x. At 2.1x, a fresh session that spends ten turns getting up to speed pays for itself after ten more. A 1M window lets a session run long. It does not make running long free.

What Happens When You Hit the Claude Context Window Limit

Everything counts toward the window: system prompt, every message, tool results, images, documents, tool definitions, and the output Claude generates on the turn, including thinking. Cached tokens count too. Caching changes what you pay for them, not whether they occupy space.

If the input alone exceeds the window, the API returns a 400 invalid_request_error ("prompt is too long") on every model. If input plus max_tokens exceeds the window on Claude 4.5 and newer models, the API accepts the request. Generation stops with stop_reason: "model_context_window_exceeded" if it reaches the limit. Older models return a validation error instead, unless you send the model-context-window-exceeded-2025-08-26 beta header.

For conversations that keep growing, Anthropic offers server-side compaction in beta for Claude 4.6 and later models. It summarizes earlier turns on the server so the conversation continues past the limit. Since September 14, 2026, a separate beta (compact-2026-09-04) lets you trigger compaction on demand with a top-level compaction parameter, which returns a signed compaction block. Context editing can also clear old tool results or thinking blocks. In claude.ai, paid plans with code execution enabled get automatic summarization of earlier messages near the limit.

Context Window vs Effective Context

A context window is how many tokens you can send. Effective context is how many of those tokens the model uses reliably. Anthropic's own docs say it directly: as token count grows, accuracy and recall degrade, a phenomenon known as context rot.

Chroma published one of the most cited studies of this gap on July 14, 2025, testing 18 models including Claude 4, GPT-4.1, Gemini 2.5, and Qwen3. Performance dropped as input length grew on every model, even on simple tasks.

Adobe Research's NoLiMa benchmark (arXiv 2502.05167) sharpens the point. It extends needle-in-a-haystack with needles that share almost no words with the question, so the model has to infer the link instead of string-matching. Of 13 models that claim at least 128K context, 11 dropped below 50 percent of their short-context baseline at 32K tokens. GPT-4o fell from 99.3 percent to 69.7 percent.

The Lost-in-the-Middle Problem

LLMs attend most strongly to tokens at the beginning and end of the context. Information in the middle gets less weight. Put the instructions and the question at the edges of a long prompt, and put the bulk material between them.

Chroma's study added a counterintuitive result: models performed better on shuffled haystacks than on logically structured ones. Coherent documents made retrieval harder, not easier.

Half
Of 17 models held up at 32K (NVIDIA RULER)
11 of 13
Models below half their base score by 32K (NoLiMa)
76%
Opus 4.6 MRCR v2 (8 needles, 1M tokens)
18.5%
Sonnet 4.5 on same MRCR v2 test

NVIDIA's RULER benchmark (arXiv 2404.06654) tested 17 long-context models across 13 tasks. All claimed at least 32K tokens; only half held satisfactory performance at 32K. Anthropic's own long-context number is from the Opus 4.6 launch: 76% on MRCR v2 with 8 needles across 1M tokens, against 18.5% for Sonnet 4.5. Anthropic called it "a qualitative shift in how much context a model can actually use while maintaining peak performance." The Opus 5.5 announcement published no new long-context benchmark.

Context rot comes from how attention works

Attention cost grows quadratically with sequence length, so every added token adds pairwise relationships the model has to track. Semantically similar distractors interfere with retrieval. Positional encoding degrades at distances rarely seen in training. A bigger window does not remove these effects. The useful question is which models degrade slowest on your task.

How Claude Compares to GPT-4.1, Gemini, and Others

Window size no longer separates the frontier. Every current Claude model except Haiku 4.5 sits at 1M, and open-weight models match it: Morph serves GLM-5.3 at 1M and DeepSeek V4 Flash at 1M context. GPT-4.1 is several generations old. OpenAI's current API models are the GPT-6 family.

Frontier Model Context Windows (September 22, 2026)
ModelContext WindowMax OutputLong-context pricing
Claude Opus 5.5 / Fable 5.1 / Sonnet 51M128KNone: 1M billed at standard rates
GPT-6 Astra / Sol / Luna1.05M128KOver 272K input: 2x input, 1.5x output on the whole request
Gemini 3.8 Flash1,048,57665,536One rate (introductory through Dec 31 2026)
Gemini 3.1 Pro Preview1,048,57665,536Over 200K: $4/$18 instead of $2/$12
Grok 4.7500KNot listedAt 200K+ prompt: $4/$12 instead of $2/$6 on all tokens
DeepSeek V4.1 Flash1M384KNone listed
Kimi K31,048,576Not listedNone listed

GPT-6 Astra launched September 3, 2026, and GPT-6 Sol and Luna followed on September 22. Gemini 3.8 Flash went GA on September 2. Grok 4.7 shipped September 21. DeepSeek V4.1 Flash shipped September 10 and replaced V4 Flash on DeepSeek's own API. Kimi K3 reached Moonshot's API in July. OpenAI, Gemini 3.1 Pro, and xAI charge more once a prompt passes a threshold. Claude bills its full 1M window at one rate. DeepSeek V4.1 Flash has the largest output cap at 384K. Grok 4.7 has the smallest window at 500K. Sources: OpenAI pricing, Google's Gemini 3.8 Flash and Gemini 3.1 Pro model pages, xAI models, DeepSeek, and Kimi pricing, September 22, 2026.

Size is close to a tie. The comparison that matters is behavior at length, and the best independent data is still the 2025 studies.

Independent Long-Context Studies
StudyPublishedModels testedFinding
Chroma context rotJul 14 202518, incl. Claude 4, GPT-4.1, Gemini 2.5Every model degrades as input grows; Claude lowest hallucination rate, GPT highest
Adobe NoLiMaFeb 202513 models claiming 128K+11 of 13 below half their short-context score at 32K
NVIDIA RULERApr 202417 models claiming 32K+Only half held satisfactory performance at 32K
Anthropic MRCR v2 (8-needle, 1M)Opus 4.6 launchOpus 4.6 vs Sonnet 4.576% vs 18.5%

Chroma's model-level results split by family. "Claude models consistently exhibit the lowest hallucination rates" and lean conservative under ambiguity, while "GPT models show the highest rates of hallucination." On the repeated-words task, GPT-4.1 refused 2.55% of attempts and Claude Opus 4 refused about 2.89%. Gemini 2.5 Pro was the most erratic, generating random text starting around 500 to 750 words.

Vendor benchmarks tend to favor the vendor. Needle-in-a-haystack scores measure retrieval of one planted fact, not reasoning across facts spread through the document. Run your own eval at the lengths you actually use.

Prompt Caching: 90% Cost Reduction (97.5% on Fable 5.1)

Prompt caching reuses a processed prompt prefix across API calls. The standard cache read costs 10% of base input. Two new models go lower: Opus 5.5 reads cache at 5% of base input ($0.20/MTok), and Fable 5.1 and Mythos 5.1 at 2.5% ($0.25/MTok). On a 1M-context agent that resends the same repository every turn, cache reads are most of the input bill, so the multiplier matters more than the headline rate.

0.1x
Cache read vs base input (most models)
0.05x
Cache read on Opus 5.5
0.025x
Cache read on Fable 5.1 and Mythos 5.1
2 calls
Break-even for 5-min cache (1 write + 1 read)
Prompt Caching Pricing per MTok (September 2026)
ModelBase input5-min write (1.25x)1-hour write (2x)Cache read
Opus 5.5$4$5$8$0.20
Fable 5.1$10$12.50$20$0.25
Sonnet 5$2$2.50$4$0.20
Opus 5 / Opus 4.8$5$6.25$10$0.50
Haiku 4.5$1$1.25$2$0.10

Anthropic's launch example: a 100K-token book prompt dropped from 11.5 seconds to 2.4 seconds with caching, at 90% lower cost. For a worked number on Sonnet 5, take 100 calls in an hour sharing a 50K-token prefix with the cache kept warm. Uncached, that prefix costs $10.00 per hour. Cached, it costs one $0.125 write plus 99 reads at $0.01 each, about $1.12. Opus 5.5 and Sonnet 5 now charge the same $0.20/MTok for cache reads, so on cache-heavy agents the gap between them is mostly output price. Caching multipliers stack with the 50% Batch discount and the 1.1x US-only inference_geo multiplier.

Prompt caching is context engineering

When repeated context is almost free, the budget question moves from "what can I afford to send" to "what should the model see." Cached tokens still occupy the window and still compete for attention, so a cheap 800K prefix can still hurt answer quality.

Extended Thinking and Context (Adaptive Thinking on Opus 5.5 and Sonnet 5)

Thinking tokens are part of max_tokens, bill as output, and count toward the window on the turn they are generated. What happens to them on the next turn depends on the model, and this changed with recent Claude generations.

On Opus 4.5 and later Opus models, Sonnet 4.6 and later Sonnet models, and the Fable and Mythos 5 models, the API keeps previous thinking blocks by default. They count toward the window like any other input and bill as input on later requests. On earlier Opus and Sonnet models and all Haiku models, the API strips previous thinking blocks automatically. If you were counting on stripped thinking to keep a long Opus conversation small, that assumption no longer holds. Use thinking block clearing through context editing to override the default.

Adaptive Thinking Replaces budget_tokens

Opus 5.5 and Fable 5.1 use adaptive thinking that is always on; Sonnet 5 and Opus 5 use adaptive thinking steered by the effort parameter. The manual budget_tokens mode is deprecated on Opus 4.6 and Sonnet 4.6 and not accepted on later models.

Context Awareness (Sonnet and Haiku only)

Sonnet 5, Sonnet 4.6, Sonnet 4.5, and Haiku 4.5 get an injected token budget and a remaining-tokens update after each tool call. Opus 4.7 and later Opus models and Fable/Mythos 5 do not; use task budgets (beta) instead.

Thinking cannot be turned off on Opus 5.5. Sending thinking: disabled, or enabled with budget_tokens, returns a 400. Opus 5 still allowed disabled at effort high or below. Opus 5.5 defaults to effort medium, while Fable 5.1, Sonnet 5, and Opus 5 default to high. Higher effort means more thinking tokens per turn. All 1M models support 128K output per request, and up to 300K on the Batch API with the output-300k-2026-03-24 header on the Opus 4.6+ and Sonnet 4.6+ models listed above.

How Much Code Fits in 200K Tokens (and 1M)

Anthropic's rule of thumb is about 4 characters or 0.75 words per token for English. Code is denser in symbols and whitespace, so treat these as estimates. Assuming about 40 characters per line, a line costs about 10 tokens on the older tokenizer and about 13 on the Claude 4.7+ tokenizer, which produces roughly 30% more tokens.

Code Capacity at Different Context Sizes (estimates)
Context SizeLines of Code (est.)Equivalent Project
32K tokens~2,500-3,200 linesA few large modules
128K tokens~10,000-13,000 linesSmall application or library
200K tokens~15,000-20,000 linesMedium service or framework package
1M tokens~75,000-100,000 linesLarge application; not most monorepos

A React project with 200 files averaging 75 lines (15,000 lines) fits in 200K. A 100,000-line codebase does not fit in 200K. On Opus 5.5 it runs to about 1.3M tokens at 13 tokens per line, so it does not fit in 1M either, before you add the system prompt, tool definitions, and output. Measure with the token counting API before you design around a number.

Cognition (makers of Devin) measured that their coding agent spent 60% of its time searching for code before making changes. The bottleneck was finding the right code to put in context, not the size of the window.

Context quality over context quantity

Putting 200K tokens of code into context when you need 5K tokens of relevant code produces worse results than sending the 5K. Context rot means more input tokens degrade output quality. The goal is to fill the window with the right content, not to fill it.

Context Engineering > Context Size

Windows went from 4K to 200K to 1M. Chroma's results show performance falling with length on all 18 models they tested. Bigger windows give you capacity, not quality.

Context engineering is choosing what goes into the window: the relevant code, the right instructions, and nothing that competes with them for attention. Anthropic's context-window docs make the same point: "more context isn't automatically better," and curating what is in context matters as much as how much space is available.

Practical strategies

Semantic Search Before Context Loading

Find the relevant code first, then load it. WarpGrep uses RL-trained search to reach 0.73 F1 in 3.8 steps, versus 12.4 steps for baseline agentic search, which cuts both latency and wasted context.

Prompt Caching for Stable Context

Cache system prompts, docs, and reference code. Pay the write once, then read at 10% of input, or 5% on Opus 5.5 and 2.5% on Fable 5.1.

Efficient Code Edits

Code edits are among the highest-token operations in an agent loop. Fast Apply merges edits at 10,500 tok/s with 98% accuracy, so the frontier model writes a short edit snippet instead of the whole file.

Subagent Isolation

Send search, testing, and analysis to subagents with their own context windows. The main agent gets back only the summary, so file reads made by the subagent never land in its window.

Context Compression

When a session grows past the point where quality holds, compress it instead of truncating. Morph Compact shrinks context 50-70% at 33,000 tok/s and keeps surviving sentences verbatim, so a 1.5-minute compaction pass takes about 2.5 seconds.

A 200K window filled with the right code beats a 1M window filled with a whole repository. The tooling that picks and shrinks context decides outcomes more than the window does. Morph Compact does the shrinking. If you want a raw long window on an open-weight model, Morph serves GLM-5.3 (1M) and DeepSeek V4 Flash (1M) through one OpenAI-compatible API.

Frequently Asked Questions

What is Claude's context window size?

On the Claude API, Claude Opus 5.5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, Sonnet 4.6, Fable 5.1, and Fable 5 have a 1 million token context window. Claude Haiku 4.5, Opus 4.5, and Sonnet 4.5 have 200,000 tokens. The 1M window needs no beta header and bills at standard per-token pricing. On the current tokenizer, 1M tokens is about 555,000 words; 200K is about 150,000 words. Source: Anthropic models overview and context windows docs, September 2026.

What is the Claude Opus 5.5 context window?

Claude Opus 5.5 (claude-opus-5-5, released September 22, 2026) has a 1M-token context window and a 128K max output on the Messages API, or 300K output on the Message Batches API with the output-300k-2026-03-24 beta header. It costs $4 per million input tokens and $20 per million output tokens, and cache reads cost $0.20, which is 5% of input instead of the usual 10%.

What is the Claude Opus 5 context window? Is it 200K?

On the API, Claude Opus 5 has a 1M-token context window and 128K max output at $5/$25 per million tokens. The 200K figure comes from the Claude apps: at launch in July 2026, paid-plan users reported Opus 5 defaulting to a 0.2M window until they enabled usage credits. Anthropic's support table now lists Opus 5 at 1M in claude.ai chat, Claude Code, and Cowork.

What is the claude.ai context window size on Pro and Max plans?

In claude.ai chat on paid plans, Fable 5.1, Opus 5.5, Opus 5, and Sonnet 5 get 1M tokens. Fable 5, Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6 get 500K. Other models get 200K. With code execution enabled, claude.ai summarizes earlier messages as a chat nears the limit, so long conversations can continue.

What is the Claude Code context window size?

Claude Code runs Fable 5.1, Fable 5, Opus 5.5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, and Sonnet 4.6 with a 1M-token window. For Opus 4.6 and Sonnet 4.6 you select the [1m] variant with /model (for example claude-sonnet-4-6[1m]) and need usage credits enabled (for Opus 4.6, only on Pro; for Sonnet 4.6, on every plan except usage-based Enterprise). Before your first prompt, the system prompt, CLAUDE.md, memory, skill listings, and MCP tool names are already in context.

What is the Claude Haiku 4.5 context window?

Claude Haiku 4.5 has a 200,000 token context window and a 64K max output at $1/$5 per million tokens. It is the only current Claude model without a 1M window. A 200K model accepts up to 100 images or PDF pages per request, compared with 600 on 1M models.

How does Claude's context window compare to GPT-4.1 and Gemini?

Raw window size is no longer the differentiator. Claude's current models all sit at 1M except Haiku 4.5. As of September 22, 2026, OpenAI's GPT-6 Astra, Sol, and Luna have 1.05M with 128K output, but prompts over 272K cost 2x on input. Gemini 3.8 Flash and Gemini 3.1 Pro Preview have 1,048,576 with 65,536 output. Grok 4.7 has 500K. DeepSeek V4.1 Flash has 1M with 384K output, and Kimi K3 has 1,048,576. GPT-4.1 is several generations old. Effective context is the deciding metric. Chroma's July 2025 context-rot study of 18 models, including Claude 4, GPT-4.1, and Gemini 2.5, found performance drops as input grows on every model; Claude models had the lowest hallucination rates and GPT models the highest.

What is prompt caching in Claude?

Prompt caching reuses a processed prompt prefix across API calls. A 5-minute cache write costs 1.25x base input and a 1-hour write costs 2x. Cache reads cost 0.1x base input on most models, 0.05x on Opus 5.5, and 0.025x on Fable 5.1 and Mythos 5.1. Anthropic's example 100K-token book prompt went from 11.5s to 2.4s with a 90% cost cut.

Does Claude's performance degrade with longer contexts?

Yes. Anthropic's own context-window docs call it context rot: as token count grows, accuracy and recall degrade. Claude Opus 4.6 scored 76% on the MRCR v2 8-needle test at 1M tokens, where Sonnet 4.5 scored 18.5%, but degradation still occurs as input grows on every model tested by Chroma, RULER, and NoLiMa.

How much code fits in Claude's 200K token context window?

Roughly 15,000 to 20,000 lines of code, assuming about 40 characters per line and Anthropic's rule of thumb of about 4 characters per token. On Claude 4.7 and later models the tokenizer produces about 30% more tokens, so the same window holds fewer lines. 1M tokens holds roughly 75,000 to 100,000 lines. Use the token counting API for an exact number.

What is the pricing for Claude's 1M token context window?

Claude 4.6 and later models bill the full 1M window at standard rates: a 900K-token request costs the same per token as a 9K one. Current rates per million input/output tokens: Opus 5.5 $4/$20, Opus 5 and Opus 4.8 $5/$25, Sonnet 5 $2/$10, Fable 5.1 $10/$50, Haiku 4.5 $1/$5. Batch is 50% off, and caching multipliers apply across the full window.

Related Reading

Better Context, Not More Context

WarpGrep finds relevant code in 3.8 steps (0.73 F1) so your agent fills its context window with signal, not noise. Fast Apply merges code edits at 10,500 tok/s. Both work with Claude, GPT, Gemini, or any LLM.