LLM API Providers (2026): 12 APIs Compared by Price per 1M Tokens, Rate Limits, and Context

LLM API pricing for 12 providers, verified September 22, 2026: GPT-6 Astra $10/$50, GPT-6 Sol $2/$10, GPT-6 Luna $0.10/$0.50, Claude Opus 5.5 $4/$20, Fable 5.1 $10/$50, Gemini 3.8 Flash $0.75/$3.75, DeepSeek V4.1 Flash $0.30/$1.20 per 1M tokens. The cheapest LLM API, rate limits, free tiers, and context windows.

June 28, 2026 · 1 min read
LLM API Providers (2026): 12 APIs Compared by Price per 1M Tokens, Rate Limits, and Context

An LLM API is an HTTP endpoint for a hosted large language model: you POST a prompt to a URL such as /v1/chat/completions and pay per million input and output tokens. As of September 22, 2026, output prices on general models run from $0.50/M (GPT-6 Luna) to $50/M (Claude Fable 5.1, GPT-6 Astra). Most providers accept the OpenAI request format, so switching is a base_url change.

Quick answer

Cheapest current-generation LLM API: Morph's morph-glm53flash at $0.20/$0.70 per 1M tokens, then GPT-6 Luna at $0.10/$0.50, morph-dsv4flash at $0.142/$0.40, Qwen3.8 Flash at $0.15/$0.47, and GLM-5.3-Flash at $0.15/$0.50. DeepSeek's own V4.1 Flash is $0.30/$1.20 at peak and $0.15/$0.60 off-peak. Best value at the frontier: GPT-6 Sol and Claude Sonnet 5, both $2/$10. Top models: Claude Fable 5.1 ($10/$50), GPT-6 Astra ($10/$50), and Claude Opus 5.5 ($4/$20), which Anthropic released today as its recommended default. Most production systems route each request to the cheapest model that can handle it. See the Morph model router and the model list at Morph models.

What Is an LLM API

An LLM API is an HTTP interface to a hosted large language model. You send a prompt (plus optional tools, images, or files) to an endpoint such as /v1/chat/completions and receive generated tokens back, billed per million tokens (MTok) of input and output. It replaces running model weights on your own GPUs with a metered service: you run no infrastructure, manage no model updates, and pay only for tokens used.

Every provider on this page works that way. The differences are price, rate limits, context window, and model quality, and they are large. Output tokens cost between $0.50/M (GPT-6 Luna) and $50/M (Fable 5.1, GPT-6 Astra), a 100x spread. Seven models on the llm-stats tracker sit between 80.2% and 80.6% on SWE-bench Verified, and the four of them with list API prices span $1.20 to $12 per million output tokens. The cheapest rows are mostly open-weight models; the open source LLM guide covers their licenses and the point at which self-hosting beats any API rate.

100x
Output price spread ($0.50 GPT-6 Luna to $50 Fable 5.1 per MTok)
1.05M
GPT-6 context window (922K max input)
95.0%
Top SWE-bench Verified (Claude Fable 5, llm-stats)
$0
GLM-4.7-Flash API price (free)

What Changed in LLM API Pricing in September 2026

Every price below was re-checked on the provider's official pricing page on September 22, 2026. Since the September 2 revision of this page:

  • OpenAI GPT-6. GPT-6 Astra shipped September 3 at $10/$50. GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50) shipped September 22, half the price of GPT-5.6 Sol ($4/$20) and Luna ($0.20/$1.20). OpenAI's pricing page says GPT-5.6 Sol's promotional $4/$20 price runs at least through November 21, 2026. GPT-5.6 and GPT-6 also add a cache-write price column at 1.25x input.
  • Claude Opus 5.5. Released September 22 at $4/$20, down from Opus 5's $5/$25, with cache reads at 0.05x ($0.20/M). Anthropic now recommends it as the starting model and lists Opus 5 as legacy. Anthropic says it costs 40% less to run than Opus 5 on typical workloads and generates output more than 30% faster.
  • DeepSeek V4.1 Flash. Released September 10 under the model ID deepseek-flash at $0.30/$1.20 peak and $0.15/$0.60 off-peak. V4 Flash is retired.
  • Gemini 3.7 and 3.8 Flash. Google shipped Gemini 3.8 Flash on September 2. Gemini 3.6, 3.7, and 3.8 Flash are all $0.75/$3.75 through December 31, 2026, then $1.50/$7.50.
  • Anthropic rate limits. Anthropic replaced its four numbered tiers with Start, Build, and Scale on June 26, 2026. The entry tier allows 1,000 RPM on every current model, and Opus 5.5 gets its own bucket.
  • Z.AI. The GLM-5.3-Flash 50% promotion ended; it is $0.15/$0.50. GLM-5.3-FlashX was added at $0.37/$1.25. GLM-5.3 requests that send thinking.type: "disabled" fail; send enabled with reasoning_effort: low instead.
  • xAI Grok 4.7. Released September 21 at $2/$6 per 1M tokens ($4/$12 on long context) with a 500K context. xAI's docs recommend it for code. The 2x-priced Grok 4.7 Fast is only in Cursor and Grok Build, not the public API.
  • Retired model IDs. Moonshot retired kimi-k2.5 and every moonshot-v1 model in August 2026; they now return 404. Anthropic retired Claude Opus 4.1 on August 5. DeepSeek's deepseek-chat and deepseek-reasoner names were scheduled for discontinuation on July 24 and no longer appear on its pricing page.

LLM API Pricing Table: Flagship and Cheapest Model per Provider

Prices are $ per 1M tokens, input / output, standard (non-batch) tier, sorted by output price. Context is the maximum window where the provider publishes it.

ProviderModelInput/MTokOutput/MTokContext
AnthropicClaude Fable 5.1 / Fable 5$10.00$50.001M
OpenAIGPT-6 Astra$10.00$50.001.05M
OpenAIGPT-5.5$5.00$30.00long-context tier
AnthropicClaude Opus 5 / Opus 4.8 (legacy)$5.00$25.001M
AnthropicClaude Opus 5.5$4.00$20.001M
OpenAIGPT-5.6 Sol$4.00$20.00long-context tier
MoonshotKimi K3$3.00$15.001,048,576
AnthropicClaude Sonnet 4.6 (legacy)$3.00$15.001M
Morphmorph-kimik3 (Kimi K3, 16-bit)$2.50$14.001M
GoogleGemini 3.1 Pro (preview, ≤200K prompts)$2.00$12.00n/a
OpenAIGPT-5.6 Terra$2.00$12.00long-context tier
OpenAIGPT-6 Sol$2.00$10.001.05M
AnthropicClaude Sonnet 5$2.00$10.001M
GoogleGemini 3.5 Flash$1.50$9.00n/a
AlibabaQwen3.7 Max$2.50$7.501M
AlibabaQwen3.8 Max$2.00$6.001M
xAIGrok 4.7 (reference, released Sep 21)$2.00$6.00500K
AnthropicClaude Haiku 4.5$1.00$5.00200K
Z.AIGLM-5.3 / GLM-5.2$1.40$4.401M
MoonshotKimi K2.6$0.95$4.00262,144
DeepSeekV4 Pro (peak / off-peak)$1.32 / $0.66$3.96 / $1.981M
GoogleGemini 3.8 / 3.7 / 3.6 Flash (through Dec 31)$0.75$3.75n/a
Morphmorph-glm53-744b (GLM-5.3, 16-bit)$1.19$3.741M
DeepSeekV4.1 Flash (peak / off-peak)$0.30 / $0.15$1.20 / $0.601M
MiniMaxM3 (up to 512K input)$0.30$1.20priced to 1M
Morphmorph-dsv41flash (DeepSeek V4.1 Flash)$0.15$0.601M
OpenAIGPT-6 Luna$0.10$0.501.05M
Z.AIGLM-5.3-Flash$0.15$0.501M
AlibabaQwen3.8 Flash$0.15$0.471M
Morphmorph-dsv4flash (DeepSeek V4 Flash 0731, 16-bit)$0.142$0.401M
Morphmorph-glm53flash (GLM-5.3-Flash, 16-bit)$0.20$0.701M
Z.AIGLM-4.7-Flash / GLM-4.5-FlashFreeFreen/a

xAI is not one of the 12 providers profiled below; its Grok 4.7 row is there for price reference. Notes that change effective cost: GPT-6 bills 2x input and 1.5x output on prompts over 272K input tokens, so Astra becomes $20/$75, Sol $4/$15, and Luna $0.20/$0.75. GPT-5.6 has a similar long-context tier ($8/$30 Sol, $4/$18 Terra, $0.40/$1.80 Luna), and GPT-5.5 goes to $10/$45. Gemini 3.1 Pro rises to $4/$18 above 200K tokens. MiniMax M3 is $0.30/$1.20 up to 512K input and $0.60/$2.40 above, after a permanent 50% discount off list. DeepSeek bills peak rates 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays (excluding Chinese public holidays) and half price at all other hours. Anthropic charges the same per-token rate at any context length on Claude 4.6 and later (a 900K-token request bills like a 9K one), but Opus 4.7 and later use a tokenizer that produces about 30% more tokens for the same text.

Batch and cache discounts: OpenAI and Anthropic batch is 50% off. OpenAI cached input is 10% of base ($0.20/M on GPT-6 Sol, $1/M on Astra), and cache writes on GPT-5.6 and GPT-6 bill at 1.25x input. Anthropic cache reads are 0.1x base input, 0.05x on Opus 5.5 ($0.20/M), and 0.025x on Fable 5.1 ($0.25/M). GLM-5.3 cached input is $0.26/M. Kimi K3 cache hits are $0.30/M. DeepSeek V4.1 Flash cache hits are $0.006/M at peak, 50x below its $0.30 cache-miss rate. If your workload re-sends the same system prompt or file context, the cache column matters more than the headline price. On paper, a few older small models undercut everything above: Z.AI's GLM-4.7-FlashX is $0.07/$0.40 and GLM-4-32B-0414-128K is $0.10/$0.10.

Provider-by-Provider Breakdown

OpenAI

GPT-6 is the new top of OpenAI's price list in three sizes: Astra ($10/$50, September 3), Sol ($2/$10), and Luna ($0.10/$0.50), the latter two released September 22. All three list a 1,050,000-token context window with 922,000 max input and 128,000 max output. GPT-5.6 Sol ($4/$20), Terra ($2/$12), and Luna ($0.20/$1.20) and GPT-5.5 ($5/$30) stay on the list. There is no GPT-6 Terra. Fast mode (renamed from priority processing on July 30, 2026) costs 2x standard on GPT-6 and GPT-5.6 and 2.5x on GPT-5.5. Regional data residency adds 10% on models released on or after March 5, 2026, and GPT-6 EU data residency is Standard tier only. See Codex pricing for the subscription side.

ModelInput/MTokCached InOutput/MTokOver 272K input (In/Out)
GPT-6 Astra$10.00$1.00$50.00$20.00 / $75.00
GPT-5.5$5.00$0.50$30.00$10.00 / $45.00
GPT-5.6 Sol$4.00$0.40$20.00$8.00 / $30.00
GPT-5.6 Terra$2.00$0.20$12.00$4.00 / $18.00
GPT-6 Sol$2.00$0.20$10.00$4.00 / $15.00
GPT-5.6 Luna$0.20$0.02$1.20$0.40 / $1.80
GPT-6 Luna$0.10$0.01$0.50$0.20 / $0.75

A gotcha on long prompts: past 272K input tokens GPT-6 Sol costs $4/$15, while Claude Opus 5.5 is a flat $4/$20 at any length up to 1M. A Hacker News commenter on the Sol launch thread made the same point: once you pass 272K, Sol is roughly Opus-priced and Astra roughly Fable-priced. If your agent routinely sends 300K+ token prompts, model the long-context tier, not the headline.

Anthropic (Claude)

Anthropic's current lineup is Claude Fable 5.1 ($10/$50, released September 1, 2026), Claude Opus 5.5 ($4/$20, released September 22, 2026), Claude Sonnet 5 ($2/$10), and Claude Haiku 4.5 ($1/$5). Anthropic's models overview now says to start with Opus 5.5 for most workloads and use Fable 5.1 for demanding reasoning and long-horizon agentic work. Opus 5.5 reports 66.4% on Terminal-Bench 4.0, which Anthropic says matches GPT-6 Astra at about 40% of the cost, and 57.8% on CursorBench 4.0. Sonnet 5's $2/$10 is the standard price; the planned September 1 increase to $3/$15 was cancelled. Opus 5.5 defaults to medium effort (Opus 5 defaulted to high), so an integration that omits effort gets a different cost profile. Anthropic says Sonnet 5.5 and Haiku 5.5 follow in the coming weeks. Fable 5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Opus 4.5, Sonnet 4.6, and Sonnet 4.5 remain available as legacy models. Opus 4.1 was retired on August 5, 2026. Sonnet 4.5 can retire as early as September 29, 2026, and Haiku 4.5 as early as October 15, 2026. Mythos 5.1 ($10/$50) is limited availability. Fable 5.1, Opus 5.5, and Sonnet 5 carry a 1M context and 128K max output; Haiku 4.5 is 200K and 64K. Full breakdown: Anthropic API pricing.

ModelInput/MTokCache ReadOutput/MTokContext
Claude Fable 5.1$10.00$0.25$50.001M
Claude Fable 5 (legacy)$10.00$1.00$50.001M
Claude Opus 5 / Opus 4.8 (legacy)$5.00$0.50$25.001M
Claude Opus 5.5$4.00$0.20$20.001M
Claude Sonnet 4.6 (legacy)$3.00$0.30$15.001M
Claude Sonnet 5$2.00$0.20$10.001M
Claude Haiku 4.5$1.00$0.10$5.00200K

Google (Gemini)

Gemini 3.1 Pro is still labelled preview at $2/$12 for prompts up to 200K tokens ($4/$18 above), has no free tier, and scores 80.6% on SWE-bench Verified. The Flash line moved fast: Gemini 3.8 Flash (September 2), 3.7 Flash, and 3.6 Flash are all stable at a promotional $0.75/$3.75 through December 31, 2026, rising to $1.50/$7.50 after. Google positions 3.8 Flash for long-horizon software engineering and agents. Gemini 3.5 Flash ($1.50/$9) remains listed, and Gemini 3.5 Flash-Lite ($0.30/$2.50) and 3.1 Flash-Lite ($0.25/$1.50) are the budget tiers. Every Flash and Flash-Lite model has a free tier.

DeepSeek

DeepSeek V4.1 Flash (model ID deepseek-flash) launched September 10, 2026: an MoE with 552B backbone parameters (about 763B in total counting its 196B Engram memory and vision encoder), 8B active parameters for input and 16B for output, native image input, a 1M context, and 384K max output. It bills $0.30/$1.20 at peak and $0.15/$0.60 off-peak, with cache hits at $0.006 and $0.003. DeepSeek says its KV cache needs a quarter of the HBM of the previous generation and that it beats V4 Pro on DeepSeek's benchmarks. V4 Pro (deepseek-v4-pro) is listed at $1.32/$3.96 peak and $0.66/$1.98 off-peak. The API accepts OpenAI format at api.deepseek.com and Anthropic format at api.deepseek.com/anthropic. Deep dive: DeepSeek V4.

Gotcha: old DeepSeek model IDs now hit a different model

DeepSeek retired V4 Flash on September 10. Requests to deepseek-v4-flash still succeed, but they are served by V4.1 Flash and billed at the Flash price, so an eval pinned to the old ID silently changed models. The release note first said deepseek-v4-pro would also route to V4.1 Flash from September 14. DeepSeek reversed that on its updates page: V4 Pro stays on the API after September 14 with billing unchanged. The legacy deepseek-v4-flash-vision-exp ID also routes to V4.1 Flash. One HN reader who went through the model card described V4.1 Flash as tuned for agentic tool calling, cheap prefill on 8B active parameters, and moderate output, at some cost to knowledge. If you need the V4 Flash behavior you validated against, Morph still serves the DeepSeek-V4-Flash-0731 checkpoint as morph-dsv4flash.

Where you run an open-weight model changes its output and its price. Morph serves open-weight models (GLM-5.3, GLM-5.3-Flash, Kimi K3, DeepSeek V4 Flash 0731, DeepSeek V4.1 Flash) with 16-bit (bf16) activations and does not quantize them, so output matches the released weights. morph-dsv4flash is $0.142/M input and $0.40/M output with no peak surcharge, morph-dsv41flash is $0.15/$0.60, and morph-glm53-744b is $1.19/$3.74 with a 1M context. For coding agents, Morph adds codegen-tuned speculative decoding and custom inference kernels. See Morph models and pricing.

Moonshot (Kimi)

Kimi K3 ($3.00/$15.00, cache hits $0.30, cache writes $3.00 at the default 5-minute TTL or $6.00 at 1 hour) is Moonshot's flagship with a 1,048,576-token context. The API now lives at platform.kimi.ai. In August 2026 Moonshot retired kimi-k2.5 and all moonshot-v1 models, which now return 404, so pin a current ID. In September it added a Web Search API and began deducting usage 50% from cash and 50% from vouchers (contract customers excepted). Kimi K2.6 ($0.95/$4.00, cache hits $0.16, 262,144 context, 80.2% on SWE-bench Verified) and the coding-tuned Kimi K2.7 Code ($0.95/$4.00, or $1.90/$8.00 for the high-speed variant) remain available. Morph serves Kimi K3 as morph-kimik3 at $2.50/$14.00 with 16-bit activations. Deep dive: Kimi K3.

Z.AI (GLM)

GLM-5.3 ($1.40/$4.40, cached input $0.26/M, 1M context, 128K max output) is Z.AI's open-weights flagship, priced the same as GLM-5.2 and GLM-5.1. It takes text input only. Reasoning is mandatory at low, high, or max effort, with max the default: thinking.type accepts only enabled, and a request that sends disabled fails. Z.AI reports 34.5% task completion on its Code Bench at max effort using about 75K output tokens, against 23.4% at 96K tokens for GLM-5.2. GLM-5.3-Flash is $0.15/$0.50 now that its launch promotion has ended, and GLM-5.3-FlashX is $0.37/$1.25. GLM-4.7-Flash and GLM-4.5-Flash are free. Morph serves morph-glm53-744b at $1.19/$3.74 and morph-glm53flash at $0.20/$0.70 with 16-bit activations. Deep dive: GLM-5 family.

MiniMax

MiniMax M3 is $0.30/$1.20 up to 512K input tokens, a permanent 50% discount off the $0.60/$2.40 list price, and $0.60/$2.40 above 512K; cache reads are $0.06. It scores 80.5% on SWE-bench Verified, the cheapest 80%+ model on a hosted API: about one-twentieth of Opus 4.8's output price for a score 8 points lower. A priority tier runs 1.5x standard ($0.45/$1.80). MiniMax M2.7 ($0.30/$1.20) is the prior generation.

Alibaba (Qwen)

Qwen3.8 Max ($2.00/$6.00 in the Singapore region, up to 1M context) is Alibaba's proprietary flagship and undercuts Qwen3.7 Max, which lists at $2.50/$7.50 and scores 80.4% on SWE-bench Verified. Qwen3.8 Flash ($0.15/$0.47, 1M context) is the budget tier, built on the open Qwen3.8-Flash-Next architecture. Qwen3.6 Plus ($0.50/$3.00, tiered at 256K) remains available, and Qwen3.7 Plus is 20% off list for a limited time in Singapore. New Model Studio users get 1M free tokens valid 90 days. The Qwen API guide has model IDs, base URLs by region, and the open-weight self-hosting path.

Aggregators and Cloud Resellers

OpenRouter fronts many providers' models behind one OpenAI-compatible key, useful for evaluation and failover. Claude models are also sold through Amazon Bedrock and Google Cloud (billed by the cloud provider, with a 10% premium on regional endpoints) and through Claude Platform on AWS and Claude in Microsoft Foundry, which bill Anthropic's standard rates through the cloud marketplace. OpenAI models have been on Amazon Bedrock since June 1, 2026, at the same price as direct. If you need a routing layer rather than a reseller, see LLM gateways and LLM routers.

Morph (specialized)

Beyond the open-weight chat models above, Morph serves task-specific models: Fast Apply merges code edits at 10,500 tok/s, and WarpGrep does agentic codebase search at $0.80 per 100K tokens. Covered in Specialized APIs below.

LLM API Rate Limits by Provider and Tier

Rate limits decide whether your launch survives traffic, and most comparison pages skip them. Here is what each provider enforces, from official docs as of September 22, 2026.

Anthropic: Start, Build, and Scale tiers, cache-aware token counting

Anthropic replaced its numbered tiers with Start ($500 monthly spend cap), Build ($1,000), and Scale ($200,000), plus a Custom tier with no cap. Organizations move up automatically with usage history; new organizations may start on an Evaluation tier below these limits. Limits are per model, in requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Cache reads do not count toward ITPM: with a 2M ITPM limit and an 80% cache hit rate you can process 10M total input tokens per minute.

ModelStart (RPM / ITPM / OTPM)BuildScale
Claude Fable 5.x (combined)1,000 / 500K / 100K2,000 / 1.5M / 300K4,000 / 4M / 800K
Claude Opus 5.51,000 / 2M / 400K5,000 / 5M / 1M10,000 / 10M / 2M
Claude Opus 51,000 / 2M / 400K5,000 / 5M / 1M10,000 / 10M / 2M
Claude Opus 4.x (combined)1,000 / 2M / 400K5,000 / 5M / 1M10,000 / 10M / 2M
Claude Sonnet 51,000 / 2M / 400K5,000 / 5M / 1M10,000 / 10M / 2M
Claude Haiku 4.51,000 / 2M / 400K5,000 / 5M / 1M10,000 / 10M / 2M

Two things to plan around. Fable 5.x gets a quarter of the entry-tier input throughput of every other current model (500K vs 2M ITPM), shared across Fable 5.1 and Fable 5. Opus 5.5 and Opus 5 each have their own bucket, separate from the combined Opus 4.x bucket, so spreading traffic across Opus versions raises your effective ceiling. Hitting the monthly spend cap returns HTTP 429 with the same rate_limit_error type as a rate limit but no retry-after header and error.details.error_code set to enforced_spend_limit_reached. SDK auto-retries keep failing until 00:00 UTC on the first of the next month, so check that field before you back off.

OpenAI: spend-unlocked tiers, per-model limits

TierQualificationMonthly usage capGPT-6 Astra RPM / TPM
FreeAllowed geography$100not available
Tier 1$5 paid$100500 / 500K
Tier 2$50 paid$5005,000 / 1M
Tier 3$100 paid$1,0005,000 / 2M
Tier 4$250 paid$5,00010,000 / 4M
Tier 5$1,000 paid$200,00015,000 / 40M

Per-model RPM/TPM numbers live on each model's page. GPT-6 Sol and Luna also start at 500 RPM and 500K TPM on Tier 1, with batch queue limits of 1.5M and 5M tokens.

DeepSeek: concurrency caps instead of token budgets

DeepSeek publishes no RPM/TPM limits. It caps concurrent requests: 2,500 in flight on deepseek-flash, 500 on deepseek-v4-pro. For batch-style pipelines this is friendlier than token-per-minute budgets; for bursty single requests it makes no difference.

Others

Moonshot, Z.AI, MiniMax, and Alibaba show per-account limits in their consoles. MiniMax sells a priority tier at 1.5x the standard price for latency-sensitive traffic. Bedrock and Google Cloud enforce cloud-account quotas that you raise through the cloud provider.

OpenAI-Compatibility Matrix

Most providers accept OpenAI-format /v1/chat/completions requests, so switching is a base_url and api_key change. The exceptions matter when you build against provider-specific features.

ProviderOpenAI chat formatAnthropic formatStreamingTool calls
OpenAINativeNoYesYes
AnthropicVia OpenAI SDK compat layerNativeYesYes
Google GeminiCompat endpointNoYesYes
DeepSeekYesYes (/anthropic)YesYes
Moonshot (Kimi)YesNoYesYes
Z.AI (GLM)YesNoYesYes
MiniMaxYesNoYesYes
Alibaba (Qwen)Compat modeNoYesYes
OpenRouterNative (aggregator)NoYesYes
MorphYesNoYesYes (chat models)

DeepSeek is the first-party provider that documents both OpenAI and Anthropic request formats (https://api.deepseek.com and https://api.deepseek.com/anthropic), which lets it drop into Claude-Code-style agents without a proxy. If you point a coding agent at a custom provider, the agent must support it; Codex provider configuration covers the Codex side.

Benchmarks vs Price: What You Get per Dollar

SWE-bench Verified (real GitHub issues, verified fixes) is the most cited coding benchmark. Scores below are from the llm-stats tracker on September 22, 2026; prices are official API rates. Claude Fable 5.1, Opus 5.5, Opus 5, GPT-6, GPT-5.6, DeepSeek V4.1 Flash, Gemini 3.8 Flash, Kimi K3, GLM-5.3, and Qwen3.8 Max have no entry on the tracker yet.

ModelSWE-bench VerifiedInput/MTokOutput/MTokValue ($/MTok out per point)
Claude Fable 595.0%$10.00$50.00$0.53
Claude Mythos Preview (limited)93.9%restrictedrestrictedn/a
Claude Opus 4.888.6%$5.00$25.00$0.28
Claude Opus 4.787.6%$5.00$25.00$0.29
Claude Sonnet 585.2%$2.00$10.00$0.12
DeepSeek V4 Pro Max80.6%open weightsopen weightsopen weights
Gemini 3.1 Pro80.6%$2.00$12.00$0.15
MiniMax M380.5%$0.30$1.20$0.015
Qwen3.7 Max80.4%$2.50$7.50$0.09
Kimi K2.680.2%$0.95$4.00$0.05
DeepSeek V4 Flash Max79.0%open weightsopen weightsopen weights

An independent run tells a different story. Vals.ai scores SWE-bench Verified with one bash-only harness and archived the benchmark on September 1, 2026, because the top models saturated it. Its top entries: Claude Opus 5 at 97.0%, DeepSeek V4 Pro 0813 at 96.4%, GPT-5.6 Sol at 96.2%, GLM-5.3 at 95.4%, Claude Fable 5 at 95.0%, and Kimi K3 at 93.4%. On that harness MiniMax M3 scores 75.0%, below its self-reported 80.5%. DeepSeek V4.1 Flash shipped after the archive date and has no score.

The 80% club spans a 10x price range

Seven models on the tracker score between 80.2% and 80.6% on SWE-bench Verified. Of the four priced in the table, output runs from $1.20/M (MiniMax M3) to $12/M (Gemini 3.1 Pro), and DeepSeek V4 Pro Max ships the same score in open weights. Claude Sonnet 5 adds 4.6 points for $10/M output. The step to Opus 4.8 (88.6%) costs $25/M, and Fable 5 (95.0%) costs $50/M, a 42x step from MiniMax M3. Whether those points are worth it depends on whether your tasks live in the gap.

Newer models are reporting on different benchmarks. Anthropic reports Opus 5.5 at 66.4% on Terminal-Bench 4.0 and 57.8% on CursorBench 4.0 and says it matches GPT-6 Astra on Terminal-Bench 4.0 at about 40% of the cost. For the harder, standardized-harness view, see SWE-bench Pro, and compare scores only within the same harness.

Read SWE-bench numbers as vendor claims

The llm-stats tracker lists 116 SWE-bench Verified results, all self-reported and 0 independently verified. Treat the rankings above as vendor claims and run your own prompts before committing volume.

LLM API Free Tiers: Exact Amounts

ProviderFree offerLimit
Z.AIGLM-4.7-Flash, GLM-4.5-FlashFree models, no token charge
Alibaba Model Studio1M tokens for new users90-day validity
OpenAI APIFree tier in allowed geographies$100/month usage cap
Google Gemini APIFree tier on 3.x Flash and Flash-LiteNot on Gemini 3.1 Pro Preview
AnthropicSmall free credit for new usersThen Start tier, $500/month cap
Morph WarpGrepNone$0.80 per 100K tokens

For experimentation, the practical order is: GLM Flash models (free, no token charge), Gemini Flash free tier (current 3.8 Flash included), then Qwen's 1M tokens.

Cost Calculator: Real Workloads

Per-token prices mean nothing until mapped to usage. Two reference workloads, 30-day months, standard tier, no cache discounts (caching reduces all of these), all prompts under the long-context thresholds.

Coding agent: 50M input / 5M output tokens per day

ModelDaily costMonthly cost
Claude Fable 5.1$750.00$22,500
GPT-6 Astra$750.00$22,500
Claude Opus 5$375.00$11,250
Claude Opus 5.5$300.00$9,000
Kimi K3$225.00$6,750
Gemini 3.1 Pro$160.00$4,800
GPT-6 Sol$150.00$4,500
Claude Sonnet 5$150.00$4,500
GLM-5.3$92.00$2,760
Gemini 3.8 Flash (promo)$56.25$1,688
DeepSeek V4.1 Flash (peak)$21.00$630
MiniMax M3$21.00$630
GLM-5.3-Flash$10.00$300
Morph morph-dsv4flash$9.10$273
GPT-6 Luna$7.50$225
Morph morph-glm53flash$13.50$405

Support chatbot: 20M input / 5M output tokens per day

ModelDaily costMonthly cost
Claude Haiku 4.5$45.00$1,350
Gemini 3.8 Flash (promo)$33.75$1,013
Gemini 3.5 Flash-Lite$18.50$555
DeepSeek V4.1 Flash (peak)$12.00$360
MiniMax M3$12.00$360
GLM-5.3-Flash$5.50$165
Morph morph-dsv4flash$4.84$145
GPT-6 Luna$4.50$135
Morph morph-glm53flash$7.50$225
The 56x gap on identical traffic

The same coding-agent workload costs $22,500/month on Claude Fable 5.1 or GPT-6 Astra and $405/month on Morph's morph-glm53flash, a 56x gap. The production answer is rarely either extreme: route routine edits to a cheap model, escalate multi-file refactors to a frontier one, and cache aggressively (DeepSeek V4.1 Flash cache hits bill input at $0.006/M at peak; Anthropic cache reads are 0.1x or less and do not count against rate limits). Model the routing split with the LLM cost calculator.

Latency and Throughput

Two metrics matter: time to first token (how fast streaming starts) and tokens per second (how fast it finishes). Provider speed claims vary with load and region, so measure on your own traffic. What the official docs establish:

  • Speed is now a paid SKU at both frontier labs. OpenAI Fast mode costs 2x standard on GPT-6 and GPT-5.6. Anthropic fast mode (research preview, Claude API only, not with batch) costs $8/$40 on Opus 5.5 and $10/$50 on Opus 5 and Opus 4.8.
  • Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5 at standard speed.
  • Reasoning adds latency by design. Fable 5.1 and Opus 5.5 run adaptive thinking that is always on; GLM-5.3 reasoning is mandatory and defaults to max effort. Thinking tokens are generated and billed before visible output.
  • MiniMax sells a priority tier at 1.5x standard for faster scheduling on M3.
  • Specialized models break the general-purpose ceiling: Morph Fast Apply sustains 10,500 tok/s on code-edit application because the task (merging an edit into a file) does not need frontier reasoning.

Specialized APIs: When General-Purpose Falls Short

Coding agents spend much of their compute on two mechanical operations: searching codebases for context and applying edits to files. Both run through general-purpose LLMs by default, at general-purpose prices and speeds.

Morph Fast Apply

Code-edit application at 10,500+ tok/s with 98% accuracy and sub-second cold starts. The agent outputs a lazy edit snippet; Fast Apply merges it into the full file. OpenAI-compatible /v1/chat/completions endpoint.

Morph WarpGrep

Agentic codebase search, #1 on SWE-Bench Pro. $0.80 per 100K tokens. Ships as an MCP server for any agent.

The pattern generalizes: a frontier model reasons and decides what to change; narrow, fast models execute the mechanical steps. The frontier model's output shrinks (edit snippets instead of whole files), and output is the token class that costs $10 to $50/M. See Fast Apply and WarpGrep.

Best LLM API by Use Case

One ranking does not fit every job. Pick by the constraint that binds you. Each pick below comes from the data already on this page.

Best for reasoning and hard coding

Claude Fable 5.1 ($10/$50; Fable 5 scores 95.0% on SWE-bench Verified) and GPT-6 Astra ($10/$50) lead. Claude Opus 5.5 ($4/$20) is Anthropic's recommended default and, per Anthropic, matches Astra on Terminal-Bench 4.0 at about 40% of the cost. Use these for multi-file refactors and tasks where a cheaper model measurably fails.

Best value

GPT-6 Sol and Claude Sonnet 5 are both $2/$10. Sonnet 5 scores 85.2% on SWE-bench Verified at $0.12 of output per benchmark point versus $0.28 for Opus 4.8 and $0.53 for Fable 5. MiniMax M3 ($0.30/$1.20) scores 80.5%, the cheapest model above 80%, at $0.015 per point. DeepSeek V4 Pro Max (80.6%, open weights) matches it for self-hosting.

Best for speed

For code edits, Morph Fast Apply: 10,500 tok/s with 98% accuracy, because merging an edit into a file does not need frontier reasoning. For frontier chat speed, both OpenAI (Fast mode, 2x price) and Anthropic (fast mode on Opus 5.5, $8/$40) sell paid fast tiers.

Morph Fast Apply: best for code edits

10,500 tok/s, 98% accuracy. The frontier model decides what to change; Fast Apply merges the edit snippet into the full file behind an OpenAI-compatible endpoint. See /products/fastapply.

Morph model router

Route each request to the cheapest model that can handle it, billed at $0.005 per request. One key, OpenAI-compatible. See /llm-router.

Best on a budget

Morph morph-glm53flash ($0.20/$0.70, 1M context, GLM-5.3-Flash at 16-bit) is the price floor among current-generation models, followed by GPT-6 Luna ($0.10/$0.50), morph-dsv4flash ($0.142/$0.40), Qwen3.8 Flash ($0.15/$0.47), and GLM-5.3-Flash on Z.AI ($0.15/$0.50). DeepSeek's V4.1 Flash is $0.30/$1.20 at peak and $0.15/$0.60 off-peak. GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.AI API. For high-volume simple traffic, start here and escalate only the requests that measurably need a frontier model.

How to Choose an LLM API

Decision framework
  • 1. Establish your quality floor cheaply. Run your real prompts through morph-glm53flash ($0.70/M out), GPT-6 Luna ($0.50/M), DeepSeek V4.1 Flash ($1.20/M peak), and MiniMax M3 ($1.20/M). If one passes, you are done at a fraction of frontier cost.
  • 2. Escalate only measured gaps. Move to GPT-6 Sol or Claude Sonnet 5 ($10/M), Gemini 3.1 Pro ($12/M), Opus 5.5 ($20/M), or Fable 5.1 and GPT-6 Astra ($50/M) for the tasks where the cheap tier measurably fails.
  • 3. Check the constraint that binds you. Rate-limited on day one? Anthropic's Start tier gives Opus 5.5 2M ITPM but Fable 5.x only 500K. Prompts over 272K? GPT-6's long-context tier applies. Need self-hosting or data control? DeepSeek, GLM-5.3, and MiniMax M3 are open weights. Compliance? Bedrock, Google Cloud, or Foundry.
  • 4. Use specialized APIs for mechanical steps. Edit application, search, embeddings, and reranking all have purpose-built models that beat $10-50/M generalists on both speed and cost.
  • 5. Pin model IDs and re-run evals on release days. DeepSeek's deepseek-v4-flash ID now serves a different model. Aliases move.

The most common mistake is anchoring on one provider's flagship and never testing down. MiniMax M3, Kimi K2.6, Qwen3.7 Max, and Gemini 3.1 Pro all score 80.2% to 80.6% on SWE-bench Verified for $1.20 to $12 per million output; the frontier costs $20 to $50. The second most common is ignoring caching: Anthropic bills cached reads at 0.1x (0.05x on Opus 5.5, 0.025x on Fable 5.1) and exempts them from rate limits, and DeepSeek V4.1 Flash drops cached input to $0.006/M. For long-context workloads, compare windows in detail at LLM context window comparison.

Private deployments

The fastest endpoints are private deployments

Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 100 billion tokens per day run this way.

Talk to us about a private deployment

Frequently Asked Questions

What is an LLM API?

An LLM API is an HTTP interface to a hosted large language model. You send a prompt (plus optional tools, images, or files) to an endpoint such as /v1/chat/completions and receive generated tokens back, billed per million tokens of input and output. It replaces running model weights on your own GPUs with a metered, pay-per-token service.

What are the best LLM API providers in 2026?

The major first-party providers are OpenAI (GPT-6 Astra $10/$50, GPT-6 Sol $2/$10, GPT-6 Luna $0.10/$0.50 per 1M tokens), Anthropic (Claude Fable 5.1 $10/$50, Opus 5.5 $4/$20, Sonnet 5 $2/$10), Google (Gemini 3.1 Pro $2/$12, Gemini 3.8 Flash $0.75/$3.75 through December 31, 2026), DeepSeek (V4.1 Flash $0.30/$1.20 at peak), Moonshot (Kimi K3 $3/$15), Z.AI (GLM-5.3 $1.40/$4.40), MiniMax (M3 $0.30/$1.20), and Alibaba (Qwen3.8 Max $2/$6). OpenRouter aggregates them behind one key; Bedrock, Google Cloud, and Microsoft Foundry resell Claude with cloud billing. Morph serves open-weight models such as GLM-5.3 ($1.19/$3.74) and Kimi K3 ($2.50/$14.00) at 16-bit precision behind one OpenAI-compatible key.

What is the cheapest LLM API in 2026?

Among current-generation models, Morph's morph-glm53flash (GLM-5.3-Flash at 16-bit) at $0.20/M input and $0.70/M output, then OpenAI's GPT-6 Luna at $0.10/$0.50 (released September 22, 2026), Morph's morph-dsv4flash at $0.142/$0.40, Alibaba's Qwen3.8 Flash at $0.15/$0.47, and GLM-5.3-Flash on Z.AI at $0.15/$0.50. DeepSeek V4.1 Flash is $0.30/$1.20 at peak and $0.15/$0.60 off-peak. MiniMax M3 at $0.30/$1.20 is the cheapest model with a published SWE-bench Verified score above 80%. GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.AI API.

Which LLM API is best for coding?

Claude Fable 5 has the top SWE-bench Verified score on the llm-stats tracker at 95.0% ($10/$50); its successor Fable 5.1 (September 1, 2026, same price) has no tracker score yet. Claude Opus 5.5 (September 22, 2026, $4/$20) scores 66.4% on Terminal-Bench 4.0, which Anthropic says matches GPT-6 Astra at about 40% of the cost. Claude Sonnet 5 scores 85.2% at $2/$10. On a budget, MiniMax M3 (80.5%, $0.30/$1.20) and open-weights DeepSeek V4 Pro Max (80.6%) are the value picks. For applying code edits, Morph Fast Apply runs at 10,500 tok/s with 98% accuracy.

What rate limits do LLM APIs have?

Anthropic now uses Start, Build, and Scale tiers. On the Start tier, Opus 5.5, Opus 5, Sonnet 5, and Haiku 4.5 each get 1,000 RPM and 2M input tokens per minute, while Fable 5.x gets 1,000 RPM and 500K ITPM. Scale gives 10,000 RPM and 10M ITPM. Cached input tokens do not count toward Anthropic's limits. OpenAI tiers unlock by spend ($5 to $1,000 paid); GPT-6 Astra runs 500 RPM / 500K TPM at Tier 1 and 15,000 RPM / 40M TPM at Tier 5. DeepSeek caps concurrency instead: 2,500 parallel requests on deepseek-flash, 500 on deepseek-v4-pro.

How do I handle LLM API rate limits?

Read the 429 response before retrying. Rate-limit 429s carry a retry-after header; honor it with exponential backoff. Anthropic's monthly spend-cap 429 has the same error type but no retry-after header and error.details.error_code set to enforced_spend_limit_reached, so SDK auto-retries keep failing until the first of the next month. Cache repeated context (Anthropic does not count cache reads toward ITPM), ramp traffic gradually to avoid acceleration limits, and keep a second provider behind an OpenAI-compatible router for overflow.

Which LLM APIs have a free tier?

Z.AI serves GLM-4.7-Flash and GLM-4.5-Flash free. Alibaba Model Studio gives new users 1M free tokens valid for 90 days. OpenAI has a Free API tier in allowed geographies with a $100/month usage cap. Google offers a free tier on Gemini 3.8, 3.7, 3.6, and 3.5 Flash and Flash-Lite, but not on Gemini 3.1 Pro Preview. Anthropic gives new users a small amount of free credits to test the API.

Which LLM API has the largest context window?

GPT-6 Astra, Sol, and Luna list a 1,050,000-token context window (922,000 max input, 128,000 max output). Claude Fable 5.1, Opus 5.5, and Sonnet 5 have 1M tokens at standard pricing with no long-context surcharge. Kimi K3 is 1,048,576 tokens. DeepSeek V4.1 Flash, GLM-5.3, and Qwen3.8 Max support 1M. Morph serves GLM-5.3 and DeepSeek V4 Flash with 1M and 1M contexts. Below 1M: Kimi K2.6 and K2.7 Code at 262,144 and Claude Haiku 4.5 at 200K.

Are LLM APIs OpenAI-compatible?

Mostly yes. DeepSeek, Moonshot (Kimi), Z.AI (GLM), MiniMax, Alibaba (Qwen), OpenRouter, and Morph accept OpenAI-format /v1/chat/completions requests, so switching providers is usually a base_url and api_key change. Google exposes an OpenAI-compatible endpoint alongside its native API. Anthropic's native API is Messages-format with an OpenAI SDK compatibility layer. DeepSeek also accepts Anthropic-format requests at api.deepseek.com/anthropic.

Why are my LLM API costs so high?

Four usual causes. Output tokens cost 3x to 6x input, and reasoning models bill thinking tokens as output. Long prompts cross price thresholds: GPT-6 bills 2x input and 1.5x output above 272K input tokens, and Gemini 3.1 Pro rises from $2/$12 to $4/$18 above 200K. Cache misses: a cache hit is 10x cheaper on OpenAI and Anthropic and 50x cheaper on DeepSeek V4.1 Flash. And newer Claude models (Opus 4.7 onward) use a tokenizer that produces about 30% more tokens for the same text.

Should I use one LLM API provider or multiple?

Production systems usually route between two or more. Send high-volume simple traffic to a sub-$1.50/MTok model (GPT-6 Luna, DeepSeek V4.1 Flash, MiniMax M3) and escalate hard tasks to a frontier model (Fable 5.1, Opus 5.5, GPT-6 Astra). Because most providers are OpenAI-compatible, a router or gateway makes the switch a config change rather than a rewrite.

Which LLM API is best by use case?

Best for reasoning and hard coding: Claude Fable 5.1 ($10/$50), Claude Opus 5.5 ($4/$20), and GPT-6 Astra ($10/$50). Best value: Claude Sonnet 5 (85.2% SWE-bench Verified, $2/$10), GPT-6 Sol ($2/$10), and MiniMax M3 ($0.30/$1.20). Best for speed: Morph Fast Apply for code edits at 10,500 tok/s with 98% accuracy. Best on a budget: Morph morph-glm53flash ($0.20/$0.70), GPT-6 Luna ($0.10/$0.50), or the free GLM-4.7-Flash and GLM-4.5-Flash on Z.AI.

What changed in LLM API pricing in September 2026?

OpenAI released GPT-6 Astra ($10/$50) on September 3 and GPT-6 Sol ($2/$10) and Luna ($0.10/$0.50) on September 22, half the price of GPT-5.6 Sol and Luna. Anthropic released Claude Fable 5.1 ($10/$50) on September 1 and Claude Opus 5.5 ($4/$20) on September 22, now its recommended default. DeepSeek retired V4 Flash for V4.1 Flash on September 10 at $0.30/$1.20 peak. Google shipped Gemini 3.8 Flash on September 2 at $0.75/$3.75 through December 31, 2026. xAI released Grok 4.7 on September 21 at $2/$6 with a 500K context.

Sources

Every third-party price, limit, and benchmark on this page traces to a primary source, checked September 22, 2026. Morph prices come from the live Morph model list.

Related Resources

Code Editing at 10,500 tok/s

Frontier LLM APIs bill $10-50 per million output tokens to rewrite whole files. Morph Fast Apply merges edit snippets into files at 10,500 tok/s with 98% accuracy, behind an OpenAI-compatible API.