An LLM API is an HTTP endpoint for a hosted large language model: you POST a prompt to a URL such as /v1/chat/completions and pay per million input and output tokens. As of September 22, 2026, output prices on general models run from $0.50/M (GPT-6 Luna) to $50/M (Claude Fable 5.1, GPT-6 Astra). Most providers accept the OpenAI request format, so switching is a base_url change.
Cheapest current-generation LLM API: Morph's morph-glm53flash at $0.20/$0.70 per 1M tokens, then GPT-6 Luna at $0.10/$0.50, morph-dsv4flash at $0.142/$0.40, Qwen3.8 Flash at $0.15/$0.47, and GLM-5.3-Flash at $0.15/$0.50. DeepSeek's own V4.1 Flash is $0.30/$1.20 at peak and $0.15/$0.60 off-peak. Best value at the frontier: GPT-6 Sol and Claude Sonnet 5, both $2/$10. Top models: Claude Fable 5.1 ($10/$50), GPT-6 Astra ($10/$50), and Claude Opus 5.5 ($4/$20), which Anthropic released today as its recommended default. Most production systems route each request to the cheapest model that can handle it. See the Morph model router and the model list at Morph models.
What Is an LLM API
An LLM API is an HTTP interface to a hosted large language model. You send a prompt (plus optional tools, images, or files) to an endpoint such as /v1/chat/completions and receive generated tokens back, billed per million tokens (MTok) of input and output. It replaces running model weights on your own GPUs with a metered service: you run no infrastructure, manage no model updates, and pay only for tokens used.
Every provider on this page works that way. The differences are price, rate limits, context window, and model quality, and they are large. Output tokens cost between $0.50/M (GPT-6 Luna) and $50/M (Fable 5.1, GPT-6 Astra), a 100x spread. Seven models on the llm-stats tracker sit between 80.2% and 80.6% on SWE-bench Verified, and the four of them with list API prices span $1.20 to $12 per million output tokens. The cheapest rows are mostly open-weight models; the open source LLM guide covers their licenses and the point at which self-hosting beats any API rate.
What Changed in LLM API Pricing in September 2026
Every price below was re-checked on the provider's official pricing page on September 22, 2026. Since the September 2 revision of this page:
- OpenAI GPT-6. GPT-6 Astra shipped September 3 at $10/$50. GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50) shipped September 22, half the price of GPT-5.6 Sol ($4/$20) and Luna ($0.20/$1.20). OpenAI's pricing page says GPT-5.6 Sol's promotional $4/$20 price runs at least through November 21, 2026. GPT-5.6 and GPT-6 also add a cache-write price column at 1.25x input.
- Claude Opus 5.5. Released September 22 at $4/$20, down from Opus 5's $5/$25, with cache reads at 0.05x ($0.20/M). Anthropic now recommends it as the starting model and lists Opus 5 as legacy. Anthropic says it costs 40% less to run than Opus 5 on typical workloads and generates output more than 30% faster.
- DeepSeek V4.1 Flash. Released September 10 under the model ID
deepseek-flashat $0.30/$1.20 peak and $0.15/$0.60 off-peak. V4 Flash is retired. - Gemini 3.7 and 3.8 Flash. Google shipped Gemini 3.8 Flash on September 2. Gemini 3.6, 3.7, and 3.8 Flash are all $0.75/$3.75 through December 31, 2026, then $1.50/$7.50.
- Anthropic rate limits. Anthropic replaced its four numbered tiers with Start, Build, and Scale on June 26, 2026. The entry tier allows 1,000 RPM on every current model, and Opus 5.5 gets its own bucket.
- Z.AI. The GLM-5.3-Flash 50% promotion ended; it is $0.15/$0.50. GLM-5.3-FlashX was added at $0.37/$1.25. GLM-5.3 requests that send
thinking.type: "disabled"fail; sendenabledwithreasoning_effort: lowinstead. - xAI Grok 4.7. Released September 21 at $2/$6 per 1M tokens ($4/$12 on long context) with a 500K context. xAI's docs recommend it for code. The 2x-priced Grok 4.7 Fast is only in Cursor and Grok Build, not the public API.
- Retired model IDs. Moonshot retired
kimi-k2.5and everymoonshot-v1model in August 2026; they now return 404. Anthropic retired Claude Opus 4.1 on August 5. DeepSeek'sdeepseek-chatanddeepseek-reasonernames were scheduled for discontinuation on July 24 and no longer appear on its pricing page.
LLM API Pricing Table: Flagship and Cheapest Model per Provider
Prices are $ per 1M tokens, input / output, standard (non-batch) tier, sorted by output price. Context is the maximum window where the provider publishes it.
| Provider | Model | Input/MTok | Output/MTok | Context |
|---|---|---|---|---|
| Anthropic | Claude Fable 5.1 / Fable 5 | $10.00 | $50.00 | 1M |
| OpenAI | GPT-6 Astra | $10.00 | $50.00 | 1.05M |
| OpenAI | GPT-5.5 | $5.00 | $30.00 | long-context tier |
| Anthropic | Claude Opus 5 / Opus 4.8 (legacy) | $5.00 | $25.00 | 1M |
| Anthropic | Claude Opus 5.5 | $4.00 | $20.00 | 1M |
| OpenAI | GPT-5.6 Sol | $4.00 | $20.00 | long-context tier |
| Moonshot | Kimi K3 | $3.00 | $15.00 | 1,048,576 |
| Anthropic | Claude Sonnet 4.6 (legacy) | $3.00 | $15.00 | 1M |
| Morph | morph-kimik3 (Kimi K3, 16-bit) | $2.50 | $14.00 | 1M |
| Gemini 3.1 Pro (preview, ≤200K prompts) | $2.00 | $12.00 | n/a | |
| OpenAI | GPT-5.6 Terra | $2.00 | $12.00 | long-context tier |
| OpenAI | GPT-6 Sol | $2.00 | $10.00 | 1.05M |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | 1M |
| Gemini 3.5 Flash | $1.50 | $9.00 | n/a | |
| Alibaba | Qwen3.7 Max | $2.50 | $7.50 | 1M |
| Alibaba | Qwen3.8 Max | $2.00 | $6.00 | 1M |
| xAI | Grok 4.7 (reference, released Sep 21) | $2.00 | $6.00 | 500K |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | 200K |
| Z.AI | GLM-5.3 / GLM-5.2 | $1.40 | $4.40 | 1M |
| Moonshot | Kimi K2.6 | $0.95 | $4.00 | 262,144 |
| DeepSeek | V4 Pro (peak / off-peak) | $1.32 / $0.66 | $3.96 / $1.98 | 1M |
| Gemini 3.8 / 3.7 / 3.6 Flash (through Dec 31) | $0.75 | $3.75 | n/a | |
| Morph | morph-glm53-744b (GLM-5.3, 16-bit) | $1.19 | $3.74 | 1M |
| DeepSeek | V4.1 Flash (peak / off-peak) | $0.30 / $0.15 | $1.20 / $0.60 | 1M |
| MiniMax | M3 (up to 512K input) | $0.30 | $1.20 | priced to 1M |
| Morph | morph-dsv41flash (DeepSeek V4.1 Flash) | $0.15 | $0.60 | 1M |
| OpenAI | GPT-6 Luna | $0.10 | $0.50 | 1.05M |
| Z.AI | GLM-5.3-Flash | $0.15 | $0.50 | 1M |
| Alibaba | Qwen3.8 Flash | $0.15 | $0.47 | 1M |
| Morph | morph-dsv4flash (DeepSeek V4 Flash 0731, 16-bit) | $0.142 | $0.40 | 1M |
| Morph | morph-glm53flash (GLM-5.3-Flash, 16-bit) | $0.20 | $0.70 | 1M |
| Z.AI | GLM-4.7-Flash / GLM-4.5-Flash | Free | Free | n/a |
xAI is not one of the 12 providers profiled below; its Grok 4.7 row is there for price reference. Notes that change effective cost: GPT-6 bills 2x input and 1.5x output on prompts over 272K input tokens, so Astra becomes $20/$75, Sol $4/$15, and Luna $0.20/$0.75. GPT-5.6 has a similar long-context tier ($8/$30 Sol, $4/$18 Terra, $0.40/$1.80 Luna), and GPT-5.5 goes to $10/$45. Gemini 3.1 Pro rises to $4/$18 above 200K tokens. MiniMax M3 is $0.30/$1.20 up to 512K input and $0.60/$2.40 above, after a permanent 50% discount off list. DeepSeek bills peak rates 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays (excluding Chinese public holidays) and half price at all other hours. Anthropic charges the same per-token rate at any context length on Claude 4.6 and later (a 900K-token request bills like a 9K one), but Opus 4.7 and later use a tokenizer that produces about 30% more tokens for the same text.
Batch and cache discounts: OpenAI and Anthropic batch is 50% off. OpenAI cached input is 10% of base ($0.20/M on GPT-6 Sol, $1/M on Astra), and cache writes on GPT-5.6 and GPT-6 bill at 1.25x input. Anthropic cache reads are 0.1x base input, 0.05x on Opus 5.5 ($0.20/M), and 0.025x on Fable 5.1 ($0.25/M). GLM-5.3 cached input is $0.26/M. Kimi K3 cache hits are $0.30/M. DeepSeek V4.1 Flash cache hits are $0.006/M at peak, 50x below its $0.30 cache-miss rate. If your workload re-sends the same system prompt or file context, the cache column matters more than the headline price. On paper, a few older small models undercut everything above: Z.AI's GLM-4.7-FlashX is $0.07/$0.40 and GLM-4-32B-0414-128K is $0.10/$0.10.
Provider-by-Provider Breakdown
OpenAI
GPT-6 is the new top of OpenAI's price list in three sizes: Astra ($10/$50, September 3), Sol ($2/$10), and Luna ($0.10/$0.50), the latter two released September 22. All three list a 1,050,000-token context window with 922,000 max input and 128,000 max output. GPT-5.6 Sol ($4/$20), Terra ($2/$12), and Luna ($0.20/$1.20) and GPT-5.5 ($5/$30) stay on the list. There is no GPT-6 Terra. Fast mode (renamed from priority processing on July 30, 2026) costs 2x standard on GPT-6 and GPT-5.6 and 2.5x on GPT-5.5. Regional data residency adds 10% on models released on or after March 5, 2026, and GPT-6 EU data residency is Standard tier only. See Codex pricing for the subscription side.
| Model | Input/MTok | Cached In | Output/MTok | Over 272K input (In/Out) |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | $20.00 / $75.00 |
| GPT-5.5 | $5.00 | $0.50 | $30.00 | $10.00 / $45.00 |
| GPT-5.6 Sol | $4.00 | $0.40 | $20.00 | $8.00 / $30.00 |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | $4.00 / $18.00 |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | $4.00 / $15.00 |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | $0.40 / $1.80 |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | $0.20 / $0.75 |
A gotcha on long prompts: past 272K input tokens GPT-6 Sol costs $4/$15, while Claude Opus 5.5 is a flat $4/$20 at any length up to 1M. A Hacker News commenter on the Sol launch thread made the same point: once you pass 272K, Sol is roughly Opus-priced and Astra roughly Fable-priced. If your agent routinely sends 300K+ token prompts, model the long-context tier, not the headline.
Anthropic (Claude)
Anthropic's current lineup is Claude Fable 5.1 ($10/$50, released September 1, 2026), Claude Opus 5.5 ($4/$20, released September 22, 2026), Claude Sonnet 5 ($2/$10), and Claude Haiku 4.5 ($1/$5). Anthropic's models overview now says to start with Opus 5.5 for most workloads and use Fable 5.1 for demanding reasoning and long-horizon agentic work. Opus 5.5 reports 66.4% on Terminal-Bench 4.0, which Anthropic says matches GPT-6 Astra at about 40% of the cost, and 57.8% on CursorBench 4.0. Sonnet 5's $2/$10 is the standard price; the planned September 1 increase to $3/$15 was cancelled. Opus 5.5 defaults to medium effort (Opus 5 defaulted to high), so an integration that omits effort gets a different cost profile. Anthropic says Sonnet 5.5 and Haiku 5.5 follow in the coming weeks. Fable 5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Opus 4.5, Sonnet 4.6, and Sonnet 4.5 remain available as legacy models. Opus 4.1 was retired on August 5, 2026. Sonnet 4.5 can retire as early as September 29, 2026, and Haiku 4.5 as early as October 15, 2026. Mythos 5.1 ($10/$50) is limited availability. Fable 5.1, Opus 5.5, and Sonnet 5 carry a 1M context and 128K max output; Haiku 4.5 is 200K and 64K. Full breakdown: Anthropic API pricing.
| Model | Input/MTok | Cache Read | Output/MTok | Context |
|---|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 | 1M |
| Claude Fable 5 (legacy) | $10.00 | $1.00 | $50.00 | 1M |
| Claude Opus 5 / Opus 4.8 (legacy) | $5.00 | $0.50 | $25.00 | 1M |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 | 1M |
| Claude Sonnet 4.6 (legacy) | $3.00 | $0.30 | $15.00 | 1M |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | 1M |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | 200K |
Google (Gemini)
Gemini 3.1 Pro is still labelled preview at $2/$12 for prompts up to 200K tokens ($4/$18 above), has no free tier, and scores 80.6% on SWE-bench Verified. The Flash line moved fast: Gemini 3.8 Flash (September 2), 3.7 Flash, and 3.6 Flash are all stable at a promotional $0.75/$3.75 through December 31, 2026, rising to $1.50/$7.50 after. Google positions 3.8 Flash for long-horizon software engineering and agents. Gemini 3.5 Flash ($1.50/$9) remains listed, and Gemini 3.5 Flash-Lite ($0.30/$2.50) and 3.1 Flash-Lite ($0.25/$1.50) are the budget tiers. Every Flash and Flash-Lite model has a free tier.
DeepSeek
DeepSeek V4.1 Flash (model ID deepseek-flash) launched September 10, 2026: an MoE with 552B backbone parameters (about 763B in total counting its 196B Engram memory and vision encoder), 8B active parameters for input and 16B for output, native image input, a 1M context, and 384K max output. It bills $0.30/$1.20 at peak and $0.15/$0.60 off-peak, with cache hits at $0.006 and $0.003. DeepSeek says its KV cache needs a quarter of the HBM of the previous generation and that it beats V4 Pro on DeepSeek's benchmarks. V4 Pro (deepseek-v4-pro) is listed at $1.32/$3.96 peak and $0.66/$1.98 off-peak. The API accepts OpenAI format at api.deepseek.com and Anthropic format at api.deepseek.com/anthropic. Deep dive: DeepSeek V4.
DeepSeek retired V4 Flash on September 10. Requests to deepseek-v4-flash still succeed, but they are served by V4.1 Flash and billed at the Flash price, so an eval pinned to the old ID silently changed models. The release note first said deepseek-v4-pro would also route to V4.1 Flash from September 14. DeepSeek reversed that on its updates page: V4 Pro stays on the API after September 14 with billing unchanged. The legacy deepseek-v4-flash-vision-exp ID also routes to V4.1 Flash. One HN reader who went through the model card described V4.1 Flash as tuned for agentic tool calling, cheap prefill on 8B active parameters, and moderate output, at some cost to knowledge. If you need the V4 Flash behavior you validated against, Morph still serves the DeepSeek-V4-Flash-0731 checkpoint as morph-dsv4flash.
Where you run an open-weight model changes its output and its price. Morph serves open-weight models (GLM-5.3, GLM-5.3-Flash, Kimi K3, DeepSeek V4 Flash 0731, DeepSeek V4.1 Flash) with 16-bit (bf16) activations and does not quantize them, so output matches the released weights. morph-dsv4flash is $0.142/M input and $0.40/M output with no peak surcharge, morph-dsv41flash is $0.15/$0.60, and morph-glm53-744b is $1.19/$3.74 with a 1M context. For coding agents, Morph adds codegen-tuned speculative decoding and custom inference kernels. See Morph models and pricing.
Moonshot (Kimi)
Kimi K3 ($3.00/$15.00, cache hits $0.30, cache writes $3.00 at the default 5-minute TTL or $6.00 at 1 hour) is Moonshot's flagship with a 1,048,576-token context. The API now lives at platform.kimi.ai. In August 2026 Moonshot retired kimi-k2.5 and all moonshot-v1 models, which now return 404, so pin a current ID. In September it added a Web Search API and began deducting usage 50% from cash and 50% from vouchers (contract customers excepted). Kimi K2.6 ($0.95/$4.00, cache hits $0.16, 262,144 context, 80.2% on SWE-bench Verified) and the coding-tuned Kimi K2.7 Code ($0.95/$4.00, or $1.90/$8.00 for the high-speed variant) remain available. Morph serves Kimi K3 as morph-kimik3 at $2.50/$14.00 with 16-bit activations. Deep dive: Kimi K3.
Z.AI (GLM)
GLM-5.3 ($1.40/$4.40, cached input $0.26/M, 1M context, 128K max output) is Z.AI's open-weights flagship, priced the same as GLM-5.2 and GLM-5.1. It takes text input only. Reasoning is mandatory at low, high, or max effort, with max the default: thinking.type accepts only enabled, and a request that sends disabled fails. Z.AI reports 34.5% task completion on its Code Bench at max effort using about 75K output tokens, against 23.4% at 96K tokens for GLM-5.2. GLM-5.3-Flash is $0.15/$0.50 now that its launch promotion has ended, and GLM-5.3-FlashX is $0.37/$1.25. GLM-4.7-Flash and GLM-4.5-Flash are free. Morph serves morph-glm53-744b at $1.19/$3.74 and morph-glm53flash at $0.20/$0.70 with 16-bit activations. Deep dive: GLM-5 family.
MiniMax
MiniMax M3 is $0.30/$1.20 up to 512K input tokens, a permanent 50% discount off the $0.60/$2.40 list price, and $0.60/$2.40 above 512K; cache reads are $0.06. It scores 80.5% on SWE-bench Verified, the cheapest 80%+ model on a hosted API: about one-twentieth of Opus 4.8's output price for a score 8 points lower. A priority tier runs 1.5x standard ($0.45/$1.80). MiniMax M2.7 ($0.30/$1.20) is the prior generation.
Alibaba (Qwen)
Qwen3.8 Max ($2.00/$6.00 in the Singapore region, up to 1M context) is Alibaba's proprietary flagship and undercuts Qwen3.7 Max, which lists at $2.50/$7.50 and scores 80.4% on SWE-bench Verified. Qwen3.8 Flash ($0.15/$0.47, 1M context) is the budget tier, built on the open Qwen3.8-Flash-Next architecture. Qwen3.6 Plus ($0.50/$3.00, tiered at 256K) remains available, and Qwen3.7 Plus is 20% off list for a limited time in Singapore. New Model Studio users get 1M free tokens valid 90 days. The Qwen API guide has model IDs, base URLs by region, and the open-weight self-hosting path.
Aggregators and Cloud Resellers
OpenRouter fronts many providers' models behind one OpenAI-compatible key, useful for evaluation and failover. Claude models are also sold through Amazon Bedrock and Google Cloud (billed by the cloud provider, with a 10% premium on regional endpoints) and through Claude Platform on AWS and Claude in Microsoft Foundry, which bill Anthropic's standard rates through the cloud marketplace. OpenAI models have been on Amazon Bedrock since June 1, 2026, at the same price as direct. If you need a routing layer rather than a reseller, see LLM gateways and LLM routers.
Morph (specialized)
Beyond the open-weight chat models above, Morph serves task-specific models: Fast Apply merges code edits at 10,500 tok/s, and WarpGrep does agentic codebase search at $0.80 per 100K tokens. Covered in Specialized APIs below.
LLM API Rate Limits by Provider and Tier
Rate limits decide whether your launch survives traffic, and most comparison pages skip them. Here is what each provider enforces, from official docs as of September 22, 2026.
Anthropic: Start, Build, and Scale tiers, cache-aware token counting
Anthropic replaced its numbered tiers with Start ($500 monthly spend cap), Build ($1,000), and Scale ($200,000), plus a Custom tier with no cap. Organizations move up automatically with usage history; new organizations may start on an Evaluation tier below these limits. Limits are per model, in requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Cache reads do not count toward ITPM: with a 2M ITPM limit and an 80% cache hit rate you can process 10M total input tokens per minute.
| Model | Start (RPM / ITPM / OTPM) | Build | Scale |
|---|---|---|---|
| Claude Fable 5.x (combined) | 1,000 / 500K / 100K | 2,000 / 1.5M / 300K | 4,000 / 4M / 800K |
| Claude Opus 5.5 | 1,000 / 2M / 400K | 5,000 / 5M / 1M | 10,000 / 10M / 2M |
| Claude Opus 5 | 1,000 / 2M / 400K | 5,000 / 5M / 1M | 10,000 / 10M / 2M |
| Claude Opus 4.x (combined) | 1,000 / 2M / 400K | 5,000 / 5M / 1M | 10,000 / 10M / 2M |
| Claude Sonnet 5 | 1,000 / 2M / 400K | 5,000 / 5M / 1M | 10,000 / 10M / 2M |
| Claude Haiku 4.5 | 1,000 / 2M / 400K | 5,000 / 5M / 1M | 10,000 / 10M / 2M |
Two things to plan around. Fable 5.x gets a quarter of the entry-tier input throughput of every other current model (500K vs 2M ITPM), shared across Fable 5.1 and Fable 5. Opus 5.5 and Opus 5 each have their own bucket, separate from the combined Opus 4.x bucket, so spreading traffic across Opus versions raises your effective ceiling. Hitting the monthly spend cap returns HTTP 429 with the same rate_limit_error type as a rate limit but no retry-after header and error.details.error_code set to enforced_spend_limit_reached. SDK auto-retries keep failing until 00:00 UTC on the first of the next month, so check that field before you back off.
OpenAI: spend-unlocked tiers, per-model limits
| Tier | Qualification | Monthly usage cap | GPT-6 Astra RPM / TPM |
|---|---|---|---|
| Free | Allowed geography | $100 | not available |
| Tier 1 | $5 paid | $100 | 500 / 500K |
| Tier 2 | $50 paid | $500 | 5,000 / 1M |
| Tier 3 | $100 paid | $1,000 | 5,000 / 2M |
| Tier 4 | $250 paid | $5,000 | 10,000 / 4M |
| Tier 5 | $1,000 paid | $200,000 | 15,000 / 40M |
Per-model RPM/TPM numbers live on each model's page. GPT-6 Sol and Luna also start at 500 RPM and 500K TPM on Tier 1, with batch queue limits of 1.5M and 5M tokens.
DeepSeek: concurrency caps instead of token budgets
DeepSeek publishes no RPM/TPM limits. It caps concurrent requests: 2,500 in flight on deepseek-flash, 500 on deepseek-v4-pro. For batch-style pipelines this is friendlier than token-per-minute budgets; for bursty single requests it makes no difference.
Others
Moonshot, Z.AI, MiniMax, and Alibaba show per-account limits in their consoles. MiniMax sells a priority tier at 1.5x the standard price for latency-sensitive traffic. Bedrock and Google Cloud enforce cloud-account quotas that you raise through the cloud provider.
OpenAI-Compatibility Matrix
Most providers accept OpenAI-format /v1/chat/completions requests, so switching is a base_url and api_key change. The exceptions matter when you build against provider-specific features.
| Provider | OpenAI chat format | Anthropic format | Streaming | Tool calls |
|---|---|---|---|---|
| OpenAI | Native | No | Yes | Yes |
| Anthropic | Via OpenAI SDK compat layer | Native | Yes | Yes |
| Google Gemini | Compat endpoint | No | Yes | Yes |
| DeepSeek | Yes | Yes (/anthropic) | Yes | Yes |
| Moonshot (Kimi) | Yes | No | Yes | Yes |
| Z.AI (GLM) | Yes | No | Yes | Yes |
| MiniMax | Yes | No | Yes | Yes |
| Alibaba (Qwen) | Compat mode | No | Yes | Yes |
| OpenRouter | Native (aggregator) | No | Yes | Yes |
| Morph | Yes | No | Yes | Yes (chat models) |
DeepSeek is the first-party provider that documents both OpenAI and Anthropic request formats (https://api.deepseek.com and https://api.deepseek.com/anthropic), which lets it drop into Claude-Code-style agents without a proxy. If you point a coding agent at a custom provider, the agent must support it; Codex provider configuration covers the Codex side.
Benchmarks vs Price: What You Get per Dollar
SWE-bench Verified (real GitHub issues, verified fixes) is the most cited coding benchmark. Scores below are from the llm-stats tracker on September 22, 2026; prices are official API rates. Claude Fable 5.1, Opus 5.5, Opus 5, GPT-6, GPT-5.6, DeepSeek V4.1 Flash, Gemini 3.8 Flash, Kimi K3, GLM-5.3, and Qwen3.8 Max have no entry on the tracker yet.
| Model | SWE-bench Verified | Input/MTok | Output/MTok | Value ($/MTok out per point) |
|---|---|---|---|---|
| Claude Fable 5 | 95.0% | $10.00 | $50.00 | $0.53 |
| Claude Mythos Preview (limited) | 93.9% | restricted | restricted | n/a |
| Claude Opus 4.8 | 88.6% | $5.00 | $25.00 | $0.28 |
| Claude Opus 4.7 | 87.6% | $5.00 | $25.00 | $0.29 |
| Claude Sonnet 5 | 85.2% | $2.00 | $10.00 | $0.12 |
| DeepSeek V4 Pro Max | 80.6% | open weights | open weights | open weights |
| Gemini 3.1 Pro | 80.6% | $2.00 | $12.00 | $0.15 |
| MiniMax M3 | 80.5% | $0.30 | $1.20 | $0.015 |
| Qwen3.7 Max | 80.4% | $2.50 | $7.50 | $0.09 |
| Kimi K2.6 | 80.2% | $0.95 | $4.00 | $0.05 |
| DeepSeek V4 Flash Max | 79.0% | open weights | open weights | open weights |
An independent run tells a different story. Vals.ai scores SWE-bench Verified with one bash-only harness and archived the benchmark on September 1, 2026, because the top models saturated it. Its top entries: Claude Opus 5 at 97.0%, DeepSeek V4 Pro 0813 at 96.4%, GPT-5.6 Sol at 96.2%, GLM-5.3 at 95.4%, Claude Fable 5 at 95.0%, and Kimi K3 at 93.4%. On that harness MiniMax M3 scores 75.0%, below its self-reported 80.5%. DeepSeek V4.1 Flash shipped after the archive date and has no score.
Seven models on the tracker score between 80.2% and 80.6% on SWE-bench Verified. Of the four priced in the table, output runs from $1.20/M (MiniMax M3) to $12/M (Gemini 3.1 Pro), and DeepSeek V4 Pro Max ships the same score in open weights. Claude Sonnet 5 adds 4.6 points for $10/M output. The step to Opus 4.8 (88.6%) costs $25/M, and Fable 5 (95.0%) costs $50/M, a 42x step from MiniMax M3. Whether those points are worth it depends on whether your tasks live in the gap.
Newer models are reporting on different benchmarks. Anthropic reports Opus 5.5 at 66.4% on Terminal-Bench 4.0 and 57.8% on CursorBench 4.0 and says it matches GPT-6 Astra on Terminal-Bench 4.0 at about 40% of the cost. For the harder, standardized-harness view, see SWE-bench Pro, and compare scores only within the same harness.
The llm-stats tracker lists 116 SWE-bench Verified results, all self-reported and 0 independently verified. Treat the rankings above as vendor claims and run your own prompts before committing volume.
LLM API Free Tiers: Exact Amounts
| Provider | Free offer | Limit |
|---|---|---|
| Z.AI | GLM-4.7-Flash, GLM-4.5-Flash | Free models, no token charge |
| Alibaba Model Studio | 1M tokens for new users | 90-day validity |
| OpenAI API | Free tier in allowed geographies | $100/month usage cap |
| Google Gemini API | Free tier on 3.x Flash and Flash-Lite | Not on Gemini 3.1 Pro Preview |
| Anthropic | Small free credit for new users | Then Start tier, $500/month cap |
| Morph WarpGrep | None | $0.80 per 100K tokens |
For experimentation, the practical order is: GLM Flash models (free, no token charge), Gemini Flash free tier (current 3.8 Flash included), then Qwen's 1M tokens.
Cost Calculator: Real Workloads
Per-token prices mean nothing until mapped to usage. Two reference workloads, 30-day months, standard tier, no cache discounts (caching reduces all of these), all prompts under the long-context thresholds.
Coding agent: 50M input / 5M output tokens per day
| Model | Daily cost | Monthly cost |
|---|---|---|
| Claude Fable 5.1 | $750.00 | $22,500 |
| GPT-6 Astra | $750.00 | $22,500 |
| Claude Opus 5 | $375.00 | $11,250 |
| Claude Opus 5.5 | $300.00 | $9,000 |
| Kimi K3 | $225.00 | $6,750 |
| Gemini 3.1 Pro | $160.00 | $4,800 |
| GPT-6 Sol | $150.00 | $4,500 |
| Claude Sonnet 5 | $150.00 | $4,500 |
| GLM-5.3 | $92.00 | $2,760 |
| Gemini 3.8 Flash (promo) | $56.25 | $1,688 |
| DeepSeek V4.1 Flash (peak) | $21.00 | $630 |
| MiniMax M3 | $21.00 | $630 |
| GLM-5.3-Flash | $10.00 | $300 |
| Morph morph-dsv4flash | $9.10 | $273 |
| GPT-6 Luna | $7.50 | $225 |
| Morph morph-glm53flash | $13.50 | $405 |
Support chatbot: 20M input / 5M output tokens per day
| Model | Daily cost | Monthly cost |
|---|---|---|
| Claude Haiku 4.5 | $45.00 | $1,350 |
| Gemini 3.8 Flash (promo) | $33.75 | $1,013 |
| Gemini 3.5 Flash-Lite | $18.50 | $555 |
| DeepSeek V4.1 Flash (peak) | $12.00 | $360 |
| MiniMax M3 | $12.00 | $360 |
| GLM-5.3-Flash | $5.50 | $165 |
| Morph morph-dsv4flash | $4.84 | $145 |
| GPT-6 Luna | $4.50 | $135 |
| Morph morph-glm53flash | $7.50 | $225 |
The same coding-agent workload costs $22,500/month on Claude Fable 5.1 or GPT-6 Astra and $405/month on Morph's morph-glm53flash, a 56x gap. The production answer is rarely either extreme: route routine edits to a cheap model, escalate multi-file refactors to a frontier one, and cache aggressively (DeepSeek V4.1 Flash cache hits bill input at $0.006/M at peak; Anthropic cache reads are 0.1x or less and do not count against rate limits). Model the routing split with the LLM cost calculator.
Latency and Throughput
Two metrics matter: time to first token (how fast streaming starts) and tokens per second (how fast it finishes). Provider speed claims vary with load and region, so measure on your own traffic. What the official docs establish:
- Speed is now a paid SKU at both frontier labs. OpenAI Fast mode costs 2x standard on GPT-6 and GPT-5.6. Anthropic fast mode (research preview, Claude API only, not with batch) costs $8/$40 on Opus 5.5 and $10/$50 on Opus 5 and Opus 4.8.
- Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5 at standard speed.
- Reasoning adds latency by design. Fable 5.1 and Opus 5.5 run adaptive thinking that is always on; GLM-5.3 reasoning is mandatory and defaults to
maxeffort. Thinking tokens are generated and billed before visible output. - MiniMax sells a priority tier at 1.5x standard for faster scheduling on M3.
- Specialized models break the general-purpose ceiling: Morph Fast Apply sustains 10,500 tok/s on code-edit application because the task (merging an edit into a file) does not need frontier reasoning.
Specialized APIs: When General-Purpose Falls Short
Coding agents spend much of their compute on two mechanical operations: searching codebases for context and applying edits to files. Both run through general-purpose LLMs by default, at general-purpose prices and speeds.
Morph Fast Apply
Code-edit application at 10,500+ tok/s with 98% accuracy and sub-second cold starts. The agent outputs a lazy edit snippet; Fast Apply merges it into the full file. OpenAI-compatible /v1/chat/completions endpoint.
Morph WarpGrep
Agentic codebase search, #1 on SWE-Bench Pro. $0.80 per 100K tokens. Ships as an MCP server for any agent.
The pattern generalizes: a frontier model reasons and decides what to change; narrow, fast models execute the mechanical steps. The frontier model's output shrinks (edit snippets instead of whole files), and output is the token class that costs $10 to $50/M. See Fast Apply and WarpGrep.
Best LLM API by Use Case
One ranking does not fit every job. Pick by the constraint that binds you. Each pick below comes from the data already on this page.
Best for reasoning and hard coding
Claude Fable 5.1 ($10/$50; Fable 5 scores 95.0% on SWE-bench Verified) and GPT-6 Astra ($10/$50) lead. Claude Opus 5.5 ($4/$20) is Anthropic's recommended default and, per Anthropic, matches Astra on Terminal-Bench 4.0 at about 40% of the cost. Use these for multi-file refactors and tasks where a cheaper model measurably fails.
Best value
GPT-6 Sol and Claude Sonnet 5 are both $2/$10. Sonnet 5 scores 85.2% on SWE-bench Verified at $0.12 of output per benchmark point versus $0.28 for Opus 4.8 and $0.53 for Fable 5. MiniMax M3 ($0.30/$1.20) scores 80.5%, the cheapest model above 80%, at $0.015 per point. DeepSeek V4 Pro Max (80.6%, open weights) matches it for self-hosting.
Best for speed
For code edits, Morph Fast Apply: 10,500 tok/s with 98% accuracy, because merging an edit into a file does not need frontier reasoning. For frontier chat speed, both OpenAI (Fast mode, 2x price) and Anthropic (fast mode on Opus 5.5, $8/$40) sell paid fast tiers.
Morph Fast Apply: best for code edits
10,500 tok/s, 98% accuracy. The frontier model decides what to change; Fast Apply merges the edit snippet into the full file behind an OpenAI-compatible endpoint. See /products/fastapply.
Morph model router
Route each request to the cheapest model that can handle it, billed at $0.005 per request. One key, OpenAI-compatible. See /llm-router.
Best on a budget
Morph morph-glm53flash ($0.20/$0.70, 1M context, GLM-5.3-Flash at 16-bit) is the price floor among current-generation models, followed by GPT-6 Luna ($0.10/$0.50), morph-dsv4flash ($0.142/$0.40), Qwen3.8 Flash ($0.15/$0.47), and GLM-5.3-Flash on Z.AI ($0.15/$0.50). DeepSeek's V4.1 Flash is $0.30/$1.20 at peak and $0.15/$0.60 off-peak. GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.AI API. For high-volume simple traffic, start here and escalate only the requests that measurably need a frontier model.
How to Choose an LLM API
- 1. Establish your quality floor cheaply. Run your real prompts through morph-glm53flash ($0.70/M out), GPT-6 Luna ($0.50/M), DeepSeek V4.1 Flash ($1.20/M peak), and MiniMax M3 ($1.20/M). If one passes, you are done at a fraction of frontier cost.
- 2. Escalate only measured gaps. Move to GPT-6 Sol or Claude Sonnet 5 ($10/M), Gemini 3.1 Pro ($12/M), Opus 5.5 ($20/M), or Fable 5.1 and GPT-6 Astra ($50/M) for the tasks where the cheap tier measurably fails.
- 3. Check the constraint that binds you. Rate-limited on day one? Anthropic's Start tier gives Opus 5.5 2M ITPM but Fable 5.x only 500K. Prompts over 272K? GPT-6's long-context tier applies. Need self-hosting or data control? DeepSeek, GLM-5.3, and MiniMax M3 are open weights. Compliance? Bedrock, Google Cloud, or Foundry.
- 4. Use specialized APIs for mechanical steps. Edit application, search, embeddings, and reranking all have purpose-built models that beat $10-50/M generalists on both speed and cost.
- 5. Pin model IDs and re-run evals on release days. DeepSeek's
deepseek-v4-flashID now serves a different model. Aliases move.
The most common mistake is anchoring on one provider's flagship and never testing down. MiniMax M3, Kimi K2.6, Qwen3.7 Max, and Gemini 3.1 Pro all score 80.2% to 80.6% on SWE-bench Verified for $1.20 to $12 per million output; the frontier costs $20 to $50. The second most common is ignoring caching: Anthropic bills cached reads at 0.1x (0.05x on Opus 5.5, 0.025x on Fable 5.1) and exempts them from rate limits, and DeepSeek V4.1 Flash drops cached input to $0.006/M. For long-context workloads, compare windows in detail at LLM context window comparison.
The fastest endpoints are private deployments
Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 100 billion tokens per day run this way.
Frequently Asked Questions
What is an LLM API?
An LLM API is an HTTP interface to a hosted large language model. You send a prompt (plus optional tools, images, or files) to an endpoint such as /v1/chat/completions and receive generated tokens back, billed per million tokens of input and output. It replaces running model weights on your own GPUs with a metered, pay-per-token service.
What are the best LLM API providers in 2026?
The major first-party providers are OpenAI (GPT-6 Astra $10/$50, GPT-6 Sol $2/$10, GPT-6 Luna $0.10/$0.50 per 1M tokens), Anthropic (Claude Fable 5.1 $10/$50, Opus 5.5 $4/$20, Sonnet 5 $2/$10), Google (Gemini 3.1 Pro $2/$12, Gemini 3.8 Flash $0.75/$3.75 through December 31, 2026), DeepSeek (V4.1 Flash $0.30/$1.20 at peak), Moonshot (Kimi K3 $3/$15), Z.AI (GLM-5.3 $1.40/$4.40), MiniMax (M3 $0.30/$1.20), and Alibaba (Qwen3.8 Max $2/$6). OpenRouter aggregates them behind one key; Bedrock, Google Cloud, and Microsoft Foundry resell Claude with cloud billing. Morph serves open-weight models such as GLM-5.3 ($1.19/$3.74) and Kimi K3 ($2.50/$14.00) at 16-bit precision behind one OpenAI-compatible key.
What is the cheapest LLM API in 2026?
Among current-generation models, Morph's morph-glm53flash (GLM-5.3-Flash at 16-bit) at $0.20/M input and $0.70/M output, then OpenAI's GPT-6 Luna at $0.10/$0.50 (released September 22, 2026), Morph's morph-dsv4flash at $0.142/$0.40, Alibaba's Qwen3.8 Flash at $0.15/$0.47, and GLM-5.3-Flash on Z.AI at $0.15/$0.50. DeepSeek V4.1 Flash is $0.30/$1.20 at peak and $0.15/$0.60 off-peak. MiniMax M3 at $0.30/$1.20 is the cheapest model with a published SWE-bench Verified score above 80%. GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.AI API.
Which LLM API is best for coding?
Claude Fable 5 has the top SWE-bench Verified score on the llm-stats tracker at 95.0% ($10/$50); its successor Fable 5.1 (September 1, 2026, same price) has no tracker score yet. Claude Opus 5.5 (September 22, 2026, $4/$20) scores 66.4% on Terminal-Bench 4.0, which Anthropic says matches GPT-6 Astra at about 40% of the cost. Claude Sonnet 5 scores 85.2% at $2/$10. On a budget, MiniMax M3 (80.5%, $0.30/$1.20) and open-weights DeepSeek V4 Pro Max (80.6%) are the value picks. For applying code edits, Morph Fast Apply runs at 10,500 tok/s with 98% accuracy.
What rate limits do LLM APIs have?
Anthropic now uses Start, Build, and Scale tiers. On the Start tier, Opus 5.5, Opus 5, Sonnet 5, and Haiku 4.5 each get 1,000 RPM and 2M input tokens per minute, while Fable 5.x gets 1,000 RPM and 500K ITPM. Scale gives 10,000 RPM and 10M ITPM. Cached input tokens do not count toward Anthropic's limits. OpenAI tiers unlock by spend ($5 to $1,000 paid); GPT-6 Astra runs 500 RPM / 500K TPM at Tier 1 and 15,000 RPM / 40M TPM at Tier 5. DeepSeek caps concurrency instead: 2,500 parallel requests on deepseek-flash, 500 on deepseek-v4-pro.
How do I handle LLM API rate limits?
Read the 429 response before retrying. Rate-limit 429s carry a retry-after header; honor it with exponential backoff. Anthropic's monthly spend-cap 429 has the same error type but no retry-after header and error.details.error_code set to enforced_spend_limit_reached, so SDK auto-retries keep failing until the first of the next month. Cache repeated context (Anthropic does not count cache reads toward ITPM), ramp traffic gradually to avoid acceleration limits, and keep a second provider behind an OpenAI-compatible router for overflow.
Which LLM APIs have a free tier?
Z.AI serves GLM-4.7-Flash and GLM-4.5-Flash free. Alibaba Model Studio gives new users 1M free tokens valid for 90 days. OpenAI has a Free API tier in allowed geographies with a $100/month usage cap. Google offers a free tier on Gemini 3.8, 3.7, 3.6, and 3.5 Flash and Flash-Lite, but not on Gemini 3.1 Pro Preview. Anthropic gives new users a small amount of free credits to test the API.
Which LLM API has the largest context window?
GPT-6 Astra, Sol, and Luna list a 1,050,000-token context window (922,000 max input, 128,000 max output). Claude Fable 5.1, Opus 5.5, and Sonnet 5 have 1M tokens at standard pricing with no long-context surcharge. Kimi K3 is 1,048,576 tokens. DeepSeek V4.1 Flash, GLM-5.3, and Qwen3.8 Max support 1M. Morph serves GLM-5.3 and DeepSeek V4 Flash with 1M and 1M contexts. Below 1M: Kimi K2.6 and K2.7 Code at 262,144 and Claude Haiku 4.5 at 200K.
Are LLM APIs OpenAI-compatible?
Mostly yes. DeepSeek, Moonshot (Kimi), Z.AI (GLM), MiniMax, Alibaba (Qwen), OpenRouter, and Morph accept OpenAI-format /v1/chat/completions requests, so switching providers is usually a base_url and api_key change. Google exposes an OpenAI-compatible endpoint alongside its native API. Anthropic's native API is Messages-format with an OpenAI SDK compatibility layer. DeepSeek also accepts Anthropic-format requests at api.deepseek.com/anthropic.
Why are my LLM API costs so high?
Four usual causes. Output tokens cost 3x to 6x input, and reasoning models bill thinking tokens as output. Long prompts cross price thresholds: GPT-6 bills 2x input and 1.5x output above 272K input tokens, and Gemini 3.1 Pro rises from $2/$12 to $4/$18 above 200K. Cache misses: a cache hit is 10x cheaper on OpenAI and Anthropic and 50x cheaper on DeepSeek V4.1 Flash. And newer Claude models (Opus 4.7 onward) use a tokenizer that produces about 30% more tokens for the same text.
Should I use one LLM API provider or multiple?
Production systems usually route between two or more. Send high-volume simple traffic to a sub-$1.50/MTok model (GPT-6 Luna, DeepSeek V4.1 Flash, MiniMax M3) and escalate hard tasks to a frontier model (Fable 5.1, Opus 5.5, GPT-6 Astra). Because most providers are OpenAI-compatible, a router or gateway makes the switch a config change rather than a rewrite.
Which LLM API is best by use case?
Best for reasoning and hard coding: Claude Fable 5.1 ($10/$50), Claude Opus 5.5 ($4/$20), and GPT-6 Astra ($10/$50). Best value: Claude Sonnet 5 (85.2% SWE-bench Verified, $2/$10), GPT-6 Sol ($2/$10), and MiniMax M3 ($0.30/$1.20). Best for speed: Morph Fast Apply for code edits at 10,500 tok/s with 98% accuracy. Best on a budget: Morph morph-glm53flash ($0.20/$0.70), GPT-6 Luna ($0.10/$0.50), or the free GLM-4.7-Flash and GLM-4.5-Flash on Z.AI.
What changed in LLM API pricing in September 2026?
OpenAI released GPT-6 Astra ($10/$50) on September 3 and GPT-6 Sol ($2/$10) and Luna ($0.10/$0.50) on September 22, half the price of GPT-5.6 Sol and Luna. Anthropic released Claude Fable 5.1 ($10/$50) on September 1 and Claude Opus 5.5 ($4/$20) on September 22, now its recommended default. DeepSeek retired V4 Flash for V4.1 Flash on September 10 at $0.30/$1.20 peak. Google shipped Gemini 3.8 Flash on September 2 at $0.75/$3.75 through December 31, 2026. xAI released Grok 4.7 on September 21 at $2/$6 with a 500K context.
Sources
Every third-party price, limit, and benchmark on this page traces to a primary source, checked September 22, 2026. Morph prices come from the live Morph model list.
- OpenAI API pricing and models: developers.openai.com pricing, GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, rate limits
- GPT-6 Sol and Luna release coverage: The Next Web; long-context discussion on Hacker News
- Anthropic: pricing, models overview, rate limits, Introducing Claude Opus 5.5
- Google Gemini API pricing: ai.google.dev/gemini-api/docs/pricing; Gemini 3.8 Flash launch: Google blog
- DeepSeek: pricing, V4.1 Flash release note, API docs; model-card discussion on Hacker News
- Z.AI: pricing, GLM-5.3
- Moonshot Kimi pricing: platform.kimi.ai, platform changelog
- xAI pricing: docs.x.ai
- Qwen Cloud: Qwen3.8 Max, Qwen3.8 Flash
- MiniMax pay-as-you-go pricing: platform.minimax.io
- Alibaba Model Studio pricing: alibabacloud.com
- Morph model pricing: morphllm.com/api/models/json
- SWE-bench Verified leaderboard: llm-stats.com; independent harness: Vals.ai
Related Resources
Code Editing at 10,500 tok/s
Frontier LLM APIs bill $10-50 per million output tokens to rewrite whole files. Morph Fast Apply merges edit snippets into files at 10,500 tok/s with 98% accuracy, behind an OpenAI-compatible API.
