Usage-based pricing · start free
Pricing
- Per-token
- usage-based billing
- No limits
- practically no rate limits
- 200 req
- free every month
Open Source Models
Frontier open weights, served on custom kernels for codegen.
Chat
Kimi K3 2.8TFeatured
morph-kimik3Speed~100 tok/sec
Input$2.90/1M
Cache read$0.29/1M
Output$14.00/1M
Context1M
Kimi K3 2.8T, latency-tuned
morph-kimik3-fastSpeedfastest
Input$6.00/1M
Cache read$0.60/1M
Output$22.50/1M
Context1M
GLM-5.2 744B
morph-glm52-744bSpeed~80 tok/sec
Input$1.10/1M
Cache read$0.22/1M
Output$4.10/1M
Context1M
MiniMax M3
morph-minimax3-428bSpeed~90 tok/sec
Input$0.30/1M
Output$1.20/1M
Context256k
Qwen 3.5 397B
morph-qwen35-397bSpeed~180 tok/sec
Input$0.50/1M
Cache read$0.30/1M
Output$3.50/1M
Context262k
DeepSeek V4 Flash 0731
morph-dsv4flashSpeed~150 tok/sec
Input$0.14/1M
Cache read$0.03/1M
Output$0.28/1M
Context1M
Gemma 4 31B
morph-gemma4-31bSpeed~120 tok/sec
Input$0.14/1M
Cache read$0.08/1M
Output$0.40/1M
Context175k
Qwen 3.6 27B
morph-qwen36-27bSpeed~100 tok/sec
Input$0.29/1M
Output$2.40/1M
Context131k
| Model | Speed | Input | Cache Read | Output | Context | |
|---|---|---|---|---|---|---|
Kimi K3 2.8TFeatured morph-kimik3 | ~100 tok/sec | $2.90/1M | $0.29/1M | $14.00/1M | 1M | |
Kimi K3 2.8T, latency-tuned morph-kimik3-fast | fastest | $6.00/1M | $0.60/1M | $22.50/1M | 1M | |
GLM-5.2 744B morph-glm52-744b | ~80 tok/sec | $1.10/1M | $0.22/1M | $4.10/1M | 1M | |
MiniMax M3 morph-minimax3-428b | ~90 tok/sec | $0.30/1M | — | $1.20/1M | 256k | |
Qwen 3.5 397B morph-qwen35-397b | ~180 tok/sec | $0.50/1M | $0.30/1M | $3.50/1M | 262k | |
DeepSeek V4 Flash 0731 morph-dsv4flash | ~150 tok/sec | $0.14/1M | $0.03/1M | $0.28/1M | 1M | |
Gemma 4 31B morph-gemma4-31b | ~120 tok/sec | $0.14/1M | $0.08/1M | $0.40/1M | 175k | |
Qwen 3.6 27B morph-qwen36-27b | ~100 tok/sec | $0.29/1M | — | $2.40/1M | 131k |
Specialized Models
Purpose-built APIs that offload the bottlenecks: code editing, search, context management, and per-turn classification.
Fast Apply
Merges LLM code edits into files at 98% accuracy.Fastest model
morph-v3-fastSpeed10,500+ tok/sec
Price
$0.80/1M in$1.20/1M out
Context262k
Most diversePopular
morph-v3-largeSpeed5,000+ tok/sec
Price
$0.90/1M in$1.90/1M out
Context262k
Code Search
Agentic repo search. #1 on SWE-Bench Pro.Fast context for agentsNew
morph-warp-grep-v2Price
$0.80/100K
Context100K (1M for Pro)
Compaction
Verbatim context compression, shrinking context 50–70%.Verbatim context compaction
morph-compactSpeed< 2s P99
Price
$0.20/1M in$0.50/1M out
Context1M
Reflex
Per-turn classifiers for jailbreaks, looping, and frustration.Realtime per-turn classifiers
morph-reflex-v1Speed< 90ms
Price
$0.001/event$0.0005 over 1M/mo
Context64K
Batch (offline)
morph-reflex-v1Price
$0.0005/event$0.00025 over 1M/mo
Context64K
| Model | Speed | Price | Context | |
|---|---|---|---|---|
Fast Apply Merges LLM code edits into files at 98% accuracy. | ||||
Fastest model morph-v3-fast | 10,500+ tok/sec | $0.80/1M in$1.20/1M out | 262k | |
Most diversePopular morph-v3-large | 5,000+ tok/sec | $0.90/1M in$1.90/1M out | 262k | |
Code Search Agentic repo search. #1 on SWE-Bench Pro. | ||||
Fast context for agentsNew morph-warp-grep-v2 | — | $0.80/100K | 100K (1M for Pro) | |
Compaction Verbatim context compression, shrinking context 50–70%. | ||||
Verbatim context compaction morph-compact | < 2s P99 | $0.20/1M in$0.50/1M out | 1M | |
Reflex Per-turn classifiers for jailbreaks, looping, and frustration. | ||||
Realtime per-turn classifiers morph-reflex-v1 | < 90ms | $0.001/event$0.0005 over 1M/mo | 64K | |
Batch (offline) morph-reflex-v1 | — | $0.0005/event$0.00025 over 1M/mo | 64K | |
Router
What is the Router?Difficulty-based model routing
morph-routerPrice$0.005/request
Context—
| Model | Price | Context | |
|---|---|---|---|
Difficulty-based model routing morph-router | $0.005/request | — |
Billing
Pay per token, or go flat-rate with Scale.
Pay as you go
Top up credits, use any model
From $10
First top-up+$5 free
LimitsPractically no rate limits
Get Started
Scale
For individuals who can't stop coding
$200/month
Credits40M($400)
LimitsPractically no rate limits
Dedicated Deployments
Single-tenant GPUs running your model on custom kernels. Billed by the minute.
Dedicated GPUs
B200
NVIDIA Blackwell$0.1663/min
GB300Fastest
NVIDIA Grace Blackwell| GPU | Accelerator | Price |
|---|---|---|
B200 | NVIDIA Blackwell | $0.1663/min |
GB300Fastest | NVIDIA Grace Blackwell | Contact us |
Our fastest endpoints are private deployments
Over 100 billion tokens per day served on dedicated capacity with custom speculators and caching, at large discounts over the public pricing above. Includes SSO, custom rate limits, and priority support.
Get in touch