Models
Every model Morph serves, always current. Frontier open weights on custom kernels, plus the specialized agent APIs that offload an agent's bottlenecks. One endpoint, one key.
Open-source coding models
Frontier open weights, served on Morph's custom kernels for codegen. Run the primary agent loop.
morph-kimik3-fastThe same Kimi K3, served on hardware tuned for latency over throughput.morph-dsv4flashFast MoE with compressed attention for coding, reasoning, tool use.| Model | Input | Cache Read | Output | Context | Max Out |
|---|---|---|---|---|---|
| $2.80/1M | $0.29/1M | $14.00/1M | 1M | 1M | |
Kimi K3 2.8T Fast morph-kimik3-fastThe same Kimi K3, served on hardware tuned for latency over throughput. | $6.00/1M | $0.60/1M | $22.50/1M | 1M | 1M |
| $1.25/1M | $0.26/1M | $4.40/1M | 1M | 1M | |
| $0.15/1M | $0.01/1M | $0.42/1M | 1M | 1M | |
DeepSeek V4 Flash 0731 morph-dsv4flashFast MoE with compressed attention for coding, reasoning, tool use. | $0.10/1M | — | $0.28/1M | 1M | 1M |
Specialized agent models
Purpose-built APIs that offload the bottlenecks: code editing, search, context management, per-turn classification, and routing.
| Model | Product | Speed | Price | Context |
|---|---|---|---|---|
Fast Apply | 10,500+ tok/s | $0.80 in $1.20 out /1M | 262k | |
Fast Apply | 5,000+ tok/s | $0.90 in $1.90 out /1M | 262k | |
Code Search | — | $0/100k · $1/1M req | — | |
Compaction | 33,000 tok/s | $0.20 in $0.50 out /1M | 1M | |
Reflex | < 90ms | $0.001/event · realtime | 64k | |
Router | — | $0.005/request | — |
Models we've analyzed
Deep-dives on models Morph doesn't serve on public endpoints. Architecture, benchmarks, and where to run them.
The fastest endpoints are private deployments
Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 100 billion tokens per day run this way.