MorphMorph
Models
Open Source Models

Kimi K3

Fable-tier, 100 tok/s. #1 on Design Arena.

GLM-5.2

Opus-tier. 744B MoE, 1M context.

Qwen

27B dense, low latency.

MiniMax

230B MoE for agentic workflows.

DeepSeek

1M-context model, served fast.

Specialized

Reflex

Classify every agent trace for any behavior that matters, in under 90ms.

Fast Apply

Merge AI-generated code edits instantly.

WarpGrep

AI search subagent with sub-6s searches.

Compact

Verbatim context compaction for long-running agents.

Model Router

Auto-route each prompt to the best model.

SDK

DocsPricingMCP
Resources

Blog

Engineering deep dives and product updates.

Startup Credits

Up to $5K in API credits for startups.

Contact Us

Talk to the team about your use case.

About

How we train and deploy models.

Careers

Join a small team shipping daily.

Usage-based pricing · start free

Pricing

Per-token
usage-based billing
No limits
practically no rate limits
200 req
free every month

Open Source Models

Frontier open weights, served on custom kernels for codegen.

Chat
Kimi K3 2.8TFeatured
morph-kimik3
Speed~100 tok/sec
Input$2.90/1M
Cache read$0.29/1M
Output$14.00/1M
Context1M
Kimi K3 2.8T, latency-tuned
morph-kimik3-fast
Speedfastest
Input$6.00/1M
Cache read$0.60/1M
Output$22.50/1M
Context1M
GLM-5.2 744B
morph-glm52-744b
Speed~80 tok/sec
Input$1.10/1M
Cache read$0.22/1M
Output$4.10/1M
Context1M
MiniMax M3
morph-minimax3-428b
Speed~90 tok/sec
Input$0.30/1M
Output$1.20/1M
Context256k
Qwen 3.5 397B
morph-qwen35-397b
Speed~180 tok/sec
Input$0.50/1M
Cache read$0.30/1M
Output$3.50/1M
Context262k
DeepSeek V4 Flash 0731
morph-dsv4flash
Speed~150 tok/sec
Input$0.14/1M
Cache read$0.03/1M
Output$0.28/1M
Context1M
Gemma 4 31B
morph-gemma4-31b
Speed~120 tok/sec
Input$0.14/1M
Cache read$0.08/1M
Output$0.40/1M
Context175k
Qwen 3.6 27B
morph-qwen36-27b
Speed~100 tok/sec
Input$0.29/1M
Output$2.40/1M
Context131k
ModelSpeedInputCache ReadOutputContext
Kimi K3 2.8TFeatured
morph-kimik3
~100 tok/sec$2.90/1M$0.29/1M$14.00/1M1M
Kimi K3 2.8T, latency-tuned
morph-kimik3-fast
fastest$6.00/1M$0.60/1M$22.50/1M1M
GLM-5.2 744B
morph-glm52-744b
~80 tok/sec$1.10/1M$0.22/1M$4.10/1M1M
MiniMax M3
morph-minimax3-428b
~90 tok/sec$0.30/1M—$1.20/1M256k
Qwen 3.5 397B
morph-qwen35-397b
~180 tok/sec$0.50/1M$0.30/1M$3.50/1M262k
DeepSeek V4 Flash 0731
morph-dsv4flash
~150 tok/sec$0.14/1M$0.03/1M$0.28/1M1M
Gemma 4 31B
morph-gemma4-31b
~120 tok/sec$0.14/1M$0.08/1M$0.40/1M175k
Qwen 3.6 27B
morph-qwen36-27b
~100 tok/sec$0.29/1M—$2.40/1M131k

Specialized Models

Purpose-built APIs that offload the bottlenecks: code editing, search, context management, and per-turn classification.

Fast Apply
Merges LLM code edits into files at 98% accuracy.
What is Fast Apply?
Fastest model
morph-v3-fast
Try
Speed10,500+ tok/sec
Price
$0.80/1M in$1.20/1M out
Context262k
Most diversePopular
morph-v3-large
Try
Speed5,000+ tok/sec
Price
$0.90/1M in$1.90/1M out
Context262k
Code Search
Agentic repo search. #1 on SWE-Bench Pro.
What is WarpGrep?
Fast context for agentsNew
morph-warp-grep-v2
Try
Price
$0.80/100K
Context100K (1M for Pro)
Compaction
Verbatim context compression, shrinking context 50–70%.
What is Compact?
Verbatim context compaction
morph-compact
Try
Speed< 2s P99
Price
$0.20/1M in$0.50/1M out
Context1M
Reflex
Per-turn classifiers for jailbreaks, looping, and frustration.
What is Reflex?
Realtime per-turn classifiers
morph-reflex-v1
Try
Speed< 90ms
Price
$0.001/event$0.0005 over 1M/mo
Context64K
Batch (offline)
morph-reflex-v1
Try
Price
$0.0005/event$0.00025 over 1M/mo
Context64K
ModelSpeedPriceContext
Fast Apply
Merges LLM code edits into files at 98% accuracy.
What is Fast Apply?
Fastest model
morph-v3-fast
10,500+ tok/sec
$0.80/1M in$1.20/1M out
262k
Try
Most diversePopular
morph-v3-large
5,000+ tok/sec
$0.90/1M in$1.90/1M out
262k
Try
Code Search
Agentic repo search. #1 on SWE-Bench Pro.
What is WarpGrep?
Fast context for agentsNew
morph-warp-grep-v2
—
$0.80/100K
100K (1M for Pro)
Try
Compaction
Verbatim context compression, shrinking context 50–70%.
What is Compact?
Verbatim context compaction
morph-compact
< 2s P99
$0.20/1M in$0.50/1M out
1M
Try
Reflex
Per-turn classifiers for jailbreaks, looping, and frustration.
What is Reflex?
Realtime per-turn classifiers
morph-reflex-v1
< 90ms
$0.001/event$0.0005 over 1M/mo
64K
Try
Batch (offline)
morph-reflex-v1
—
$0.0005/event$0.00025 over 1M/mo
64K
Try
Router
What is the Router?
Difficulty-based model routing
morph-router
Price$0.005/request
Context—
ModelPriceContext
Difficulty-based model routing
morph-router
$0.005/request—

Billing

Pay per token, or go flat-rate with Scale.

Pay as you go

Top up credits, use any model

From $10
First top-up+$5 free
LimitsPractically no rate limits
Get Started

Scale

For individuals who can't stop coding

$200/month
Credits40M($400)
LimitsPractically no rate limits

Dedicated Deployments

Single-tenant GPUs running your model on custom kernels. Billed by the minute.

Dedicated GPUs
B200
NVIDIA Blackwell
$0.1663/min
GB300Fastest
NVIDIA Grace Blackwell
Contact us
GPUAcceleratorPrice
B200
NVIDIA Blackwell$0.1663/min
GB300Fastest
NVIDIA Grace BlackwellContact us

Our fastest endpoints are private deployments

Over 100 billion tokens per day served on dedicated capacity with custom speculators and caching, at large discounts over the public pricing above. Includes SSO, custom rate limits, and priority support.

Get in touch
MorphMorph

Applied research building for the future of codegen.

© 2026 AutoInfra, Inc. All rights reserved.

Y
Backed by Y Combinator
  • Documentation
  • Blog
  • Trust Center
  • CareersWe're Hiring!
  • Privacy Policy
  • Terms of Service
  • EULA
  • Service Status
  • Book a Call