Usage-based pricing · start free

Pricing

Per-token
usage-based billing
No limits
practically no rate limits
200 req
free every month

Join the teams behind 400+ production agents.

Read customer stories
JetBrains
Vercel
Onlook
Webflow
Databutton
Warp
Zo Computer
Anything
Inference at work today
Private tokens / dayProduction agentsFast Apply throughputDedicated availability
Private tokens / dayProduction agentsFast Apply throughputDedicated availability

Open Source Models

Frontier open weights, served on custom kernels for codegen.

Chat
Kimi K3 2.8TFeatured
morph-kimik3
Speed~100 tok/sec
Input$2.80/1M
Cache read$0.29/1M
Output$14.00/1M
Context1M
Kimi K3 2.8T, latency-tuned
morph-kimik3-fast
Speedfastest
Input$6.00/1M
Cache read$0.60/1M
Output$22.50/1M
Context1M
GLM-5.2 744B
morph-glm52-744b
Speed~80 tok/sec
Input$1.10/1M
Cache read$0.22/1M
Output$4.10/1M
Context1M
MiniMax M3
morph-minimax3-428b
Speed~90 tok/sec
Input$0.30/1M
Output$1.20/1M
Context256k
Qwen 3.5 397B
morph-qwen35-397b
Speed~180 tok/sec
Input$0.50/1M
Cache read$0.30/1M
Output$3.50/1M
Context262k
DeepSeek V4 Flash 0731
morph-dsv4flash
Speed~150 tok/sec
Input$0.14/1M
Cache read$0.07/1M
Output$0.28/1M
Context1M
Gemma 4 31B
morph-gemma4-31b
Speed~120 tok/sec
Input$0.14/1M
Cache read$0.08/1M
Output$0.40/1M
Context175k
Qwen 3.6 27B
morph-qwen36-27b
Speed~100 tok/sec
Input$0.29/1M
Output$2.40/1M
Context131k

Specialized Models

Purpose-built APIs that offload the bottlenecks: code editing, search, context management, and per-turn classification.

Fast Apply
What is Fast Apply?
Fastest model
morph-v3-fast
Speed10,500+ tok/sec
Price
$0.80/1M in$1.20/1M out
Context262k
Most diversePopular
morph-v3-large
Speed5,000+ tok/sec
Price
$0.90/1M in$1.90/1M out
Context262k
Code Search
What is WarpGrep?
Fast context for agentsNew
morph-warp-grep-v2
Price
$0.80/100K
Context100K (1M for Pro)
Compaction
What is Compact?
Verbatim context compaction
morph-compact
Speed< 2s P99
Price
$0.20/1M in$0.50/1M out
Context1M
Reflex
What is Reflex?
Realtime per-turn classifiers
morph-reflex-v1
Speed< 90ms
Price
$0.001/event$0.0005 over 1M/mo
Context64K
Batch (offline)
morph-reflex-v1
Price
$0.0005/event$0.00025 over 1M/mo
Context64K
Difficulty-based model routing
morph-router
Price$0.005/request
Context

Billing

Pay per token, or go flat-rate with Scale.

Pay as you go

Top up credits, use any model

From $10
First top-up+$5 free
LimitsPractically no rate limits
Get Started

Scale

For individuals who can't stop coding

$200/month
Credits40M($400)
LimitsPractically no rate limits

Dedicated Endpoints

Reserve B200 or B300 capacity by the GPU-hour, from $9.50 per reserved B200-hour, behind a 99.9% monthly availability SLA. Paying for GPU time instead of tokens runs up to 7× cheaper than DeepSeek API pricing.

2× B200

$9.06/GPU-hr

up to 6.6x cheaper than token-based pricing

  • 2× B200-equivalent reserved capacity
  • Dedicated endpoint with zero data retention
  • 99.9% availability SLA
  • Migration services
  • Slack support
  • Dedicated customer success
  • 24/7 Incident monitoring

4× B200

$8.84/GPU-hr

up to 6.8x cheaper than token-based pricing

  • Everything in 2× B200
  • 2× the reserved capacity
  • Lower GPU-hour rate
  • ~2× modeled token capacity

8× B200

$8.55/GPU-hr

up to 7.0x cheaper than token-based pricing

  • Everything in 4× B200
  • 2× the reserved capacity
  • Lowest GPU-hour rate
  • ~2× modeled token capacity
  • Full-node-equivalent capacity envelope

Compare DeepSeek V4 Flash 0731

Adjust useful utilization and cache share to compare hourly capacity with the selected model's pay-per-token API rates.

Workload assumptions

Model the effective token price of hourly capacity.

Capacity utilization100%
Input served from cache65%
Hourly purchase vs. API pricing

Openrouter per-token reference is $0.14/M input, $0.0028/M cached input, and $0.28/M output.

Reserved capacityTokens / hrHourlyAPI pricingSavings
2× B200
$20.14/endpoint-hr
855.4M
760.4M in · 95M out
$20.14/hr
$0.02354/M
$0.07628/M
$65.26/hr
3.2× cheaper
4× B200
$39.31/endpoint-hr
1.71B
1.52B in · 190.1M out
$39.31/hr
$0.02298/M
$0.07628/M
$130.51/hr
3.3× cheaper
8× B200
$76.04/endpoint-hr
3.42B
3.04B in · 380.2M out
$76.04/hr
$0.02222/M
$0.07628/M
$261.03/hr
3.4× cheaper

Dedicated capacity or pay per token?

Compare Morph Dedicated with OpenRouter's pay-as-you-go API.

Economics
Billing unit
Reserved GPU-hour
Cost at sustained utilization
Effective token cost falls as utilization rises
Full-use economics
Up to 7× cheaper than per-token API pricing
Idle capacity
Reserved time remains billable
Credit purchase fee
None
Initial commitment
90 days
Best fit
Sustained traffic on a selected model
Platform
Model catalog
6 dedicated models
Serving path
Morph-operated serving and overflow
OpenAI-compatible API
Spend controls
Reliability and privacy
Zero data retention
Default for every endpoint
Prompt and response logging
Never persisted by default
Included with Morph Dedicated
Reserved B200 or B300 capacity
Isolated endpoint
No credit purchase fee
99.9% availability SLA
Migration services
Slack support
Dedicated customer success
24/7 endpoint monitoring

Dedicated FAQ

Billing, activation, ownership, and retention details.

How does hourly billing work?+

Charges accrue by reserved GPU-second while your endpoint is active, including idle time, and are invoiced monthly. Tokens are not a billing unit and there is no per-token overage charge.

Are tokens included or capped?+

No token allowance is consumed. You pay for reserved GPU time and can use the available capacity throughout that time. Calculator token figures are an illustrative economics model, not credits, measured benchmarks, or throughput guarantees.

How is the GPU-hour price calculated?+

The listed per-GPU-hour rate is the actual billing rate. The endpoint-hour rate multiplies it by the number of reserved B200-equivalents. Monthly reference costs use a standard 730-hour month; actual invoices use active reserved GPU-seconds.

Is the physical hardware exclusive to me?+

No. The endpoint, credentials, and purchased service capacity are isolated for your use, but Morph may serve other traffic on idle hardware. Your requests retain the service entitlement you purchased, and successful overflow capacity counts toward availability.

How quickly is an endpoint activated?+

Provisioning starts automatically as soon as payment is committed. An account with billing on file is charged directly; an account without one completes a bank-account checkout first. If Morph cannot activate the endpoint within two hours, the subscription is canceled and any collected payment is automatically voided or refunded.

What is the term and cancellation policy?+

The initial term is 90 days. After that, service continues month to month with 30 days’ notice. Cancellation takes effect on the first billing boundary at least 30 days after notice, without partial-month prorations.

Does Morph retain prompts or responses?+

No. Dedicated endpoints default to zero data retention: Morph does not persist prompts, responses, or model-generated content. Only bounded operational and billing metadata is retained.

Our fastest endpoints are private deployments

Over 100 billion tokens per day served on dedicated capacity with custom speculators and caching, at large discounts over the public pricing above. Includes SSO, custom rate limits, and priority support.

Get in touch