Free planning tool

LLM inference cost and capacity calculator

Compare serverless cost with dedicated capacity using your tokens, cache reuse, and required generation speed.

Workload economics
Per model decision
900M
tokens per month
100 tok/s
required generation
$63
serverless per month
$14,016
dedicated per month
Sizing review required

An exact Morph capacity measurement is required before recommending a dedicated plan.

B200 is the compatible public platform. Dedicated capacity is invoiced monthly at the beginning of the month. Tokens are not billed separately.

Difference from serverless: $13,953 more per month.

What the estimate means

Monthly demand

Input and output volume priced with your prompt cache reuse.

Capacity evidence

A dedicated recommendation only when an exact Morph measurement supports it.

Economic context

The same model and workload valued at current serverless list rates.

Use better inputs

Measure before you buy

Export a representative week of request metadata. Use actual prompt length, output length, concurrency, and cache reuse instead of averages from another product.

Choose the latency target before optimizing throughput. A configuration can process many tokens while making an interactive agent wait too long for its first token.

Learn how to read LLM inference benchmarks

Questions

How do I size an LLM inference endpoint?

Start with monthly input and output tokens, cache hit rate, and the generation speed each active user needs. Exact dedicated sizing requires a matching Morph measurement.

When is dedicated inference better than serverless?

Dedicated inference fits steady production demand that can use reserved capacity. Serverless fits new or variable workloads where avoiding idle capacity matters more.

Is this a throughput guarantee?

No. It is a planning estimate. Model architecture, context length, cache reuse, concurrency, and latency targets all change real capacity. Validate the result with production traces.