Time to first token
How long a user waits before the answer starts.
Raw throughput does not tell you how an agent feels. Read latency, speed, and capacity together.
How long a user waits before the answer starts.
How quickly the answer continues after it starts.
How many requests share the system at once.
These are independent InferenceX measurements, not Morph production benchmarks. Each row links to its source run. Similar model names are not interchangeable.
Fast for one interactive agent.
View source runMore total work, but users wait longer.
View source runResponsive generation on a long agent trace.
View source runVery fast generation on this trace.
View source runFilter attributed measurements by model, GPU, and minimum generation speed. A missing cost means the source run did not report enough information to calculate it.
AgentX coding agent trace
AgentX coding agent trace
AgentX coding agent trace
AgentX coding agent trace
8K input, 1K output
8K input, 1K output at concurrency 64
Reference data is directional. Model version, workload, context length, concurrency, cache state, precision, framework, and topology must match before a result can size a production endpoint.
A highly batched run can produce impressive total throughput while making every interactive user wait. Choose an acceptable first token delay and generation speed first. Then compare the capacity that remains inside those limits.
Context length, cache reuse, framework, precision, topology, and software version can change the answer. Always compare the same workload.
Use your team size and traffic shape to estimate monthly demand, then verify with your traces.