NVIDIA B200: Specs, Price per GPU-Hour, and B200 vs H200 (October 2026)

The NVIDIA B200 carries 180 GB of HBM3e at 8 TB/s per GPU and 9 PFLOPS of dense FP4. It rents for $6.25 to $6.89 per GPU-hour on Modal, RunPod, and Lambda in October 2026, against about $4.55 for an H200. Specs, cloud prices, B200 vs H200 vs H100, and GB200 NVL72.

October 6, 2026 · 2 min read

The NVIDIA B200 is a Blackwell data center GPU with 180 GB of HBM3e at 8 TB/s per GPU, fifth-generation NVLink at 1.8 TB/s, and hardware FP4. An eight-GPU DGX B200 holds 1,440 GB of GPU memory and delivers 72 PFLOPS of dense FP4. Renting one B200 costs $6.25 to $6.89 per GPU-hour in October 2026 (Modal, RunPod, Lambda), about 1.4x an H200, and NVIDIA publishes no list price for a single card. Per TB/s of memory bandwidth, the number that sets LLM decode speed, the B200 rents for less than the H200 or H100.

180 GB
HBM3e per B200 (1,440 GB per DGX)
8 TB/s
Memory bandwidth per GPU
9 PFLOPS
Dense FP4 per GPU (18 sparse)
$6.25/hr
Lowest on-demand B200 (Modal, Oct 6)

NVIDIA B200 Specs

NVIDIA publishes B200 numbers at the system level, for the eight-GPU DGX B200. The per-GPU column divides those totals by eight. NVIDIA marks its FP4 figure as sparse and dense and its FP8 figure as sparse only, with dense at half.

NVIDIA B200 specifications (from the DGX B200 spec sheet)
SpecPer B200 GPUDGX B200 (8 GPUs)
ArchitectureBlackwell8x Blackwell
GPU memory180 GB HBM3e1,440 GB
Memory bandwidth8 TB/s64 TB/s
FP4 Tensor Core (dense / sparse)9 / 18 PFLOPS72 / 144 PFLOPS
FP8 Tensor Core (dense / sparse)4.5 / 9 PFLOPS36 / 72 PFLOPS
NVLink5th gen, 1.8 TB/s2x NVSwitch, 14.4 TB/s aggregate
Powern/a~14.3 kW max
Host CPU and RAMn/a2x Xeon Platinum 8570 (112 cores), 2 TB
Form factorSXM10 RU
180 GB or 192 GB?
Some spec pages list the B200 at 192 GB. The shipping DGX B200 reports 1,440 GB across eight GPUs, and the Lambda and RunPod B200 instances both list 180 GB per GPU. Plan capacity on 180 GB.

NVIDIA B200 Price: Cloud Rental per GPU-Hour (October 2026)

These are the published on-demand prices for one B200, read from each provider's pricing page on October 6, 2026. Modal bills per second and charges CPU and RAM separately; RunPod and Lambda bundle host CPU, RAM, and disk into the hourly rate.

B200 price per GPU-hour by provider (October 6, 2026)
ProviderPlanPrice per GPU-hourWhat is included
ModalServerless, per second ($0.001736/s)$6.25GPU only
RunPodB200 pod, 180 GB$6.7928 vCPU, 283 GB RAM
Lambda8x B200 SXM6 instance$6.79104 vCPU, 1,440 GiB RAM, 11 TiB SSD
Lambda4x B200 SXM6 instance$6.8952 vCPU, 720 GiB RAM, 5.5 TiB SSD
Lambda1-Click Cluster, 2 weeks to 1 year$8.87 (256+ GPUs) to $9.86 (16 GPUs)InfiniBand cluster
MorphDedicated inference endpoint$9.98Served model, OpenAI-compatible API

To buy, NVIDIA does not list a price for a single B200. Northflank reported OEM quotes of $45,000 to $50,000 per GPU and complete 8x B200 servers above $500,000 in August 2025. At $6.79 per hour, $45,000 buys about 6,600 GPU-hours of rental, or nine months of one GPU running around the clock, before power, cooling, and the host server.

B200 vs H200 vs H100

The H200 and H100 share Hopper compute and differ in memory. The B200 moves to Blackwell, adds FP4, and raises memory and bandwidth again. For more on the Hopper pair, see H100 vs H200 and the NVIDIA H200 guide.

B200 vs H200 vs H100 (SXM, per GPU)
SpecB200H200H100
ArchitectureBlackwellHopperHopper
Memory180 GB HBM3e141 GB HBM3e80 GB HBM3
Bandwidth8 TB/s4.8 TB/s3.35 TB/s
FP8 (sparse)9 PFLOPS3.96 PFLOPS3.96 PFLOPS
FP4 in hardwareYesNoNo
NVLink per GPU1.8 TB/s900 GB/s900 GB/s
Rent (Oct 6, 2026)$6.25 to $6.89$4.54 to $4.59$3.95 to $4.19

H200 rental prices are Modal ($0.001261/s) and RunPod ($4.59). H100 prices are Modal ($0.001097/s) and Lambda's 8x and 4x instances ($4.09 and $4.19).

B200 Bandwidth per Dollar: Why It Wins on Decode

Each decode step reads the model weights and the KV cache from HBM, so tokens per second on a fixed model scales with memory bandwidth more than with FLOPS. Dividing the hourly price by bandwidth gives a rough cost of decode capacity. At the RunPod and Lambda prices above, the B200 is the cheapest of the three.

Rental price per TB/s of memory bandwidth (per GPU-hour)

Hourly price divided by memory bandwidth. Lower is better.

1
H100 ($4.09 / 3.35 TB/s)
$1.22
1.22$
2
H200 ($4.59 / 4.8 TB/s)
$0.96
0.96$
3
B200 ($6.79 / 8 TB/s)
$0.85
0.85$

RunPod B200 and H200, Lambda 8x H100 SXM, read October 6, 2026. Real throughput also depends on batch size, kernels, and precision; FP4 widens the B200 gap where model quality holds.

The ratio breaks down for small models at low concurrency, where none of the three GPUs is bandwidth-limited and the cheaper hour wins. Measured B200 and B300 serving numbers for open models are on the dedicated inference benchmarks page.

What Fits on 8x B200 (1,440 GB)

FP8 weights take about one byte per parameter, FP4 about half a byte. GLM-5.2 at 753B parameters needs roughly 755 GB in FP8. On an 8x B200 node that leaves about 685 GB for KV cache and activations; on an 8x H200 node (1,128 GB) it leaves about 373 GB, and it does not fit on an 8x H100 node (640 GB). A dense 70B model in FP8 fits on one B200 with about 110 GB left for context.

Approximate memory left after FP8 weights, 8-GPU node
Model size (FP8)8x B200 (1,440 GB)8x H200 (1,128 GB)8x H100 (640 GB)
70B (~70 GB)~1,370 GB~1,058 GB~570 GB
400B (~400 GB)~1,040 GB~728 GB~240 GB
753B GLM-5.2 (~755 GB)~685 GB~373 GBDoes not fit
1.2T (~1,200 GB)~240 GBDoes not fitDoes not fit

GB200 NVL72 and B300: The Other Blackwell Options

GB200 NVL72 is a liquid-cooled rack of 36 Grace CPUs and 72 Blackwell GPUs in a single NVLink domain. NVIDIA lists 130 TB/s of GPU-to-GPU bandwidth and 720 PFLOPS of dense NVFP4 for the rack. Each GB200 Superchip pairs one Grace CPU with two Blackwell GPUs. The rack exists for models too large or too communication-heavy for an eight-GPU node.

The B300 raises memory to 288 GB of HBM3e per GPU. It rents for $7.10 per GPU-hour on Modal ($0.001972/s) and $7.89 on RunPod, a 14% to 16% premium over the B200 for 1.6x the memory. NVIDIA's next generation is Rubin; its HGX page projects 10x the token throughput of HGX B200 for HGX Rubin NVL8 and labels that projection subject to change.

Rent, Buy, or Use a Dedicated B200 Endpoint

Renting by the hour suits bursty training and evaluation. Buying needs a site that can deliver 14.3 kW per DGX and enough utilization to beat about $6.79 per GPU-hour. For serving an open model in production, a dedicated endpoint removes the serving stack: Morph runs reserved B200 and B300 capacity behind an OpenAI-compatible API at $9.98 per B200-hour. See dedicated inference for models and regions.

Frequently Asked Questions

What are the NVIDIA B200 specs?

Per GPU, the B200 has 180 GB of HBM3e at 8 TB/s, fifth-generation NVLink at 1.8 TB/s, 9 PFLOPS of dense FP4 (18 with sparsity), and 4.5 PFLOPS of dense FP8 (9 with sparsity). Those figures are NVIDIA's DGX B200 system specs divided by its eight GPUs. The full system has 1,440 GB of GPU memory and 64 TB/s of bandwidth. It delivers 72 PFLOPS of dense FP4 and draws about 14.3 kW.

How much does an NVIDIA B200 cost?

To rent, $6.25 per GPU-hour on Modal, $6.79 on RunPod and on Lambda's 8x instance, and $6.89 on Lambda's 4x instance, as listed on October 6, 2026. Lambda's 1-Click Cluster contracts run $8.87 to $9.86 per GPU-hour depending on size. NVIDIA does not publish a list price for a single B200. Northflank reported OEM quotes of $45,000 to $50,000 per GPU, and full 8x B200 servers above $500,000, in August 2025.

How much memory does the B200 have?

180 GB of HBM3e per GPU in shipping systems. NVIDIA's DGX B200 lists 1,440 GB across eight GPUs, and the Lambda and RunPod B200 instances both expose 180 GB per GPU. Some third-party spec sheets quote 192 GB. Bandwidth is 8 TB/s per GPU.

B200 vs H200: which is better for LLM inference?

The B200 has 1.28x the memory (180 GB vs 141 GB), 1.67x the bandwidth (8 TB/s vs 4.8 TB/s), and FP4 support the H200 lacks. It rents for about 1.4x the price ($6.25 to $6.89 vs $4.54 to $4.59 per GPU-hour), so it is cheaper per TB/s of bandwidth: about $0.85 per TB/s-hour against $0.96 for the H200. The H200 stays the better buy for models that fit in 141 GB at moderate traffic, where the extra bandwidth goes unused.

What is the difference between the B200 and GB200?

B200 is the GPU. GB200 is a Grace Blackwell Superchip that pairs one NVIDIA Grace CPU with two Blackwell GPUs. GB200 NVL72 is a liquid-cooled rack of 36 Grace CPUs and 72 Blackwell GPUs in one NVLink domain with 130 TB/s of GPU-to-GPU bandwidth and 720 PFLOPS of dense NVFP4. HGX and DGX B200 are eight-GPU x86 servers.

Is the B300 worth it over the B200?

The B300 carries 288 GB of HBM3e per GPU, 1.6x the B200's 180 GB. It rents for $7.10 per GPU-hour on Modal and $7.89 on RunPod, against $6.25 and $6.79 for the B200. Take the B300 when the model plus KV cache overflows 180 GB per GPU and you would otherwise add GPUs only for memory.

Related Resources

Private deployments

The fastest endpoints are private deployments

Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 500 billion tokens per day run this way.

Talk to us about a private deployment

Serve Open Models on Reserved B200 Capacity

Morph runs dedicated endpoints on B200 and B300 GPUs behind one OpenAI-compatible API, billed at $9.98 per B200-hour. No cluster to size or serving stack to tune.

Sources