The NVIDIA B200 is a Blackwell data center GPU with 180 GB of HBM3e at 8 TB/s per GPU, fifth-generation NVLink at 1.8 TB/s, and hardware FP4. An eight-GPU DGX B200 holds 1,440 GB of GPU memory and delivers 72 PFLOPS of dense FP4. Renting one B200 costs $6.25 to $6.89 per GPU-hour in October 2026 (Modal, RunPod, Lambda), about 1.4x an H200, and NVIDIA publishes no list price for a single card. Per TB/s of memory bandwidth, the number that sets LLM decode speed, the B200 rents for less than the H200 or H100.
NVIDIA B200 Specs
NVIDIA publishes B200 numbers at the system level, for the eight-GPU DGX B200. The per-GPU column divides those totals by eight. NVIDIA marks its FP4 figure as sparse and dense and its FP8 figure as sparse only, with dense at half.
| Spec | Per B200 GPU | DGX B200 (8 GPUs) |
|---|---|---|
| Architecture | Blackwell | 8x Blackwell |
| GPU memory | 180 GB HBM3e | 1,440 GB |
| Memory bandwidth | 8 TB/s | 64 TB/s |
| FP4 Tensor Core (dense / sparse) | 9 / 18 PFLOPS | 72 / 144 PFLOPS |
| FP8 Tensor Core (dense / sparse) | 4.5 / 9 PFLOPS | 36 / 72 PFLOPS |
| NVLink | 5th gen, 1.8 TB/s | 2x NVSwitch, 14.4 TB/s aggregate |
| Power | n/a | ~14.3 kW max |
| Host CPU and RAM | n/a | 2x Xeon Platinum 8570 (112 cores), 2 TB |
| Form factor | SXM | 10 RU |
NVIDIA B200 Price: Cloud Rental per GPU-Hour (October 2026)
These are the published on-demand prices for one B200, read from each provider's pricing page on October 6, 2026. Modal bills per second and charges CPU and RAM separately; RunPod and Lambda bundle host CPU, RAM, and disk into the hourly rate.
| Provider | Plan | Price per GPU-hour | What is included |
|---|---|---|---|
| Modal | Serverless, per second ($0.001736/s) | $6.25 | GPU only |
| RunPod | B200 pod, 180 GB | $6.79 | 28 vCPU, 283 GB RAM |
| Lambda | 8x B200 SXM6 instance | $6.79 | 104 vCPU, 1,440 GiB RAM, 11 TiB SSD |
| Lambda | 4x B200 SXM6 instance | $6.89 | 52 vCPU, 720 GiB RAM, 5.5 TiB SSD |
| Lambda | 1-Click Cluster, 2 weeks to 1 year | $8.87 (256+ GPUs) to $9.86 (16 GPUs) | InfiniBand cluster |
| Morph | Dedicated inference endpoint | $9.98 | Served model, OpenAI-compatible API |
To buy, NVIDIA does not list a price for a single B200. Northflank reported OEM quotes of $45,000 to $50,000 per GPU and complete 8x B200 servers above $500,000 in August 2025. At $6.79 per hour, $45,000 buys about 6,600 GPU-hours of rental, or nine months of one GPU running around the clock, before power, cooling, and the host server.
B200 vs H200 vs H100
The H200 and H100 share Hopper compute and differ in memory. The B200 moves to Blackwell, adds FP4, and raises memory and bandwidth again. For more on the Hopper pair, see H100 vs H200 and the NVIDIA H200 guide.
| Spec | B200 | H200 | H100 |
|---|---|---|---|
| Architecture | Blackwell | Hopper | Hopper |
| Memory | 180 GB HBM3e | 141 GB HBM3e | 80 GB HBM3 |
| Bandwidth | 8 TB/s | 4.8 TB/s | 3.35 TB/s |
| FP8 (sparse) | 9 PFLOPS | 3.96 PFLOPS | 3.96 PFLOPS |
| FP4 in hardware | Yes | No | No |
| NVLink per GPU | 1.8 TB/s | 900 GB/s | 900 GB/s |
| Rent (Oct 6, 2026) | $6.25 to $6.89 | $4.54 to $4.59 | $3.95 to $4.19 |
H200 rental prices are Modal ($0.001261/s) and RunPod ($4.59). H100 prices are Modal ($0.001097/s) and Lambda's 8x and 4x instances ($4.09 and $4.19).
B200 Bandwidth per Dollar: Why It Wins on Decode
Each decode step reads the model weights and the KV cache from HBM, so tokens per second on a fixed model scales with memory bandwidth more than with FLOPS. Dividing the hourly price by bandwidth gives a rough cost of decode capacity. At the RunPod and Lambda prices above, the B200 is the cheapest of the three.
Rental price per TB/s of memory bandwidth (per GPU-hour)
Hourly price divided by memory bandwidth. Lower is better.
RunPod B200 and H200, Lambda 8x H100 SXM, read October 6, 2026. Real throughput also depends on batch size, kernels, and precision; FP4 widens the B200 gap where model quality holds.
The ratio breaks down for small models at low concurrency, where none of the three GPUs is bandwidth-limited and the cheaper hour wins. Measured B200 and B300 serving numbers for open models are on the dedicated inference benchmarks page.
What Fits on 8x B200 (1,440 GB)
FP8 weights take about one byte per parameter, FP4 about half a byte. GLM-5.2 at 753B parameters needs roughly 755 GB in FP8. On an 8x B200 node that leaves about 685 GB for KV cache and activations; on an 8x H200 node (1,128 GB) it leaves about 373 GB, and it does not fit on an 8x H100 node (640 GB). A dense 70B model in FP8 fits on one B200 with about 110 GB left for context.
| Model size (FP8) | 8x B200 (1,440 GB) | 8x H200 (1,128 GB) | 8x H100 (640 GB) |
|---|---|---|---|
| 70B (~70 GB) | ~1,370 GB | ~1,058 GB | ~570 GB |
| 400B (~400 GB) | ~1,040 GB | ~728 GB | ~240 GB |
| 753B GLM-5.2 (~755 GB) | ~685 GB | ~373 GB | Does not fit |
| 1.2T (~1,200 GB) | ~240 GB | Does not fit | Does not fit |
GB200 NVL72 and B300: The Other Blackwell Options
GB200 NVL72 is a liquid-cooled rack of 36 Grace CPUs and 72 Blackwell GPUs in a single NVLink domain. NVIDIA lists 130 TB/s of GPU-to-GPU bandwidth and 720 PFLOPS of dense NVFP4 for the rack. Each GB200 Superchip pairs one Grace CPU with two Blackwell GPUs. The rack exists for models too large or too communication-heavy for an eight-GPU node.
The B300 raises memory to 288 GB of HBM3e per GPU. It rents for $7.10 per GPU-hour on Modal ($0.001972/s) and $7.89 on RunPod, a 14% to 16% premium over the B200 for 1.6x the memory. NVIDIA's next generation is Rubin; its HGX page projects 10x the token throughput of HGX B200 for HGX Rubin NVL8 and labels that projection subject to change.
Rent, Buy, or Use a Dedicated B200 Endpoint
Renting by the hour suits bursty training and evaluation. Buying needs a site that can deliver 14.3 kW per DGX and enough utilization to beat about $6.79 per GPU-hour. For serving an open model in production, a dedicated endpoint removes the serving stack: Morph runs reserved B200 and B300 capacity behind an OpenAI-compatible API at $9.98 per B200-hour. See dedicated inference for models and regions.
Frequently Asked Questions
What are the NVIDIA B200 specs?
Per GPU, the B200 has 180 GB of HBM3e at 8 TB/s, fifth-generation NVLink at 1.8 TB/s, 9 PFLOPS of dense FP4 (18 with sparsity), and 4.5 PFLOPS of dense FP8 (9 with sparsity). Those figures are NVIDIA's DGX B200 system specs divided by its eight GPUs. The full system has 1,440 GB of GPU memory and 64 TB/s of bandwidth. It delivers 72 PFLOPS of dense FP4 and draws about 14.3 kW.
How much does an NVIDIA B200 cost?
To rent, $6.25 per GPU-hour on Modal, $6.79 on RunPod and on Lambda's 8x instance, and $6.89 on Lambda's 4x instance, as listed on October 6, 2026. Lambda's 1-Click Cluster contracts run $8.87 to $9.86 per GPU-hour depending on size. NVIDIA does not publish a list price for a single B200. Northflank reported OEM quotes of $45,000 to $50,000 per GPU, and full 8x B200 servers above $500,000, in August 2025.
How much memory does the B200 have?
180 GB of HBM3e per GPU in shipping systems. NVIDIA's DGX B200 lists 1,440 GB across eight GPUs, and the Lambda and RunPod B200 instances both expose 180 GB per GPU. Some third-party spec sheets quote 192 GB. Bandwidth is 8 TB/s per GPU.
B200 vs H200: which is better for LLM inference?
The B200 has 1.28x the memory (180 GB vs 141 GB), 1.67x the bandwidth (8 TB/s vs 4.8 TB/s), and FP4 support the H200 lacks. It rents for about 1.4x the price ($6.25 to $6.89 vs $4.54 to $4.59 per GPU-hour), so it is cheaper per TB/s of bandwidth: about $0.85 per TB/s-hour against $0.96 for the H200. The H200 stays the better buy for models that fit in 141 GB at moderate traffic, where the extra bandwidth goes unused.
What is the difference between the B200 and GB200?
B200 is the GPU. GB200 is a Grace Blackwell Superchip that pairs one NVIDIA Grace CPU with two Blackwell GPUs. GB200 NVL72 is a liquid-cooled rack of 36 Grace CPUs and 72 Blackwell GPUs in one NVLink domain with 130 TB/s of GPU-to-GPU bandwidth and 720 PFLOPS of dense NVFP4. HGX and DGX B200 are eight-GPU x86 servers.
Is the B300 worth it over the B200?
The B300 carries 288 GB of HBM3e per GPU, 1.6x the B200's 180 GB. It rents for $7.10 per GPU-hour on Modal and $7.89 on RunPod, against $6.25 and $6.79 for the B200. Take the B300 when the model plus KV cache overflows 180 GB per GPU and you would otherwise add GPUs only for memory.
Related Resources
The fastest endpoints are private deployments
Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 500 billion tokens per day run this way.
Serve Open Models on Reserved B200 Capacity
Morph runs dedicated endpoints on B200 and B300 GPUs behind one OpenAI-compatible API, billed at $9.98 per B200-hour. No cluster to size or serving stack to tune.
Sources
- NVIDIA: DGX B200 specifications (1,440 GB, 64 TB/s, 72/144 PFLOPS FP4, ~14.3 kW)
- NVIDIA: DGX B200 User Guide, hardware overview
- NVIDIA: GB200 NVL72 specifications
- NVIDIA: HGX platform (Rubin NVL8 vs HGX B200 projection)
- Modal pricing (B200, B300, H200, H100 per-second rates), read October 6, 2026
- RunPod pricing (B200 180 GB, B300 288 GB, H200), read October 6, 2026
- Lambda pricing (B200 SXM6 instances and 1-Click Clusters), read October 6, 2026
- Jarvislabs: B200 vs H200 vs H100 spec table (NVLink 1.8 TB/s vs 900 GB/s)
- Northflank: How much does an NVIDIA B200 GPU cost? (OEM quotes, August 2025)