Vast.ai In stock Our sponsor
USA 9 configs 1x-8x
High-end Blackwell GPU for large-scale AI training and inference.
We're tracking 28 cloud providers for the B200, and the pricing spread is significant. While the average sits at $8.53/hr, the lowest price is currently $3.75/hr per GPU (on-demand).
USA 9 configs 1x-8x
USA 5 configs 1x-8x
USA 5 configs 1x-8x
USA 2 configs 8x
UAE 4 configs 1x-8x
Singapore 3 configs 1x-8x
USA 1 config 8x
USA 4 configs 8x
Germany 9 configs 1x-8x
France 1 config 1x
USA 2 configs 1x
USA 4 configs 1x
UK 2 configs 1x
Netherlands 2 configs 1x
Poland 7 configs 1x-8x
Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.
192GB HBM3e with 1800 GB/s NVLink. Blackwell FP4 Tensor Cores for significantly improved inference throughput over H100. Designed for training and serving the largest foundation models.
Premium pricing and limited availability as the latest generation. If your workload fits on an H100 or H200, the cost difference may not be justified.
The median on-demand price across providers has risen about 24% since August 2025, from $5.75 to $7.15/hr per GPU. This reflects the market-wide median, which moves both when providers change prices and when lower or higher priced offerings enter the market.
With 192GB of VRAM, the B200 can usually run 70B-class models with headroom, and may handle much larger models in 4-bit quantized form depending on runtime overhead, KV cache, and context length.
The B200 has 192GB of VRAM. Multi-GPU setups increase total memory, but that memory is not automatically pooled across GPUs.
The B200 has 8,000 GB/s of memory bandwidth. Higher bandwidth helps with faster data transfer between GPU memory and compute cores.
The B200 supports 9 precision formats. Training: BF16, FP16, TF32, FP32. Inference: FP4, FP6, FP8, INT8. Scientific: FP64.
Yes. The B200 supports NVLink with 1800 GB/s of bidirectional bandwidth. This helps accelerate multi-GPU communication.
B200 pricing currently ranges from $3.35/hr to $16.11/hr per GPU, depending on the provider, instance type, and billing model.
At 720 hours per month, one B200 can cost between $2,412.00 to $11,599.20 per month, depending on the provider. Reserved and spot pricing can lower that further.
The B200 is available from 28 cloud providers, including Verda, Lyceum, UpCloud. Pricing and availability vary by region and billing model.
Yes. We currently track 116 B200 listings across 28 cloud providers:
| Billing type | Listings | Avg $/GPU/hr |
|---|---|---|
| On-demand | 42 | $7.23/hr |
| Reserved | 43 | $5.85/hr |
| Spot | 22 | $4.01/hr |
| Custom contract | 9 | $3.49/hr |
| Form Factor | 8x NVIDIA B200 SXM |
| FP4 Tensor Core¹ | 144 PFLOPS |
| FP8/FP6 Tensor Core¹ | 72 PFLOPS |
| INT8 Tensor Core¹ | 72 POPS |
| FP16/BF16 Tensor Core¹ | 36 PFLOPS |
| TF32 Tensor Core¹ | 18 PFLOPS |
| FP32 | 640 TFLOPS |
| FP64 | 320 TFLOPS |
| FP64 Tensor Core | 320 TFLOPS |
| Memory | Up to 1.5TB |
| NVLink | Fifth generation |
| NVIDIA NVSwitch™ | Fourth generation |
| NVSwitch GPU-to-GPU Bandwidth | 1.8TB/s |
| Total Aggregate Bandwidth | 14.4TB/s |
¹ With sparsity.
Source: official Nvidia B200 datasheet.
Last updated