Vast.ai
Our sponsor- On-Demand
- from $4.00
- Reserved
- from $3.28
- Spot
- from $0.84
Extends the H100 with doubled memory for large-model inference and training.
Weekly median price per GPU per hour · Get the data
By provider shows one card per company, placed where its best offer ranked. By configuration lists every offer.
What you can rent comes first: in stock, then waitlist, then not reported, then out of stock. Priced offers rank ahead of quote-only, and on-demand ahead of other billing types.
Within a group, five factors set the order:
We match the provider, its country, GPU model, form factor, billing type, availability and instance name. Matching is partial and case-insensitive. Everyday words work too, so "interruptible" finds spot and "sold out" finds out of stock.
No providers match .
Heads up: A provider's own page may quote a different figure, for example on tax, region or a promotion. A monthly-only plan shows a derived hourly rate. Verify before provisioning. More on how we price.
One H200 has 141 GB of VRAM. In practice, that's enough memory for roughly 220B parameters at 4-bit or 58B at 16-bit, assuming a 32K context. Below are some open-weight LLMs, with the estimated memory and GPUs each one needs.
| Model | Memory (INT4 / FP4) | H200s needed | Cost /hr | Cost /mo |
|---|---|---|---|---|
|
|
18 GB
|
1
|
$4.44
|
$3,197
|
|
|
22 GB
|
1
|
$4.44
|
$3,197
|
|
|
67 GB
|
1
|
$4.44
|
$3,197
|
|
|
178 GB
|
2
|
$8.88
|
$6,394
|
|
|
239 GB
|
2
|
$8.88
|
$6,394
|
|
|
306 GB
|
3
|
$13.32
|
$9,590
|
|
|
1,544 GB
|
13
|
–
|
–
|
Estimates based on the median on-demand rate. Memory is weights plus FP8 KV cache at 32K context per request (FP16/BF16 in the 16-bit column). GPU counts assume 90% of advertised VRAM is usable. No guarantee of runtime support, usable performance, or that a matching quantized build exists. Pricing methodology.
Nvidia H200 · Per GPU
| Compute · dense | |
|---|---|
| FP8 | 1,979 TFLOPS 3,958 with sparsity |
| FP16 / BF16 | 989.5 TFLOPS 1,979 with sparsity |
| INT8 | 1,979 TOPS 3,958 with sparsity |
| FP32 | 67 TFLOPS |
| FP64 | 34 TFLOPS |
| Precision support | FP8FP16BF16TF32FP32FP64INT8 |
| Memory | |
|---|---|
| Capacity | 141 GB HBM3e |
| Bandwidth | 4,800 GB/s |
| ECC | Yes |
| Silicon | |
|---|---|
| Architecture | Hopper |
| Process | TSMC 4N |
| Transistors | 80 billion |
| Shader cores | 16,896 CUDA cores |
| Matrix cores | 528 Tensor cores |
| Compute units | 132 SMs |
| Fabric and host | |
|---|---|
| GPU interconnect | NVLink 900 GB/s |
| Host interface | PCIe 5.0 x16 |
| Power | |
|---|---|
| Board power | 700 W |
| Platform | |
|---|---|
| Partitioning | MIG, up to 7 instances |
Source: official Nvidia H200 datasheet.
As of September 25, 2026, the median on-demand price is $4.44 per GPU per hour across 38 providers with a priced on-demand config.
| Billing type | Configs | Median /GPU/hr | Cheapest |
|---|---|---|---|
On-demand | 129 | $4.44 | $2.09 |
Reserved | 119 | $3.84 | $2.25 (1 mo) |
Spot | 40 | $2.62 | $0.84 (in stock, Vast.ai) |
8 providers also quote the H200 on a custom contract, priced per deal.
At 720 hours per month, one H200 costs an estimated $3,197 at the median on-demand price. Cheapest verified in stock: $2,160 per month on-demand, $2,362 reserved, $605 spot.
GetDeploying currently tracks H200 configs from 51 providers. The cheapest verified in-stock on-demand configs come from Lium, GPU.ai, Runpod and Massed Compute. See the full price comparison above for every provider and config.
As of September 25, 2026, the median on-demand price has risen about 3% over the past 90 days to $4.44 per GPU per hour, and about 25% over the past 12 months.
One H200 runs models up to roughly 220B parameters at 4-bit quantization or 58B at 16-bit, assuming a 32K context. Larger models run across multiple GPUs: the model table above shows the estimated memory and GPU count for popular open-weight LLMs.
141GB HBM3e doubles the memory of the H100 while maintaining NVLink 900 GB/s. Fits 70B+ models without sharding. Best for memory-bound workloads like long-context inference and large batch training.
Limited availability and premium pricing. If your model fits in 80GB or the workload isn't memory-bottlenecked, the H100 offers better price-per-GPU-hour.
Last updated