Runpod
Our sponsor- On-Demand
- from $7.89
- Reserved
- on request
Blackwell Ultra GPU for training and serving the largest foundation models.
Weekly median price per GPU per hour · Get the data
By provider shows one card per company, placed where its best offer ranked. By configuration lists every offer.
What you can rent comes first: in stock, then waitlist, then not reported, then out of stock. Priced offers rank ahead of quote-only, and on-demand ahead of other billing types.
Within a group, five factors set the order:
We match the provider, its country, GPU model, form factor, billing type, availability and instance name. Matching is partial and case-insensitive. Everyday words work too, so "interruptible" finds spot and "sold out" finds out of stock.
No providers match .
Heads up: A provider's own page may quote a different figure, for example on tax, region or a promotion. A monthly-only plan shows a derived hourly rate. Verify before provisioning. More on how we price.
One B300 has 288 GB of VRAM. In practice, that's enough memory for roughly 460B parameters at 4-bit or 125B at 16-bit, assuming a 32K context. Below are some open-weight LLMs, with the estimated memory and GPUs each one needs.
| Model | Memory (INT4 / FP4) | B300s needed | Cost /hr | Cost /mo |
|---|---|---|---|---|
|
|
18 GB
|
1
|
$7.99
|
$5,753
|
|
|
22 GB
|
1
|
$7.99
|
$5,753
|
|
|
67 GB
|
1
|
$7.99
|
$5,753
|
|
|
178 GB
|
1
|
$7.99
|
$5,753
|
|
|
239 GB
|
1
|
$7.99
|
$5,753
|
|
|
306 GB
|
2
|
$15.98
|
$11,506
|
|
|
1,544 GB
|
8
|
$63.92
|
$46,022
|
Estimates based on the median on-demand rate. Memory is weights plus FP8 KV cache at 32K context per request (FP16/BF16 in the 16-bit column). GPU counts assume 90% of advertised VRAM is usable. No guarantee of runtime support, usable performance, or that a matching quantized build exists. Pricing methodology.
Nvidia B300 · Per GPU
| Compute · dense | |
|---|---|
| FP4 | 14,000 TFLOPS |
| FP8 | 4,500 TFLOPS 9,000 with sparsity |
| FP16 / BF16 | 2,250 TFLOPS 4,500 with sparsity |
| INT8 | 153.5 TOPS 307 with sparsity |
| FP32 | 75 TFLOPS |
| FP64 | 1.2 TFLOPS |
| Precision support | FP4FP6FP8FP16BF16TF32FP32FP64INT8 |
| Memory | |
|---|---|
| Capacity | 288 GB HBM3e |
| Bandwidth | 8,000 GB/s |
| Bus width | 8,192-bit |
| ECC | Yes |
| Silicon | |
|---|---|
| Architecture | Blackwell Ultra |
| Process | TSMC 4NP |
| Transistors | 208 billion |
| Fabric and host | |
|---|---|
| GPU interconnect | NVLink 1,800 GB/s |
| Host interface | PCIe 6.0 x16 |
| Power | |
|---|---|
| Board power | 1,100 W |
| Cooling | Passive |
| Platform | |
|---|---|
| Partitioning | MIG, up to 7 instances |
Source: official Nvidia B300 datasheet.
As of September 28, 2026, the median on-demand price is $7.99 per GPU per hour across 17 providers with a priced on-demand config.
| Billing type | Configs | Median /GPU/hr | Cheapest |
|---|---|---|---|
On-demand | 44 | $7.99 | $4.99 |
Reserved | 66 | $6.50 | $4.25 |
Spot | 20 | $5.26 | $4.08 (in stock, Daytona) |
7 providers also quote the B300 on a custom contract, priced per deal.
The lowest listed on-demand price for the B300 is $4.99 per GPU per hour from Together AI, though we haven't verified current stock. Cheapest verified in stock: on-demand $7.10 and spot $4.08, both from Daytona.
At 720 hours per month, one B300 costs an estimated $5,753 at the median on-demand price. Cheapest verified in stock: $5,112 per month on-demand, $4,565 reserved, $2,938 spot.
GetDeploying currently tracks B300 configs from 34 providers. The cheapest verified in-stock on-demand configs come from Daytona, Runpod, Lyceum and Verda. See the full price comparison above for every provider and config.
As of September 28, 2026, the median on-demand price has been flat over the past 90 days, at about $7.99 per GPU per hour.
One B300 runs models up to roughly 460B parameters at 4-bit quantization or 125B at 16-bit, assuming a 32K context. Larger models run across multiple GPUs: the model table above shows the estimated memory and GPU count for popular open-weight LLMs.
288GB of HBM3e and 1,800 GB/s of NVLink, a capacity matched in the Blackwell line only by the GB300. Holds 200B-parameter models on a single GPU without quantization.
Last updated