Cloud GPU pricing
Compare 5,482 GPU prices across 85 cloud providers.
- GPU models
- 108
- 5,482 configs tracked
- Providers
- 85
- Updated 4 minutes ago
GPU models
How this list works
Order
By default, models are ranked on three things in order:
- Demand: how many people visited that model's own page here over the last 90 days.
- Coverage: how many providers list the model at all, which breaks ties between models drawing similar traffic.
- Name, for anything still tied.
Models nobody lists sort to the bottom as a block, whatever their demand.
The two prices on each row
Cheapest is the lowest on-demand listing a reader could rent today, with the provider it is at. Sold-out and waitlisted listings are left out.
Median is what a typical provider charges on demand: each provider's own median rate first, then the median across providers, so a provider listing thirty instance sizes counts once rather than thirty times. Spot and committed rates go into neither column, which is why a model's own page can show something cheaper still.
Filtering and sorting
The memory chips are floors: 80 GB+ keeps every model at or above 80 GB. The class chips separate datacenter parts, workstation cards and consumer GeForce hardware.
Sorting works on the figures the rows show. A model with nothing to show for the field you sorted on goes to the bottom, not the top.
Search
We match the model name and its memory. Matching is partial and case-insensitive.
Type more than one term and every one has to match, in any order, so
nvidia h100 and h100,nvidia are the same search.
Provider names are not searched here. Open a model to see who lists it, or find a provider on the providers page.
Transparency and funding
- Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.
- Prices: Shown in USD, converted at daily reference rates where a provider publishes in another currency. A month means 720 hours. Pricing methodology.
| GPU | Memory · Arch | Cheapest /GPU/hr | Median | Providers |
|---|---|---|---|---|
|
|
80 GB HBM3 Hopper | $1.73 /GPU/hr Vast.ai | $3.39 /GPU/hr |
|
|
|
288 GB HBM3e Blackwell Ultra | $4.99 /GPU/hr Together AI | $8.25 /GPU/hr |
|
|
|
180 GB HBM3e Blackwell | $3.75 /GPU/hr Packet·ai | $6.64 /GPU/hr |
|
|
|
96 GB GDDR7 Blackwell | $0.65 /GPU/hr Hostinger | $2.14 /GPU/hr |
|
|
|
141 GB HBM3e Hopper | $2.09 /GPU/hr Beam | $4.40 /GPU/hr |
|
|
|
32 GB GDDR7 Blackwell | $0.35 /GPU/hr HyperAI | $0.66 /GPU/hr |
|
|
|
80 GB HBM2e Ampere | $0.40 /GPU/hr Vast.ai | $1.79 /GPU/hr |
|
|
|
186 GB HBM3e Grace Blackwell | $10.50 /GPU/hr CoreWeave | $16.00 /GPU/hr |
|
|
|
24 GB GDDR6X Ada Lovelace | $0.33 /GPU/hr Salad | $0.48 /GPU/hr |
|
|
|
288 GB HBM3e CDNA 4 | $5.99 /GPU/hr Daytona | $8.32 /GPU/hr |
|
|
|
192 GB HBM3 CDNA 3 | $2.59 /GPU/hr DigitalOcean | $4.41 /GPU/hr |
|
|
|
288 GB HBM3e Grace Blackwell Ultra | $18.00 /GPU/hr Oracle Cloud | $18.00 /GPU/hr |
|
|
|
48 GB GDDR6 Ada Lovelace | $0.55 /GPU/hr Novita | $1.50 /GPU/hr |
|
|
|
24 GB GDDR6X Ampere | $0.10 /GPU/hr HyperAI | $0.18 /GPU/hr |
|
|
|
48 GB GDDR6 Ada Lovelace | $0.58 /GPU/hr Vast.ai | $0.96 /GPU/hr |
|
|
|
24 GB GDDR6 Ada Lovelace | $0.32 /GPU/hr Vast.ai | $0.88 /GPU/hr |
|
|
|
48 GB GDDR6 Ampere | $0.35 /GPU/hr Thunder Compute | $0.57 /GPU/hr |
|
|
|
32 GB HBM2 Volta | $0.09 /GPU/hr Vast.ai | $0.99 /GPU/hr |
|
|
|
256 GB HBM3e CDNA 3 | $2.25 /GPU/hr Bentaus | $3.06 /GPU/hr |
|
|
|
48 GB GDDR7 Blackwell | $0.60 /GPU/hr Vast.ai | $0.77 /GPU/hr |
|
No GPUs matching your search.
Showing 20 of 108 GPU models
Heads up: These are medians across providers, not a rate you can book: the cheapest listing sits below them, and a provider's own page may quote a different figure, for example on tax, region or a promotion. Verify before provisioning. More on how we price.
Common questions
What is the cheapest cloud GPU?
The lowest listing we currently track is the Nvidia GTX 1050 Ti at Salad, $0.020 per GPU per hour on-demand. For the cheapest GPU by VRAM requirement, class and current availability, see our cheapest cloud GPUs guide.
Why is the cheapest price so far below the median?
The cheapest column is the lowest on-demand listing for that GPU: often a marketplace host, a small configuration or a promotional rate. The median is what a typical provider charges on demand: each provider's own median first, then the median across providers, so a provider listing thirty instance sizes counts once. Budget on the median, then open the GPU's page to see who sells below it and on what terms.
How often are these GPU prices updated?
We track 5,482 GPU configurations from 85 providers. Listings from providers we scrape are re-checked as often as every 15 minutes; the rest are reviewed manually. The Providers figure at the top of this page says when we last refreshed a listing, which is not the same as a price having moved.
What does $/GPU/hr mean?
Every price on this page is normalized to one GPU for one hour, in US dollars, so instances of different sizes can be compared directly. An 8xH100 instance at $24.00 /hr is listed as $3.00 per GPU per hour. Prices in other currencies are converted at current ECB rates.
What is the difference between on-demand, spot, and reserved GPU pricing?
On-demand is the pay-as-you-go rate, billed by the second or hour with no commitment.
Spot (or preemptible, or interruptible) uses spare capacity at a discount, and the provider can reclaim the instance with little notice. Suited to checkpointed training and batch jobs, not to serving traffic.
Reserved trades a commitment of one month to several years for a lower hourly rate. Larger clusters are usually quoted as a custom contract rather than a list price.
Which GPU is best for AI training vs inference?
It depends on the workload, but the specs that decide it are the same:
- Memory capacity and bandwidth: more VRAM allows larger models and batch sizes
- Tensor cores: specialized units for the matrix math AI models run on
- FLOPS: raw compute throughput
- Interconnect: SXM and NVLink move data between GPUs faster than PCIe, which matters once a job spans several GPUs
For a deeper guide, see Tim Dettmers' GPU recommendations.
For training, VRAM and interconnect dominate. H100 and H200 are the usual choices, B200 and B300 the current Blackwell generation, and GB200 and GB300 the rack-scale systems above them. The A100 is still widely listed and costs less per hour.
For inference, a smaller GPU whose VRAM fits the model is often far more cost-effective: an RTX PRO 6000 Blackwell at 96 GB, an L40S at 48 GB, or an RTX 5090 or L4 for smaller models.
How much VRAM do I need?
As a rule of thumb for inference, add three terms: the weights, the KV cache and runtime overhead.
- Weights: about 2 GB per billion parameters at 16-bit, 1 GB at 8-bit and 0.55 GB in 4-bit quantized form
- KV cache: set by the model's layer count, KV heads and head size, and it grows with context length and with concurrent requests. Across open-weight models the rate runs from about 35 KB per token to over 300 KB, so a 32K context costs roughly 1 to 10 GB depending on the architecture
- Runtime overhead: about 2 GB for the CUDA context, activations and sampling buffers
Then leave headroom. Our fit tables treat a model as fitting when it lands within 90% of the card's advertised VRAM, because vLLM and SGLang reserve the rest by default and a model sized at 95% of a card does not start.
Two caveats: training needs several times more than inference, since optimizer states and gradients are held alongside the weights. And multi-GPU instances add total memory, but that memory is not automatically pooled: the model has to be sharded across GPUs, and the interconnect then becomes the limit.
Nvidia vs AMD: which is better for AI?
Nvidia's advantage is CUDA. It is the mature, default software stack, and nearly every framework and kernel targets it first.
AMD's advantage is often price and raw hardware specs, particularly memory capacity on the MI300X, MI325X and MI355X. Its ROCm stack has improved considerably but still lags in coverage of tooling and less common operations.
What is the difference between CUDA cores and tensor cores?
Both are processing units on Nvidia GPUs, with different jobs. CUDA cores handle general-purpose parallel math, such as rendering and simulation. Tensor cores are specialized for the matrix multiply-accumulate operations that dominate neural network training and inference, at reduced precision.
The closest AMD equivalents are stream processors and matrix cores, though they are not comparable one to one.
Can I run AI workloads on a CPU instead?
Yes. PyTorch and other frameworks run on CPUs, and Apple Silicon in particular does reasonably well on small models thanks to unified memory and its neural engine.
The gap is throughput. A GPU applies the same operation across thousands of cores at once, with far more memory bandwidth, so it is typically an order of magnitude or more faster on the matrix math these models are made of. CPUs remain viable for small models, low request volumes, and experimentation.
Open a GPU model to compare providers, or track the market on the GPU price trends.
Back to top ↑