Cloud GPU pricing

Compare 5,482 GPU prices across 85 cloud providers.

GPU price trend
+10%
Last 12 months · 22 models
Cheapest H100
$1.73 /GPU/hr
Vast.ai · on-demand
GPU models
108
5,482 configs tracked
Providers
85
Updated 4 minutes ago

GPU models

How this list works

Order

By default, models are ranked on three things in order:

  1. Demand: how many people visited that model's own page here over the last 90 days.
  2. Coverage: how many providers list the model at all, which breaks ties between models drawing similar traffic.
  3. Name, for anything still tied.

Models nobody lists sort to the bottom as a block, whatever their demand.

The two prices on each row

Cheapest is the lowest on-demand listing a reader could rent today, with the provider it is at. Sold-out and waitlisted listings are left out.

Median is what a typical provider charges on demand: each provider's own median rate first, then the median across providers, so a provider listing thirty instance sizes counts once rather than thirty times. Spot and committed rates go into neither column, which is why a model's own page can show something cheaper still.

Filtering and sorting

The memory chips are floors: 80 GB+ keeps every model at or above 80 GB. The class chips separate datacenter parts, workstation cards and consumer GeForce hardware.

Sorting works on the figures the rows show. A model with nothing to show for the field you sorted on goes to the bottom, not the top.

Search

We match the model name and its memory. Matching is partial and case-insensitive. Type more than one term and every one has to match, in any order, so nvidia h100 and h100,nvidia are the same search.

Provider names are not searched here. Open a model to see who lists it, or find a provider on the providers page.

Transparency and funding

  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.
  • Prices: Shown in USD, converted at daily reference rates where a provider publishes in another currency. A month means 720 hours. Pricing methodology.
Cloud GPU models with memory, architecture, cheapest on-demand price and median price per GPU per hour, and the providers listing each
GPU Memory · Arch Cheapest /GPU/hr Median Providers
Nvidia Nvidia H100 Released 2022 80 GB HBM3 Hopper $1.73 /GPU/hr Vast.ai $3.39 /GPU/hr Vast.ai Gcore HyperAI Sesterce Beam +52 providers
Nvidia Nvidia B300 Released 2025 288 GB HBM3e Blackwell Ultra $4.99 /GPU/hr Together AI $8.25 /GPU/hr Together AI Runpod Daytona Enverge GPU.ai +28 providers
Nvidia Nvidia B200 Released 2024 180 GB HBM3e Blackwell $3.75 /GPU/hr Packet·ai $6.64 /GPU/hr Packet·ai Together AI Beam Hostinger UpCloud +35 providers
Nvidia Nvidia RTX PRO 6000 Released 2025 96 GB GDDR7 Blackwell $0.65 /GPU/hr Hostinger $2.14 /GPU/hr Hostinger Packet·ai HyperAI Vast.ai Beam +48 providers
Nvidia Nvidia H200 Released 2023 141 GB HBM3e Hopper $2.09 /GPU/hr Beam $4.40 /GPU/hr Beam Lium Together AI Koyeb GPU.ai +46 providers
Nvidia Nvidia RTX 5090 Released 2025 32 GB GDDR7 Blackwell $0.35 /GPU/hr HyperAI $0.66 /GPU/hr HyperAI Vast.ai GPUhub GPU.ai Salad +17 providers
Nvidia Nvidia A100 Released 2020 80 GB HBM2e Ampere $0.40 /GPU/hr Vast.ai $1.79 /GPU/hr Vast.ai GPU.ai Jarvislabs Runpod Thunder Compute +43 providers
Nvidia Nvidia GB200 Released 2024 186 GB HBM3e Grace Blackwell $10.50 /GPU/hr CoreWeave $16.00 /GPU/hr CoreWeave Oracle Cloud Azure AWS Crusoe +5 providers
Nvidia Nvidia RTX 4090 Released 2022 24 GB GDDR6X Ada Lovelace $0.33 /GPU/hr Salad $0.48 /GPU/hr Novita Salad Vast.ai Runpod Lium +16 providers
AMD AMD MI355X Released 2025 288 GB HBM3e CDNA 4 $5.99 /GPU/hr Daytona $8.32 /GPU/hr Daytona Oracle Cloud Vultr DigitalOcean TensorWave +1 provider
AMD AMD MI300X Released 2023 192 GB HBM3 CDNA 3 $2.59 /GPU/hr DigitalOcean $4.41 /GPU/hr Runpod DigitalOcean Cyfuture AI Hot Aisle Crusoe +6 providers
Nvidia Nvidia GB300 Released 2025 288 GB HBM3e Grace Blackwell Ultra $18.00 /GPU/hr Oracle Cloud $18.00 /GPU/hr Verda Oracle Cloud AWS Gcore Nebius +2 providers
Nvidia Nvidia L40S Released 2023 48 GB GDDR6 Ada Lovelace $0.55 /GPU/hr Novita $1.50 /GPU/hr Novita Vast.ai Beam GPU.ai Runpod +34 providers
Nvidia Nvidia RTX 3090 Released 2020 24 GB GDDR6X Ampere $0.10 /GPU/hr HyperAI $0.18 /GPU/hr HyperAI GPU.ai Lium Vast.ai Salad +3 providers
Nvidia Nvidia RTX 6000 Ada Released 2022 48 GB GDDR6 Ada Lovelace $0.58 /GPU/hr Vast.ai $0.96 /GPU/hr Vast.ai GPU.ai Lium Runpod Novita +9 providers
Nvidia Nvidia L4 Released 2023 24 GB GDDR6 Ada Lovelace $0.32 /GPU/hr Vast.ai $0.88 /GPU/hr Vast.ai GPU.ai Hexabyte Jarvislabs Leaseweb +16 providers
Nvidia Nvidia RTX A6000 Released 2020 48 GB GDDR6 Ampere $0.35 /GPU/hr Thunder Compute $0.57 /GPU/hr Runpod Thunder Compute Vast.ai GPU.ai Hyperstack +16 providers
Nvidia Nvidia V100 Released 2017 32 GB HBM2 Volta $0.09 /GPU/hr Vast.ai $0.99 /GPU/hr Vast.ai HyperAI Geodd Runpod Verda +15 providers
AMD AMD MI325X Released 2024 256 GB HBM3e CDNA 3 $2.25 /GPU/hr Bentaus $3.06 /GPU/hr Cyfuture AI DigitalOcean Bentaus Vultr TensorWave 5 providers
Nvidia Nvidia RTX PRO 5000 Released 2025 48 GB GDDR7 Blackwell $0.60 /GPU/hr Vast.ai $0.77 /GPU/hr Vast.ai Runpod Database Mart Contabo 4 providers

No GPUs matching your search.

Showing 20 of 108 GPU models

Heads up: These are medians across providers, not a rate you can book: the cheapest listing sits below them, and a provider's own page may quote a different figure, for example on tax, region or a promotion. Verify before provisioning. More on how we price.

Common questions

What is the cheapest cloud GPU?

The lowest listing we currently track is the Nvidia GTX 1050 Ti at Salad, $0.020 per GPU per hour on-demand. For the cheapest GPU by VRAM requirement, class and current availability, see our cheapest cloud GPUs guide.

Why is the cheapest price so far below the median?

The cheapest column is the lowest on-demand listing for that GPU: often a marketplace host, a small configuration or a promotional rate. The median is what a typical provider charges on demand: each provider's own median first, then the median across providers, so a provider listing thirty instance sizes counts once. Budget on the median, then open the GPU's page to see who sells below it and on what terms.

How often are these GPU prices updated?

We track 5,482 GPU configurations from 85 providers. Listings from providers we scrape are re-checked as often as every 15 minutes; the rest are reviewed manually. The Providers figure at the top of this page says when we last refreshed a listing, which is not the same as a price having moved.

What does $/GPU/hr mean?

Every price on this page is normalized to one GPU for one hour, in US dollars, so instances of different sizes can be compared directly. An 8xH100 instance at $24.00 /hr is listed as $3.00 per GPU per hour. Prices in other currencies are converted at current ECB rates.

What is the difference between on-demand, spot, and reserved GPU pricing?

On-demand is the pay-as-you-go rate, billed by the second or hour with no commitment.

Spot (or preemptible, or interruptible) uses spare capacity at a discount, and the provider can reclaim the instance with little notice. Suited to checkpointed training and batch jobs, not to serving traffic.

Reserved trades a commitment of one month to several years for a lower hourly rate. Larger clusters are usually quoted as a custom contract rather than a list price.

Which GPU is best for AI training vs inference?

It depends on the workload, but the specs that decide it are the same:

  • Memory capacity and bandwidth: more VRAM allows larger models and batch sizes
  • Tensor cores: specialized units for the matrix math AI models run on
  • FLOPS: raw compute throughput
  • Interconnect: SXM and NVLink move data between GPUs faster than PCIe, which matters once a job spans several GPUs

For a deeper guide, see Tim Dettmers' GPU recommendations.

For training, VRAM and interconnect dominate. H100 and H200 are the usual choices, B200 and B300 the current Blackwell generation, and GB200 and GB300 the rack-scale systems above them. The A100 is still widely listed and costs less per hour.

For inference, a smaller GPU whose VRAM fits the model is often far more cost-effective: an RTX PRO 6000 Blackwell at 96 GB, an L40S at 48 GB, or an RTX 5090 or L4 for smaller models.

How much VRAM do I need?

As a rule of thumb for inference, add three terms: the weights, the KV cache and runtime overhead.

  • Weights: about 2 GB per billion parameters at 16-bit, 1 GB at 8-bit and 0.55 GB in 4-bit quantized form
  • KV cache: set by the model's layer count, KV heads and head size, and it grows with context length and with concurrent requests. Across open-weight models the rate runs from about 35 KB per token to over 300 KB, so a 32K context costs roughly 1 to 10 GB depending on the architecture
  • Runtime overhead: about 2 GB for the CUDA context, activations and sampling buffers

Then leave headroom. Our fit tables treat a model as fitting when it lands within 90% of the card's advertised VRAM, because vLLM and SGLang reserve the rest by default and a model sized at 95% of a card does not start.

Two caveats: training needs several times more than inference, since optimizer states and gradients are held alongside the weights. And multi-GPU instances add total memory, but that memory is not automatically pooled: the model has to be sharded across GPUs, and the interconnect then becomes the limit.

Nvidia vs AMD: which is better for AI?

Nvidia's advantage is CUDA. It is the mature, default software stack, and nearly every framework and kernel targets it first.

AMD's advantage is often price and raw hardware specs, particularly memory capacity on the MI300X, MI325X and MI355X. Its ROCm stack has improved considerably but still lags in coverage of tooling and less common operations.

What is the difference between CUDA cores and tensor cores?

Both are processing units on Nvidia GPUs, with different jobs. CUDA cores handle general-purpose parallel math, such as rendering and simulation. Tensor cores are specialized for the matrix multiply-accumulate operations that dominate neural network training and inference, at reduced precision.

The closest AMD equivalents are stream processors and matrix cores, though they are not comparable one to one.

Can I run AI workloads on a CPU instead?

Yes. PyTorch and other frameworks run on CPUs, and Apple Silicon in particular does reasonably well on small models thanks to unified memory and its neural engine.

The gap is throughput. A GPU applies the same operation across thousands of cores at once, with far more memory bandwidth, so it is typically an order of magnitude or more faster on the matrix math these models are made of. CPUs remain viable for small models, low request volumes, and experimentation.

Open a GPU model to compare providers, or track the market on the GPU price trends.

Back to top ↑