Cloud GPU Pricing

Compare 5,105 GPU prices across 78 cloud providers.

Last update 8 minutes ago

GPU models

107 GPU models, priced across 5,105 configs from 78 providers.

How this list works

Order

By default, models are ranked on three things in order:

  1. Demand: how many people visited that model's own page here over the last 90 days.
  2. Coverage: how many providers list the model at all, which breaks ties between models drawing similar traffic.
  3. Name, for anything still tied.

Models nobody lists sort to the bottom as a block, whatever their demand.

The price on each card

The headline figure is a median, not a floor: each provider's own median rate first, then the median across providers, so a provider listing thirty instance sizes counts once rather than thirty times.

Only on-demand rates go into the median. The range beside it covers every billing type, including spot and committed terms, which is why a model's own page often shows something cheaper.

Filtering and sorting

The class filter separates datacenter parts, workstation cards and consumer GeForce hardware.

Sorting works on the figures the cards show. A model with nothing to show for the field you sorted on goes to the bottom, not the top.

Search

We match the model name and its memory. Matching is partial and case-insensitive. Type more than one term and every one has to match, in any order, so nvidia h100 and h100,nvidia are the same search.

Provider names are not searched here. The provider list further down the page is the box for that.

Transparency and funding

  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.
  • Prices: Shown in USD, converted at daily reference rates where a provider publishes in another currency. A month means 720 hours. How we estimate costs.
Nvidia

Nvidia H100

80GB HBM3 Q3 2022

The standard data center GPU for large-scale AI training and inference.

Median $3.37 / GPU / hr Lium Vast.ai Gcore HyperAI Sesterce +48 providers
Nvidia

Nvidia B300

Up to 288GB HBM3e Q1 2025

Blackwell Ultra GPU for training and serving the largest foundation models.

Median $7.89 / GPU / hr Runpod Hyperstack Lium Enverge Verda +26 providers
Nvidia

Nvidia B200

180GB HBM3e Q1 2024

High-end Blackwell GPU for large-scale AI training and inference.

Median $6.00 / GPU / hr Packet·ai Together AI Beam UpCloud Vast.ai +28 providers
Nvidia

Nvidia RTX PRO 6000

96GB GDDR7 Q1 2025

Blackwell server GPU for AI inference and visualization.

Median $2.20 / GPU / hr Packet·ai HyperAI Vast.ai GPUhub Beam +40 providers
Nvidia

Nvidia H200

141GB HBM3e Q4 2023

Extends the H100 with doubled memory for large-model inference and training.

Median $4.29 / GPU / hr Beam Together AI Lium Koyeb GPU.ai +39 providers
Nvidia

Nvidia RTX 5090

32GB GDDR7 Q1 2025

Top Blackwell consumer GPU for local AI development and rendering.

Median $0.63 / GPU / hr HyperAI Vast.ai GPU.ai Salad GPUhub +14 providers
Nvidia

Nvidia A100

80GB HBM2e Q2 2020

Previous-gen data center workhorse for AI training and inference.

Median $1.73 / GPU / hr Vast.ai GPU.ai Runpod Civo Thunder Compute +38 providers
Nvidia

Nvidia GB200

186GB HBM3e Q1 2024

Grace CPU + Blackwell GPU superchip with unified memory architecture.

Median $16.00 / GPU / hr CoreWeave Oracle Cloud Azure AWS CUDO +3 providers
Nvidia

Nvidia RTX 4090

24GB GDDR6X Q4 2022

Top Ada Lovelace consumer GPU for local AI research and rendering.

Median $0.39 / GPU / hr Vast.ai Salad Lium GPU.ai Novita +12 providers
AMD

AMD MI300X

192GB HBM3 Q4 2023

AMD's flagship data center GPU for large-model inference.

Median $2.99 / GPU / hr Runpod DigitalOcean Cyfuture AI Hot Aisle Crusoe +5 providers

No GPUs matching your search.

Cloud GPU providers

78 providers offering 99 GPU models across 5,105 configs.

How this list works

Order

Providers are ranked by three factors:

  1. GPU demand: The GPUs readers look up most, like the H100, count for more. A short catalog of in-demand GPUs beats a long catalog of rare ones, and every model still counts for something.
  2. Availability: Confirmed stock ranks a little higher, sold-out a little lower. Unconfirmed stock is neutral, never penalized.
  3. Location: We blend a provider's HQ country and where their datacenters and capacity sit, so the nearest to you rank higher. We estimate your location from your IP address.

Providers that tie on all three are sorted alphabetically.

Sorting and searching

The sort control above the list reorders it by cheapest GPU, by how many GPU models a provider lists, or by name.

Each card shows the price of the most sought-after GPU that provider rents. Sorting by cheapest GPU switches every card to that provider's lowest price instead.

Search

We match the provider, their country, the GPU models they list and the billing types they sell. Matching is partial and case-insensitive.

Transparency and funding

  • Ads and sponsors: Paid placements sit at the top of the list and are always labeled as sponsored content. Sponsorship never influences the organic ranking itself.
  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.
  • Prices: Shown in USD, converted at daily reference rates where a provider publishes in another currency. A month means 720 hours. How we estimate costs.

No providers matching your search.

Common questions

What is the cheapest cloud GPU?

The lowest listing we currently track is the Nvidia GTX 1650 at Salad, $0.020 per GPU per hour on-demand. For the cheapest GPU by VRAM requirement, class and current availability, see our cheapest cloud GPUs guide.

How often are these GPU prices updated?

We track 5,105 GPU configurations from 78 providers. Listings from providers we scrape are re-checked as often as every 15 minutes; the rest are reviewed manually. The stamp at the top of this page shows when we last refreshed a listing, which is not the same as a price having moved.

What does $ / GPU / hr mean?

Every price on this page is normalized to one GPU for one hour, in US dollars, so instances of different sizes can be compared directly. An 8xH100 instance at $24.00/hr is listed as $3.00 per GPU per hour. Prices in other currencies are converted at current ECB rates.

What is the difference between on-demand, spot, and reserved GPU pricing?

On-demand is the pay-as-you-go rate, billed by the second or hour with no commitment.

Spot (or preemptible, or interruptible) uses spare capacity at a discount, and the provider can reclaim the instance with little notice. Suited to checkpointed training and batch jobs, not to serving traffic.

Reserved trades a commitment of one month to several years for a lower hourly rate. Larger clusters are usually quoted as a custom contract rather than a list price.

Which GPU is best for AI training vs inference?

It depends on the workload, but the specs that decide it are the same:

  • Memory capacity and bandwidth: more VRAM allows larger models and batch sizes
  • Tensor cores: specialized units for the matrix math AI models run on
  • FLOPS: raw compute throughput
  • Interconnect: SXM and NVLink move data between GPUs faster than PCIe, which matters once a job spans several GPUs

For a deeper guide, see Tim Dettmers' GPU recommendations.

For training, VRAM and interconnect dominate. H100 and H200 are the usual choices, B200 and B300 the current Blackwell generation, and GB200 and GB300 the rack-scale systems above them. The A100 is still widely listed and costs less per hour.

For inference, a smaller GPU whose VRAM fits the model is often far more cost-effective: an RTX PRO 6000 Blackwell at 96 GB, an L40S at 48 GB, or an RTX 5090 or L4 for smaller models.

How much VRAM do I need?

As a rule of thumb for inference, add three terms: the weights, the KV cache and runtime overhead.

  • Weights: about 2 GB per billion parameters at 16-bit, 1 GB at 8-bit and 0.55 GB in 4-bit quantized form
  • KV cache: set by the model's layer count, KV heads and head size, and it grows with context length and with concurrent requests. Across open-weight models the rate runs from about 35 KB per token to over 300 KB, so a 32K context costs roughly 1 to 10 GB depending on the architecture
  • Runtime overhead: about 2 GB for the CUDA context, activations and sampling buffers

Then leave headroom. Our fit tables treat a model as fitting when it lands within 90% of the card's advertised VRAM, because vLLM and SGLang reserve the rest by default and a model sized at 95% of a card does not start.

Two caveats: training needs several times more than inference, since optimizer states and gradients are held alongside the weights. And multi-GPU instances add total memory, but that memory is not automatically pooled: the model has to be sharded across GPUs, and the interconnect then becomes the limit.

Nvidia vs AMD: which is better for AI?

Nvidia's advantage is CUDA. It is the mature, default software stack, and nearly every framework and kernel targets it first.

AMD's advantage is often price and raw hardware specs, particularly memory capacity on the MI300X, MI325X and MI355X. Its ROCm stack has improved considerably but still lags in coverage of tooling and less common operations.

What is the difference between CUDA cores and tensor cores?

Both are processing units on Nvidia GPUs, with different jobs. CUDA cores handle general-purpose parallel math, such as rendering and simulation. Tensor cores are specialized for the matrix multiply-accumulate operations that dominate neural network training and inference, at reduced precision.

The closest AMD equivalents are stream processors and matrix cores, though they are not comparable one to one.

Can I run AI workloads on a CPU instead?

Yes. PyTorch and other frameworks run on CPUs, and Apple Silicon in particular does reasonably well on small models thanks to unified memory and its neural engine.

The gap is throughput. A GPU applies the same operation across thousands of cores at once, with far more memory bandwidth, so it is typically an order of magnitude or more faster on the matrix math these models are made of. CPUs remain viable for small models, low request volumes, and experimentation.

Open a GPU model to compare providers, or track the market on the GPU rental price index.

Back to top ↑