Cloud GPU Pricing

Compare 4,241 GPU prices across 74 cloud providers.

Updated 15 minutes ago

GPU models

102 GPU models, priced across 4,241 configs from 74 providers.

How this list works

Order

By default, models are ranked on three things in order:

  1. Demand: how many people visited that model's own page here over the last 90 days.
  2. Coverage: how many providers list the model at all, which breaks ties between models drawing similar traffic.
  3. Name, for anything still tied.

Models nobody lists sort to the bottom as a block, whatever their demand.

The price on each card

The headline figure is a median, not a floor: each provider's own median rate first, then the median across providers, so a provider listing thirty instance sizes counts once rather than thirty times.

Only on-demand rates go into the median. The range beside it covers every billing type, including spot and committed terms, which is why a model's own page often shows something cheaper.

Filtering and sorting

The class filter separates datacenter parts, workstation cards and consumer GeForce hardware.

Sorting works on the figures the cards show. A model with nothing to show for the field you sorted on goes to the bottom, not the top.

Search

We match the model name and its memory. Matching is partial and case-insensitive. Type more than one term and every one has to match, in any order, so nvidia h100 and h100,nvidia are the same search.

Provider names are not searched here. The provider list further down the page is the box for that.

Transparency and funding

  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.
  • Prices: Shown in USD, converted at daily reference rates where a provider publishes in another currency. A month means 720 hours. How we estimate costs.
Nvidia

Nvidia H100

80GB HBM3 Q3 2022

The standard data center GPU for large-scale AI training and inference.

Median $3.38 / GPU / hr Vast.ai GPU.ai HyperAI UpCloud Oblivus +47 providers
Nvidia

Nvidia B300

Up to 288GB HBM3e Q1 2025

Highest VRAM in the Blackwell lineup for training the largest foundation models.

Median $7.87 / GPU / hr Runpod Hyperstack Enverge Verda Nebius +19 providers
Nvidia

Nvidia B200

Up to 192GB HBM3e Q1 2024

High-end Blackwell GPU for large-scale AI training and inference.

Median $6.49 / GPU / hr Packet·ai UpCloud Koyeb Vast.ai GPU.ai +22 providers
Nvidia

Nvidia H200

141GB HBM3e Q4 2023

Extends the H100 with doubled memory for large-model inference and training.

Median $4.48 / GPU / hr Koyeb Cerebrium Geodd Civo Runpod +31 providers
Nvidia

Nvidia RTX 5090

32GB GDDR7 Q1 2025

Top Blackwell consumer GPU for local AI development and rendering.

Median $0.56 / GPU / hr Salad HyperAI GPU.ai Vast.ai GPUhub +9 providers
Nvidia

Nvidia RTX PRO 6000

96GB GDDR7 Q1 2025

Largest VRAM desktop workstation GPU for local AI and visualization.

Median $2.20 / GPU / hr Packet·ai HyperAI GPUhub Vast.ai Nova Cloud +29 providers
Nvidia

Nvidia A100

40GB / 80GB HBM2e Q2 2020

Previous-gen data center workhorse for AI training and inference.

Median $1.76 / GPU / hr Vast.ai GPU.ai Runpod Civo Thunder Compute +36 providers
Nvidia

Nvidia RTX 4090

24GB GDDR6X Q4 2022

Top Ada Lovelace consumer GPU for local AI research and rendering.

Median $0.44 / GPU / hr Salad Vast.ai Novita GPU.ai Runpod +12 providers
Nvidia

Nvidia GB200

Up to 13.4TB HBM3e Q1 2024

Grace CPU + Blackwell GPU superchip with unified memory architecture.

Median $16.00 / GPU / hr CoreWeave Oracle Cloud Azure AWS Together +3 providers
AMD

AMD MI300X

192GB HBM3 Q4 2023

AMD's flagship data center GPU for large-model inference.

Median $2.90 / GPU / hr Runpod DigitalOcean Cyfuture AI Hot Aisle Crusoe +5 providers

No GPUs matching your search.

Cloud GPU providers

74 providers offering 94 GPU models across 4,241 configs.

How this list works

Order

Providers are ranked by three factors:

  1. GPU demand: Widely offered GPUs like the H100 count for more than models only one or two providers list. A short catalog of in-demand GPUs ranks above a long catalog of rare ones.
  2. Availability: Confirmed stock ranks a little higher, and sold-out catalogs a little lower. Providers whose stock we cannot confirm are treated as neutral, never penalized.
  3. Location: A datacenter in your country ranks highest, then a datacenter on your continent, then a provider headquartered on it. We estimate your location from your IP address. Location is a small adjustment, so it never lifts a thin catalog above a deep one.

Providers that tie on all three are sorted alphabetically.

Sorting and searching

The sort control above the list reorders it by cheapest GPU, by how many GPU models a provider lists, or by name.

Each card shows the price of the GPU the widest part of the market rents. Sorting by cheapest GPU switches every card to that provider's lowest price instead.

Search

We match the provider, their country, the GPU models they list and the billing types they sell. Matching is partial and case-insensitive.

Transparency and funding

  • Ads and sponsors: Paid placements sit at the top of the list and are always labeled as sponsored content. Sponsorship never influences the organic ranking itself.
  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.
  • Prices: Shown in USD, converted at daily reference rates where a provider publishes in another currency. A month means 720 hours. How we estimate costs.

No providers matching your search.

Common questions

What is the cheapest cloud GPU?

The cheapest listing we currently track is the Nvidia GTX 1660 Super at Vast.ai, at $0.02 per GPU per hour. The cheapest option for a given workload is usually a different one: the smallest GPU whose VRAM fits your model, in a region you can actually get capacity in.

How often are these GPU prices updated?

We track 4,241 GPU configurations from 74 providers. Listings from providers we scrape are re-checked as often as every 15 minutes; the rest are reviewed manually. The stamp at the top of this page shows when we last refreshed a listing, which is not the same as a price having moved.

What does $ / GPU / hr mean?

Every price on this page is normalized to one GPU for one hour, in US dollars, so instances of different sizes can be compared directly. An 8xH100 instance at $24.00/hr is listed as $3.00 per GPU per hour. Prices in other currencies are converted at current ECB rates.

What is the difference between on-demand, spot, and reserved GPU pricing?

On-demand is the pay-as-you-go rate, billed by the second or hour with no commitment.

Spot (or preemptible, or interruptible) uses spare capacity at a discount, and the provider can reclaim the instance with little notice. Suited to checkpointed training and batch jobs, not to serving traffic.

Reserved trades a commitment of one month to several years for a lower hourly rate. Larger clusters are usually quoted as a custom contract rather than a list price.

Which GPU is best for AI training vs inference?

It depends on the workload, but the specs that decide it are the same:

  • Memory capacity and bandwidth: more VRAM allows larger models and batch sizes
  • Tensor cores: specialized units for the matrix math AI models run on
  • FLOPS: raw compute throughput
  • Interconnect: SXM and NVLink move data between GPUs faster than PCIe, which matters once a job spans several GPUs

For training, VRAM and interconnect dominate: H100, H200, B200 and A100 are the usual choices. For inference, a smaller GPU whose VRAM fits the model is often far more cost-effective, such as an L40S, A10 or RTX 4090.

For a deeper guide, see Tim Dettmers' GPU recommendations.

How much VRAM do I need?

As a rough rule for inference, a model needs about 2 GB of VRAM per billion parameters in FP16, or about 0.5 GB per billion in 4-bit quantized form, plus headroom for the KV cache and runtime overhead, which grows with context length. Training needs several times more, since optimizer states and gradients are held alongside the weights.

Multi-GPU instances add total memory, but that memory is not automatically pooled: the model has to be sharded across GPUs, and the interconnect then becomes the limit.

Nvidia vs AMD: which is better for AI?

Nvidia's advantage is CUDA. It is the mature, default software stack, and nearly every framework and kernel targets it first.

AMD's advantage is often price and raw hardware specs, particularly memory capacity on the MI300X and MI325X. Its ROCm stack has improved considerably but still lags in coverage of tooling and less common operations.

In practice, teams with the engineering capacity to work around gaps can run AMD economically. Everyone else pays a premium for CUDA and saves the time.

What is the difference between CUDA cores and tensor cores?

Both are processing units on Nvidia GPUs, with different jobs. CUDA cores handle general-purpose parallel math, such as rendering and simulation. Tensor cores are specialized for the matrix multiply-accumulate operations that dominate neural network training and inference, at reduced precision.

The closest AMD equivalents are stream processors and matrix cores, though they are not comparable one to one.

Can I run AI workloads on a CPU instead?

Yes. PyTorch and other frameworks run on CPUs, and Apple Silicon in particular does reasonably well on small models thanks to unified memory and its neural engine.

The gap is throughput. A GPU applies the same operation across thousands of cores at once, with far more memory bandwidth, so it is typically an order of magnitude or more faster on the matrix math these models are made of. CPUs remain viable for small models, low request volumes, and experimentation.

Looking for a specific configuration? Open a GPU model to compare every provider listing it, or track the market on the GPU rental price index.

Back to top ↑