Nvidia H100
80GB HBM3 Q3 2022
The standard data center GPU for large-scale AI training and inference.
Compare 4,241 GPU prices across 74 cloud providers.
Updated 15 minutes ago
102 GPU models, priced across 4,241 configs from 74 providers.
By default, models are ranked on three things in order:
Models nobody lists sort to the bottom as a block, whatever their demand.
The headline figure is a median, not a floor: each provider's own median rate first, then the median across providers, so a provider listing thirty instance sizes counts once rather than thirty times.
Only on-demand rates go into the median. The range beside it covers every billing type, including spot and committed terms, which is why a model's own page often shows something cheaper.
The class filter separates datacenter parts, workstation cards and consumer GeForce hardware.
Sorting works on the figures the cards show. A model with nothing to show for the field you sorted on goes to the bottom, not the top.
We match the model name and its memory. Matching is partial and case-insensitive.
Type more than one term and every one has to match, in any order, so
nvidia h100 and h100,nvidia are the same search.
Provider names are not searched here. The provider list further down the page is the box for that.
80GB HBM3 Q3 2022
The standard data center GPU for large-scale AI training and inference.
Up to 288GB HBM3e Q1 2025
Highest VRAM in the Blackwell lineup for training the largest foundation models.
Up to 192GB HBM3e Q1 2024
High-end Blackwell GPU for large-scale AI training and inference.
141GB HBM3e Q4 2023
Extends the H100 with doubled memory for large-model inference and training.
32GB GDDR7 Q1 2025
Top Blackwell consumer GPU for local AI development and rendering.
96GB GDDR7 Q1 2025
Largest VRAM desktop workstation GPU for local AI and visualization.
40GB / 80GB HBM2e Q2 2020
Previous-gen data center workhorse for AI training and inference.
24GB GDDR6X Q4 2022
Top Ada Lovelace consumer GPU for local AI research and rendering.
Up to 13.4TB HBM3e Q1 2024
Grace CPU + Blackwell GPU superchip with unified memory architecture.
192GB HBM3 Q4 2023
AMD's flagship data center GPU for large-model inference.
No GPUs matching your search.
74 providers offering 94 GPU models across 4,241 configs.
Providers are ranked by three factors:
Providers that tie on all three are sorted alphabetically.
The sort control above the list reorders it by cheapest GPU, by how many GPU models a provider lists, or by name.
Each card shows the price of the GPU the widest part of the market rents. Sorting by cheapest GPU switches every card to that provider's lowest price instead.
We match the provider, their country, the GPU models they list and the billing types they sell. Matching is partial and case-insensitive.
USA Founded 2024
On-demand GPUs for startups, researchers, and developers
USA Founded 2018
Cloud GPU marketplace · Real-time pricing · Per-second billing
USA Founded 2022
All-in-one platform to train, fine-tune, and deploy AI applications
France Founded 2018
French cloud provider specializing in GPU compute
USA Founded 2006
Cloud computing platform by Amazon
UAE
Automatically routes each launch to the lowest cost available GPU
USA Founded 2016
Cloud computing platform by Oracle
USA Founded 2026
GPU cloud platform with fast deployment and pay-per-use pricing
France Founded 2020
Deploy apps without complex infrastructure
USA Founded 2025
Hybrid cloud-edge platform for AI and media workloads
USA Founded 2021
GPU cloud for AI, ML, and VFX
USA Founded 2024
GPU cloud and inference platform
USA Founded 2008
Cloud computing platform by Google
USA Founded 2017
Cloud infrastructure for AI and compute workloads
Finland Founded 2018
European cloud provider for GPUs
USA Founded 2005
Affordable VPS and GPU hosting
USA Founded 2008
Cloud computing platform by Microsoft
Singapore Founded 2020
GPU marketplace aggregating capacity from vetted data center partners
USA Founded 2010
GPU cloud provider for AI applications
UK Founded 2020
Global cloud provider specializing in cloud GPUs
USA Founded 2023
Bare metal NVIDIA GPU instances and managed AI inference
Netherlands Founded 2022
AI cloud platform from the Netherlands
India Founded 2010
Cloud provider offering GPU compute and infrastructure services
USA Founded 2014
Cloud provider offering GPUs, bare-metal servers, Kubernetes and more
USA Founded 2012
On-demand GPUs for AI training and inference
USA Founded 2018
Cloud infrastructure for enterprise AI workloads
USA Founded 2023
GPU cloud platform for AI model training and inference
UK Founded 2015
Cloud provider offering Kubernetes, GPU servers, and LLM inference
Germany Founded 2025
EU-sovereign cloud for AI training and inference
USA Founded 2011
Simple, cost-effective cloud hosting services
USA Founded 2014
Train and deploy AI applications
Germany Founded 2003
Low-cost servers available in 9 regions
USA Founded 2021
Easy & affordable GPU cloud marketplace
Finland Founded 2011
European cloud provider with high-performance servers
USA Founded 2021
Serverless platform for running AI models
Germany Founded 2021
Bare metal GPU servers and S3-compatible object storage
Brazil Founded 2001
Bare metal, GPU and VM servers in 25 locations worldwide
USA Founded 2024
GPU cloud and open-model inference platform
India Founded 2001
GPU cloud and AI deployment platform based in India
UAE Founded 2008
Independent VPS and GPU cloud with 10+ global regions
USA Founded 2025
Green energy GPU cloud for AI training and fine-tuning
France Founded 1999
European cloud provider with a wide range of managed services
Hong Kong Founded 2017
GPU containers and notebooks for AI development
USA Founded 2018
Affordable cloud for AI/ML inference
Netherlands Founded 2020
Sustainable European cloud with managed services and GPUs
Sweden Founded 2012
Swedish cloud provider with managed services and GPUs
Switzerland Founded 2025
Bare-metal GPU servers hosted in Switzerland with per-minute billing.
Singapore Founded 2024
Affordable GPU instances with zero egress fees
Switzerland Founded 2011
European cloud provider with a range of managed services
Singapore Founded 2009
Cloud platform with a strong presence in Asia
USA Founded 2003
Developer-focused VPS and cloud provider
USA Founded 2022
US-based cloud hosting provider offering VPS and GPU instances
Poland Founded 2005
Polish cloud with DGX B200 GPUs and 100MW campus
UAE
Bare metal Nvidia B300 and GB300 servers
Netherlands Founded 1997
Cloud provider from the Netherlands with 20+ datacenters
Sweden Founded 2016
Swedish cloud provider with affordable VPS and object storage
Ukraine
VPS and GPU servers in Ukraine, Poland, and Sweden
USA
Global edge cloud with 350+ PoPs across 50+ countries
USA Founded 2023
AMD-powered bare metal compute
No providers matching your search.
The cheapest listing we currently track is the Nvidia GTX 1660 Super at Vast.ai, at $0.02 per GPU per hour. The cheapest option for a given workload is usually a different one: the smallest GPU whose VRAM fits your model, in a region you can actually get capacity in.
We track 4,241 GPU configurations from 74 providers. Listings from providers we scrape are re-checked as often as every 15 minutes; the rest are reviewed manually. The stamp at the top of this page shows when we last refreshed a listing, which is not the same as a price having moved.
Every price on this page is normalized to one GPU for one hour, in US dollars, so instances of different sizes can be compared directly. An 8xH100 instance at $24.00/hr is listed as $3.00 per GPU per hour. Prices in other currencies are converted at current ECB rates.
On-demand is the pay-as-you-go rate, billed by the second or hour with no commitment.
Spot (or preemptible, or interruptible) uses spare capacity at a discount, and the provider can reclaim the instance with little notice. Suited to checkpointed training and batch jobs, not to serving traffic.
Reserved trades a commitment of one month to several years for a lower hourly rate. Larger clusters are usually quoted as a custom contract rather than a list price.
It depends on the workload, but the specs that decide it are the same:
For training, VRAM and interconnect dominate: H100, H200, B200 and A100 are the usual choices. For inference, a smaller GPU whose VRAM fits the model is often far more cost-effective, such as an L40S, A10 or RTX 4090.
For a deeper guide, see Tim Dettmers' GPU recommendations.
As a rough rule for inference, a model needs about 2 GB of VRAM per billion parameters in FP16, or about 0.5 GB per billion in 4-bit quantized form, plus headroom for the KV cache and runtime overhead, which grows with context length. Training needs several times more, since optimizer states and gradients are held alongside the weights.
Multi-GPU instances add total memory, but that memory is not automatically pooled: the model has to be sharded across GPUs, and the interconnect then becomes the limit.
Nvidia's advantage is CUDA. It is the mature, default software stack, and nearly every framework and kernel targets it first.
AMD's advantage is often price and raw hardware specs, particularly memory capacity on the MI300X and MI325X. Its ROCm stack has improved considerably but still lags in coverage of tooling and less common operations.
In practice, teams with the engineering capacity to work around gaps can run AMD economically. Everyone else pays a premium for CUDA and saves the time.
Both are processing units on Nvidia GPUs, with different jobs. CUDA cores handle general-purpose parallel math, such as rendering and simulation. Tensor cores are specialized for the matrix multiply-accumulate operations that dominate neural network training and inference, at reduced precision.
The closest AMD equivalents are stream processors and matrix cores, though they are not comparable one to one.
Yes. PyTorch and other frameworks run on CPUs, and Apple Silicon in particular does reasonably well on small models thanks to unified memory and its neural engine.
The gap is throughput. A GPU applies the same operation across thousands of cores at once, with far more memory bandwidth, so it is typically an order of magnitude or more faster on the matrix math these models are made of. CPUs remain viable for small models, low request volumes, and experimentation.
Looking for a specific configuration? Open a GPU model to compare every provider listing it, or track the market on the GPU rental price index.
Back to top ↑