Cheapest Cloud GPUs: Updated Daily
Sept. 13, 2026 (updated)
We track 3,841 GPU offerings across 79 providers. This page shows the lowest current on-demand price for each of the 86 GPU models we list, normalized to USD per GPU per hour.
Cheapest GPU by VRAM
How much memory do you need? Each dot is a GPU model, the stepped line shows the cheapest option that meets progressively larger VRAM requirements.
The biggest price jump currently comes above 96 GB: the cheapest option rises 3.2x, from the RTX PRO 6000 to the H200.
At common memory requirements:
| Minimum VRAM | Cheapest card | Class | Price /GPU/hr | Provider | Availability |
|---|---|---|---|---|---|
| 16 GB and up |
RTX 2080 Ti 22GB
|
Consumer | $0.056 |
Vast.ai
|
In stock |
| 24 GB and up |
RTX 3090
|
Consumer | $0.10 |
HyperAI
|
In stock |
| 48 GB and up |
RTX A6000
|
Workstation | $0.35 |
Thunder Compute
|
In stock |
| 80 GB and up |
RTX PRO 6000
|
Workstation | $0.66 |
Packet·ai
|
In stock |
| 141 GB and up |
H200
|
Datacenter | $2.09 |
Beam
|
Not reported |
| 192 GB and up |
MI325X
|
Datacenter | $2.25 |
Bentaus
|
Not reported |
| 256 GB and up |
MI325X
|
Datacenter | $2.25 |
Bentaus
|
Not reported |
Cheapest GPUs for running LLMs
Know the model rather than its VRAM requirement? Here is the cheapest setup we estimate will fit common model sizes at 4-bit and a 32K context.
| Approx. model size | Est. memory at 4-bit | Cheapest setup | Price /hr |
|---|---|---|---|
| ~8B | 14 GB | 1× RTX 2080 Ti 22GB | $0.056 |
| ~30B | 27 GB | 1× V100 | $0.14 |
| ~70B | 49 GB | 2× V100 | $0.28 |
| ~120B | 76 GB | 4× RTX 3090 | $0.40 |
| ~200B | 120 GB | 8× RTX 3090 | $0.80 |
| ~400B | 230 GB | 8× V100 | $1.12 |
Around 30B is currently the last tier where the cheapest setup uses one GPU. Above that, the lowest-cost options use multiple cards.
Models with the same parameter count can need different amounts of memory, because KV-cache requirements vary by architecture.
| Model | Parameters | Est. memory at 4-bit | Cheapest setup | Price /hr |
|---|---|---|---|---|
Qwen3.8-27B
|
27B | 19 GB | 1× RTX 2080 Ti 22GB | $0.056 |
Gemma 4 31B
|
30.7B | 24 GB | 1× V100 | $0.14 |
GPT-OSS-120B
|
117B | 68 GB | 4× RTX 3090 | $0.40 |
GLM-5.3-Flash
|
320B | 178 GB | 8× V100 | $1.12 |
MiniMax-M3
|
428B | 239 GB | 4× RTX PRO 6000 | $3.20 |
DeepSeek V4.1 Flash
|
552B | 306 GB | 4× RTX PRO 6000 | $3.20 |
Kimi K3
|
2.8T | 1,544 GB | 8× B300 | $52.00 |
These are memory-fit estimates, not performance recommendations. Runtime support, quantized builds, multi-GPU interconnect and actual speed vary by model and provider. See how we estimate LLM fit.
Popular GPU prices
The 10 GPU models our readers look up most, comparing today's cheapest listing with the typical weekly market rate.
| GPU | Cheapest today | Typical market rate | Gap |
|---|---|---|---|
H100
|
$1.41 | $3.32 | 58% below |
B300
|
$6.50 | $7.89 | 18% below |
B200
|
$3.75 | $6.01 | 38% below |
RTX PRO 6000
|
$0.66 | $2.20 | 70% below |
H200
|
$2.09 | $4.40 | 53% below |
RTX 5090
|
$0.35 | $0.63 | 44% below |
A100 40GB
|
$0.70 | $1.76 | 60% below |
GB200
|
$10.50 | — | — |
RTX 4090
|
$0.23 | $0.42 | 44% below |
MI300X
|
$2.39 | $2.99 | 20% below |
Of the GPUs in this table, the cheapest RTX PRO 6000 is currently the furthest below its typical weekly market rate, at 70% below.
GPU price is only part of workload cost; storage, egress, CPU and RAM and idle time can change the total.
Heads up: A provider's own page may quote a different figure, for example on tax, region or a promotion. A monthly-only plan shows a derived hourly rate. The market-rate column is a cross-provider weekly median, which nobody sells at. Verify before provisioning. More on how we price.
Browse all GPU models and provider prices →
Frequently asked questions
What is the cheapest cloud GPU right now?
The cheapest cloud GPU is the Nvidia GTX 1650 at $0.020 per GPU per hour from Salad. It has 4 GB of VRAM, so it is the cheapest GPU overall rather than the cheapest option for a larger memory requirement.
What is the cheapest way to rent an H100?
The cheapest on-demand H100 is $1.41 per GPU per hour from Lium (checked Sep 13, 2026). See the H100 page for all providers, availability, spot and reserved prices.
What is the cheapest GPU with at least 16 GB, 24 GB, 48 GB or 80 GB of VRAM?
At least 16 GB: RTX 2080 Ti 22GB at $0.056/hr. At least 24 GB: RTX 3090 at $0.10/hr. At least 48 GB: RTX A6000 at $0.35/hr. At least 80 GB: RTX PRO 6000 at $0.66/hr.
What is the cheapest GPU for a 70B or 120B model?
At 4-bit and a 32K context, a representative 70B model needs about 49 GB and fits cheapest on 2x V100 at $0.28 an hour; a representative 120B model needs about 76 GB and fits cheapest on 4x RTX 3090 at $0.40 an hour. Actual requirements vary by architecture, and fitting is not the same as running well.
Are consumer GPUs cheaper than datacenter GPUs?
Often. At 24 GB the cheapest consumer option is currently the RTX 3090 at $0.10 per GPU per hour, against $0.32 for the L4. Consumer cards generally trade a lower price for fewer datacenter-oriented features and support guarantees.
What is the difference between the cheapest price and the typical market rate?
The cheapest price is the single lowest on-demand listing we currently track for a card, from one provider, and it moves as listings do. The typical market rate is the cross-provider median for the week of August 31, 2026, over every provider listing the card that week, so it is what the market charged rather than what the lowest-price seller charged. Shop on the first, budget from the second: the gap shows how exceptional the cheapest listing is relative to the broader market.
Methodology and data
Cheapest listing. The lowest current on-demand rate we track for a GPU model, normalized to USD per GPU per hour.
Per GPU, per hour. Multi-GPU instance prices are divided by GPU count; non-USD prices use current ECB exchange rates.
Typical market rate. The cross-provider median for the week of August 31, 2026: each provider's own median first, then the median across providers. Published for GPUs listed by at least 3 providers that week.
Availability. The provider's latest reported state: in stock, sold out or not reported. "Not reported" means we have no availability data, not that the GPU is unavailable.
LLM fit. Estimated memory is 4-bit weights plus the KV cache for one request at a 32K context plus 2 GB of runtime overhead, fitted into 90% of advertised VRAM. Named models use their architecture-specific cache rate; the parameter tiers use our default. Each setup is a machine a provider currently lists at that GPU count, priced at its own listing rather than a per-GPU rate multiplied out.
For full normalization and model-fit assumptions, see our cost estimate assumptions; for historical prices, the GPU Price Index; for what a rental dollar buys in VRAM and memory bandwidth, best value GPUs.