Vast.ai
In stock- On-Demand
- from $0.06
- Reserved
- from $0.06
- Spot
- from $0.0400
Turing GPU for basic local AI experimentation.
Weekly median price per GPU per hour · Get the data
By provider shows one card per company, placed where its best offer ranked. By configuration lists every offer.
What you can rent comes first: in stock, then waitlist, then not reported, then out of stock. Priced offers rank ahead of quote-only, and on-demand ahead of other billing types.
Within a group, five factors set the order:
We match the provider, its country, GPU model, form factor, billing type, availability and instance name. Matching is partial and case-insensitive. Everyday words work too, so "interruptible" finds spot and "sold out" finds out of stock.
No providers match .
Heads up: A provider's own page may quote a different figure, for example on tax, region or a promotion. A monthly-only plan shows a derived hourly rate. Verify before provisioning. More on how we price.
One RTX 2060 Super has 8 GB of VRAM. Even at 4-bit, the weights of most modern AI models exceed the card's usable VRAM, and context adds KV cache on top. Below are some open-weight LLMs, with the estimated memory and GPUs each one needs.
| Model | Memory (INT4 / FP4) | RTX 2060 Supers needed | Cost /hr | Cost /mo |
|---|---|---|---|---|
|
|
19 GB
|
3
|
–
|
–
|
|
|
24 GB
|
4
|
–
|
–
|
|
|
68 GB
|
10
|
–
|
–
|
|
|
178 GB
|
25
|
–
|
–
|
|
|
241 GB
|
34
|
–
|
–
|
|
|
306 GB
|
43
|
–
|
–
|
|
|
1,545 GB
|
215
|
–
|
–
|
Estimates based on the median on-demand rate. Memory is weights plus FP16/BF16 KV cache at 32K context per request. GPU counts assume 90% of advertised VRAM is usable. No guarantee of runtime support, usable performance, or that a matching quantized build exists. Pricing methodology.
Nvidia RTX 2060 Super · Per GPU
| Compute · dense | |
|---|---|
| FP16 / BF16 | 28.7 TFLOPS |
| INT8 | 114.8 TOPS |
| FP32 | 7.2 TFLOPS |
| Precision support | FP16FP32INT4INT8 |
| Memory | |
|---|---|
| Capacity | 8 GB GDDR6 |
| Bandwidth | 448 GB/s |
| Bus width | 256-bit |
| ECC | No |
| Silicon | |
|---|---|
| Architecture | Turing |
| Process | TSMC 12 nm FFN |
| Transistors | 10.8 billion |
| Shader cores | 2,176 CUDA cores |
| Matrix cores | 272 Tensor cores |
| Compute units | 34 SMs |
| Fabric and host | |
|---|---|
| Host interface | PCIe 3.0 x16 |
| Power | |
|---|---|
| Board power | 175 W |
| Cooling | Active |
As of September 21, 2026, we track 9 configs from 1 provider. Prices are per GPU per hour.
| Billing type | Configs | Cheapest |
|---|---|---|
On-demand | 5 | $0.06 (in stock, Vast.ai) |
Reserved | 2 | $0.06 (3 mo, in stock, Vast.ai) |
Spot | 2 | $0.04 (in stock, Vast.ai) |
No median for on-demand, reserved and spot: we only show this when at least 3 providers list the GPU on that billing type.
The cheapest verified in-stock estimate is $43 per month on-demand, $43 reserved (3 mo), $29 spot.
GetDeploying currently tracks RTX 2060 Super configs from 1 provider. Vast.ai has verified in-stock on-demand configs. See the full price comparison above for every provider and config.
More than one: even at 4-bit, the weights of most modern AI models exceed one RTX 2060 Super's usable VRAM, and context adds KV cache on top. Larger models run across multiple GPUs: the model table above shows the estimated memory and GPU count for popular open-weight LLMs.
8GB GDDR6 with Turing Tensor Cores. Very low cost for basic GPU experimentation.
8GB VRAM and Turing limitations make it a poor fit for modern AI workloads. Better suited for basic experimentation.
Last updated