AWS In stock
USA 3 configs 1x-16x
Legacy Kepler hardware, not viable for modern workloads.
The K80 is available at 2 cloud providers for $0.90/hr per GPU.
USA 3 configs 1x-16x
USA 4 configs 2x
Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.
Dual-GPU card with 24GB total VRAM (12GB per die). Extremely low cost. Only suitable for basic GPU experimentation and legacy CUDA applications.
Kepler architecture, 5+ generations behind. No FP16 Tensor Cores. Only viable for basic CUDA learning exercises or legacy applications.
The median on-demand price has held steady around $0.90/hr per GPU across providers.
With 12GB of VRAM, the K80 is best for 7B-class models in 4-bit or 8-bit quantized form, and smaller models in FP16.
The K80 has 12GB of VRAM. Multi-GPU setups increase total memory, but that memory is not automatically pooled across GPUs.
The K80 has 240 GB/s of memory bandwidth. Higher bandwidth helps with faster data transfer between GPU memory and compute cores.
The K80 supports 2 precision formats. Training: FP32. Scientific: FP64.
No. The K80 is a PCIe-only GPU with no NVLink, so it is better suited to single-GPU inference and smaller-scale workloads than large distributed training jobs.
K80 pricing currently ranges from $0.08/hr to $0.90/hr per GPU, depending on the provider, instance type, and billing model.
At 720 hours per month, one K80 can cost between $54.50 to $648.00 per month, depending on the provider. Reserved and spot pricing can lower that further.
The K80 is available from 2 cloud providers: Amazon Web Services, Database Mart.
Yes. We currently track 7 K80 listings across 2 cloud providers:
| Billing type | Listings | Avg $/GPU/hr |
|---|---|---|
| On-demand | 3 | $0.90/hr |
| Reserved | 4 | $0.08/hr |
| Tesla K40 | Tesla K80 | |
|---|---|---|
| Peak double-precision floating point performance (board) | 1.43 Tflops | 1.87 Tflops |
| Peak single-precision floating point performance (board) | 4.29 Tflops | 5.6 Tflops |
| GPU | 1 x GK110B | 2 x GK210 |
| CUDA Cores | 2,880 | 4,992 |
| Memory size per board (GDDR5) | 12 GB | 24 GB |
| Memory bandwidth for board (ECC off) | 288 Gbytes/sec | 480 Gbytes/sec |
| Architecture features | SMX, Dynamic Parallelism, Hyper-Q | SMX, Dynamic Parallelism, Hyper-Q |
| System | Servers and workstations | Servers |
Source: official Nvidia K80 datasheet.
Last updated