Vultr In stock
USA 3 configs 2x-8x
Purpose-built for virtual desktop (VDI) deployments.
We don't see much volatility for the A16 right now. Most providers are clustered between $0.47 and $0.56/hr per GPU, so availability is likely the deciding factor.
Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.
4x 16GB GPUs on a single card (64GB total). Designed for virtual desktop (VDI) and multi-user GPU sharing. Hardware video encode/decode support.
Built for VDI and multi-user GPU sharing. Individual GPU dies have limited compute, so better suited for remote desktops and lightweight graphics than AI workloads.
The median on-demand price across providers has risen about 8% since August 2025, from $0.52 to $0.56/hr per GPU. This reflects the market-wide median, which moves both when providers change prices and when lower or higher priced offerings enter the market.
With 16GB of VRAM, the A16 is best for 7B-class models in 4-bit or 8-bit quantized form, and smaller models in FP16.
The A16 has 16GB of VRAM. Multi-GPU setups increase total memory, but that memory is not automatically pooled across GPUs.
The A16 has 200 GB/s of memory bandwidth. Higher bandwidth helps with faster data transfer between GPU memory and compute cores.
The A16 supports 6 precision formats. Training: BF16, FP16, TF32, FP32. Inference: INT4, INT8.
No. The A16 is a PCIe-only GPU with no NVLink, so it is better suited to single-GPU inference and smaller-scale workloads than large distributed training jobs.
A16 pricing currently ranges from $0.47/hr to $0.56/hr per GPU, depending on the provider, instance type, and billing model.
At 720 hours per month, one A16 can cost between $339.12 to $405.90 per month, depending on the provider. Reserved and spot pricing can lower that further.
The A16 is available from 3 cloud providers: Vultr, Sesterce, Runcrate.
Yes. We currently track 6 A16 listings across 3 cloud providers:
| Billing type | Listings | Avg $/GPU/hr |
|---|---|---|
| On-demand | 6 | $0.52/hr |
| GPU Architecture | NVIDIA Ampere architecture |
| GPU Memory | 4x 16 GB GDDR6 |
| Memory Bandwidth | 4x 200 GB/s |
| Error-Correcting Code (ECC) | Yes |
| NVIDIA Ampere architecture-based CUDA Cores | 4x 1280 |
| NVIDIA Third-Generation Tensor Cores | 4x 40 |
| NVIDIA Second-Generation RT Cores | 4x 10 |
| FP32 | TF32 | TF32' (TFLOPS) | 4x 4.5, 4x 9, 4x 18 |
| FP16 | FP16' (TFLOPS) | 4x 17.9, 4x 35.9 |
| INT8 | INT8' (TOPS) | 4x 35.9, 4x 71.8 |
| System Interface | PCIe Gen4 (x16) |
| Max Power Consumption | 250W |
| Thermal Solution | Passive |
| Form Factor | Full height, full length (FHFL) Dual Slot |
| Power Connector | 8-pin CPU |
| Encode/Decode Engines | 4 NVENC, 8 NVDEC (includes AV1 decode) |
| Secure and Measured Boot with Hardware Root of Trust for GPU | Yes (optional) |
| vGPU Software Support | NVIDIA Virtual PC (vPC), NVIDIA Virtual Applications (vApps), NVIDIA RTX Virtual Workstation (vWS), NVIDIA AI Enterprise, NVIDIA Virtual Compute Server (vCS) |
| Graphics APIs | DirectX 12.07, Shader Model 5.17, OpenGL 4.68, Vulkan 1.18 |
| Compute APIs | CUDA, DirectCompute, OpenCL™, OpenACC® |
| MIG Support | No |
Source: official Nvidia A16 datasheet.
Last updated