Vultr
Out of stock- On-Demand
- from $0.47
Purpose-built for virtual desktop (VDI) deployments.
Weekly median price per GPU per hour · Get the data
By provider shows one card per company, placed where its best offer ranked. By configuration lists every offer.
What you can rent comes first: in stock, then waitlist, then not reported, then out of stock. Priced offers rank ahead of quote-only, and on-demand ahead of other billing types.
Within a group, five factors set the order:
We match the provider, its country, GPU model, form factor, billing type, availability and instance name. Matching is partial and case-insensitive. Everyday words work too, so "interruptible" finds spot and "sold out" finds out of stock.
No providers match .
Heads up: A provider's own page may quote a different figure, for example on tax, region or a promotion. A monthly-only plan shows a derived hourly rate. Verify before provisioning. More on how we price.
One A16 has 16 GB of VRAM. In practice, that's enough memory for roughly 8B parameters at 4-bit or 2B at 16-bit, assuming a 32K context. Below are some open-weight LLMs, with the estimated memory and GPUs each one needs.
| Model | Memory (INT4 / FP4) | A16s needed | Cost /hr | Cost /mo |
|---|---|---|---|---|
|
|
19 GB
|
2
|
–
|
–
|
|
|
24 GB
|
2
|
–
|
–
|
|
|
68 GB
|
8
|
–
|
–
|
|
|
178 GB
|
13
|
–
|
–
|
|
|
241 GB
|
17
|
–
|
–
|
|
|
306 GB
|
22
|
–
|
–
|
|
|
1,545 GB
|
108
|
–
|
–
|
Estimates based on the median on-demand rate. Memory is weights plus FP16/BF16 KV cache at 32K context per request. GPU counts assume 90% of advertised VRAM is usable. No guarantee of runtime support, usable performance, or that a matching quantized build exists. Pricing methodology.
Nvidia A16 · Per GPU
| Compute · dense | |
|---|---|
| FP16 / BF16 | 17.9 TFLOPS 35.9 with sparsity |
| INT8 | 35.9 TOPS 71.8 with sparsity |
| FP32 | 4.5 TFLOPS |
| Precision support | FP16BF16TF32FP32INT4INT8 |
| Memory | |
|---|---|
| Capacity | 16 GB GDDR6 |
| Bandwidth | 200 GB/s |
| Bus width | 128-bit |
| ECC | Yes |
| Silicon | |
|---|---|
| Architecture | Ampere |
| Shader cores | 1,280 CUDA cores |
| Matrix cores | 40 Tensor cores |
| Compute units | 10 SMs |
| Fabric and host | |
|---|---|
| Host interface | PCIe 4.0 x16 |
| Power | |
|---|---|
| Cooling | Passive |
Source: official Nvidia A16 datasheet.
As of October 3, 2026, we track 8 configs from 2 providers. Prices are per GPU per hour.
| Billing type | Configs | Median /GPU/hr | Cheapest |
|---|---|---|---|
On-demand | 8 | Sold out |
GetDeploying currently tracks A16 configs from 2 providers. See the full price comparison above for every provider and config.
One A16 runs models up to roughly 8B parameters at 4-bit quantization or 2B at 16-bit, assuming a 32K context. Larger models run across multiple GPUs: the model table above shows the estimated memory and GPU count for popular open-weight LLMs.
4x 16GB GPUs on a single card (64GB total). Designed for virtual desktop (VDI) and multi-user GPU sharing. Hardware video encode/decode support.
Built for VDI and multi-user GPU sharing. Individual GPU dies have limited compute, so better suited for remote desktops and lightweight graphics than AI workloads.
Last updated