Vast.ai Our sponsor
USA 2 configs 1x
Data center GPU for large-batch inference and professional visualization.
For smaller projects, the A40 is available around $0.34/hr per GPU (on-demand). This tier offers a 83% discount compared to the higher end of the market ($2.05/hr). Spot instances start lower, at $0.32/hr per GPU.
USA 2 configs 1x
USA 5 configs 1x-8x
UAE 1 config 1x
France 1 config 1x
Switzerland 4 configs 1x-8x
USA 4 configs 1x
Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.
48GB GDDR6 with Ampere Tensor Cores. Large VRAM for inference on bigger models. Supports both AI and professional visualization workloads.
No FP8 support. PCIe-only. For pure AI inference, the L40S offers better performance with Ada Lovelace architecture at similar or lower cost.
The median on-demand price across providers has risen about 28% since August 2025, from $0.58 to $0.74/hr per GPU. This reflects the market-wide median, which moves both when providers change prices and when lower or higher priced offerings enter the market.
With 48GB of VRAM, the A40 can typically run models up to about 30B parameters in FP16, or 70B-class models in 4-bit quantized form for inference.
The A40 has 48GB of VRAM. Multi-GPU setups increase total memory, but that memory is not automatically pooled across GPUs.
The A40 has 696 GB/s of memory bandwidth. Higher bandwidth helps with faster data transfer between GPU memory and compute cores.
The A40 supports 6 precision formats. Training: BF16, FP16, TF32, FP32. Inference: INT4, INT8.
No. The A40 is a PCIe-only GPU with no NVLink, so it is better suited to single-GPU inference and smaller-scale workloads than large distributed training jobs.
A40 pricing currently ranges from $0.32/hr to $2.05/hr per GPU, depending on the provider, instance type, and billing model.
At 720 hours per month, one A40 can cost between $231.98 to $1,473.12 per month, depending on the provider. Reserved and spot pricing can lower that further.
The A40 is available from 6 cloud providers, including Runpod, Exoscale, Sesterce. Pricing and availability vary by region and billing model.
Yes. We currently track 17 A40 listings across 6 cloud providers:
| Billing type | Listings | Avg $/GPU/hr |
|---|---|---|
| On-demand | 11 | $0.80/hr |
| Reserved | 5 | $0.47/hr |
| Spot | 1 | $0.32/hr |
| GPU architecture | NVIDIA Ampere architecture |
| GPU memory | 48 GB GDDR6 |
| Memory bandwidth | 696 GB/s |
| Interconnect interface | NVIDIA® NVLink® 112.5 GB/s (bidirectional), PCIe Gen4: 64GB/s |
| NVIDIA Ampere architecture-based CUDA Cores | 10,752 |
| NVIDIA second-generation RT Cores | 84 |
| NVIDIA third-generation Tensor Cores | 336 |
| Peak FP32 TFLOPS (non-Tensor) | 37.4 |
| Peak FP16 Tensor TFLOPS with FP16 Accumulate | 149.7, 299.4* |
| Peak TF32 Tensor TFLOPS | 74.8, 149.6* |
| RT Core performance TFLOPS | 73.1 |
| Peak BF16 Tensor TFLOPS with FP32 Accumulate | 149.7, 299.4* |
| Peak INT8 Tensor TOPS, Peak INT4 Tensor TOPS | 299.3, 598.6, 598.7, 1,197.4 |
| Form factor | 4.4” (H) x 10.5” (L) dual slot |
| Display ports | 3x DisplayPort 1.4**, Supports NVIDIA Mosaic and Quadro® Sync |
| Max power consumption | 300 W |
| Power connector | 8-pin CPU |
| Thermal solution | Passive |
| Virtual GPU (vGPU) software support | NVIDIA vPC/vApps, NVIDIA RTX Virtual Workstation, NVIDIA Virtual Compute Server |
| vGPU profiles supported | See the Virtual GPU Licensing Guide |
| NVENC, NVDEC | 1x, 2x (includes AV1 decode) |
| Secure and measured boot with hardware root of trust | Yes (optional) |
| NEBS ready | Level 3 |
| Compute APIs | CUDA, DirectCompute, OpenCL™, OpenACC® |
| Graphics APIs | DirectX 12.07, Shader Model 5.17, OpenGL 4.68, Vulkan 1.18 |
| MIG support | No |
Source: official Nvidia A40 datasheet.
Last updated