Runpod Our sponsor
USA 4 configs 1x-8x
Data center GPU for combined AI inference and visualization.
Listings for the L40 reach $1.64/hr, often reflecting a premium for high availability. However, you might be able to find available instances from as low as $0.58/hr per GPU (on-demand). Spot instances start lower, at $0.47/hr per GPU.
USA 4 configs 1x-8x
USA 6 configs 1x-4x
USA 1 config 1x
USA 3 configs 1x
USA 4 configs 1x-8x
France 7 configs 1x-8x
UAE 4 configs 1x-8x
Singapore 9 configs 1x-8x
USA 1 config 1x
UK 3 configs 1x
UK 1 config 1x
USA 2 configs 8x
Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.
48GB GDDR6 with Ada Lovelace architecture. Good for visualization, rendering, and inference workloads. FP8 Tensor Core support for efficient AI inference.
PCIe-only. Lower inference throughput than L40S for pure AI workloads. If you don't need the visualization features, the L40S is typically a better value.
The median on-demand price across providers has fallen about 8% since August 2025, from $1.09 to $1.00/hr per GPU. This reflects the market-wide median, which moves both when providers change prices and when lower or higher priced offerings enter the market.
With 48GB of VRAM, the L40 can typically run models up to about 30B parameters in FP16, or 70B-class models in 4-bit quantized form for inference.
The L40 has 48GB of VRAM. Multi-GPU setups increase total memory, but that memory is not automatically pooled across GPUs.
The L40 has 864 GB/s of memory bandwidth. Higher bandwidth helps with faster data transfer between GPU memory and compute cores.
The L40 supports 7 precision formats. Training: BF16, FP16, TF32, FP32. Inference: FP8, INT4, INT8.
No. The L40 is a PCIe-only GPU with no NVLink, so it is better suited to single-GPU inference and smaller-scale workloads than large distributed training jobs.
L40 pricing currently ranges from $0.47/hr to $1.64/hr per GPU, depending on the provider, instance type, and billing model.
At 720 hours per month, one L40 can cost between $339.19 to $1,180.08 per month, depending on the provider. Reserved and spot pricing can lower that further.
The L40 is available from 12 cloud providers, including Spheron, Sesterce, Runpod. Pricing and availability vary by region and billing model.
Yes. We currently track 45 L40 listings across 12 cloud providers:
| Billing type | Listings | Avg $/GPU/hr |
|---|---|---|
| On-demand | 33 | $0.98/hr |
| Reserved | 3 | $0.64/hr |
| Spot | 9 | $0.78/hr |
| GPU Architecture | NVIDIA Ada Lovelace architecture |
| GPU Memory | 48GB GDDR6 |
| Memory Bandwidth | 864GB/s |
| Interconnect Interface | PCIe Gen4 x16 (64GB/s bi-directional) |
| CUDA Cores | 18,176 |
| Third-Generation RT Cores | 142 |
| Fourth-Generation Tensor Cores | 568 |
| RT Core Performance TFLOPS | 209 |
| FP32 TFLOPS | 90.5 |
| TF32 Tensor Core TFLOPS | 90.5 |
| BFLOAT16 Tensor Core TFLOPS | 181.05 |
| FP16 Tensor Core TFLOPS | 181.05 |
| FP8 Tensor Core TFLOPS | 362 |
| Peak INT8 Tensor TOPS | 362 |
| Peak INT4 Tensor TOPS | 724 |
| Form Factor | 4.4" (H) x 10.5" (L) - dual slot |
| Display Ports | 4x DisplayPort 1.4a |
| Max Power Consumption | 300W |
| NVLink Support | No |
Source: official Nvidia L40 datasheet.
Last updated