Lyceum Our sponsor
Germany 5 configs 1x-8x
Cost-effective data center GPU for AI inference and media workloads.
Listings for the L40S reach $7.58/hr, often reflecting a premium for high availability. However, you might be able to find available instances from as low as $0.47/hr per GPU (on-demand). Spot instances start lower, at $0.33/hr per GPU.
Germany 5 configs 1x-8x
USA 1 config 1x
USA 8 configs 1x-8x
USA 7 configs 1x-8x
USA 4 configs 1x-8x
USA 1 config 1x
USA 7 configs 1x-10x
USA 24 configs 1x-8x
USA 9 configs 1x-4x
France 7 configs 1x-4x
France 1 config 1x
Finland 22 configs 1x-3x
France 4 configs 1x-8x
UAE 3 configs 1x-4x
Singapore 24 configs 1x-4x
Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.
48GB GDDR6 with Ada Lovelace FP8 support. Strong inference throughput at a lower cost than A100 or H100. AV1 hardware encoding makes it versatile for media and AI workloads.
PCIe-only with no NVLink, limiting multi-GPU training performance. GDDR6 memory bandwidth (864 GB/s) is significantly lower than HBM-based data center GPUs.
The median on-demand price has held steady around $1.45/hr per GPU across providers.
With 48GB of VRAM, the L40S can typically run models up to about 30B parameters in FP16, or 70B-class models in 4-bit quantized form for inference.
The L40S has 48GB of VRAM. Multi-GPU setups increase total memory, but that memory is not automatically pooled across GPUs.
The L40S has 864 GB/s of memory bandwidth. Higher bandwidth helps with faster data transfer between GPU memory and compute cores.
The L40S supports 7 precision formats. Training: BF16, FP16, TF32, FP32. Inference: FP8, INT4, INT8.
No. The L40S is a PCIe-only GPU with no NVLink, so it is better suited to single-GPU inference and smaller-scale workloads than large distributed training jobs.
L40S pricing currently ranges from $0.33/hr to $7.58/hr per GPU, depending on the provider, instance type, and billing model.
At 720 hours per month, one L40S can cost between $240.26 to $5,455.58 per month, depending on the provider. Reserved and spot pricing can lower that further.
The L40S is available from 32 cloud providers, including Verda, Amazon Web Services, Spheron. Pricing and availability vary by region and billing model.
Yes. We currently track 222 L40S listings across 32 cloud providers:
| Billing type | Listings | Avg $/GPU/hr |
|---|---|---|
| On-demand | 108 | $1.74/hr |
| Reserved | 75 | $1.31/hr |
| Spot | 39 | $1.11/hr |
| GPU Architecture | NVIDIA Ada Lovelace Architecture |
| GPU Memory | 48GB GDDR6 |
| Memory Bandwidth | 864GB/s |
| Interconnect Interface | PCIe Gen4 x16: 64GB/s bidirectional |
| NVIDIA Ada Lovelace Architecture-Based CUDA® Cores | 18,176 |
| NVIDIA Third-Generation RT Cores | 142 |
| NVIDIA Fourth-Generation Tensor Cores | 568 |
| RT Core Performance TFLOPS | 209 |
| FP32 TFLOPS | 91.6 |
| TF32 Tensor Core TFLOPS | 183 |
| BFLOAT16 Tensor Core TFLOPS | 362.05 |
| FP16 Tensor Core | 362.05 |
| FP8 Tensor Core | 733 |
| Peak INT8 Tensor TOPS | 733 |
| Peak INT4 Tensor TOPS | 733 |
| Form Factor | 4.4" (H) x 10.5" (L), dual slot |
| Display Ports | 4x DisplayPort 1.4a |
| Max Power Consumption | 350W |
| Power Connector | 16-pin |
| Thermal | Passive |
| Virtual GPU (vGPU) Software Support | Yes |
| vGPU Profiles Supported | See the virtual GPU licensing guide |
| NVENC, NVDEC | 3x, 3x (includes AV1 encode and decode) |
| Secure Boot With Root of Trust | Yes |
| NEBS Ready | Level 3 |
| MIG Support | No |
| NVIDIA® NVLink® Support | No |
Source: official Nvidia L40S datasheet.
Last updated