Lyceum
Our sponsor- On-Demand
- from $1.19
- Reserved
- on request
Cost-effective data center GPU for AI inference and media workloads.
Weekly median price per GPU per hour · Get the data
By provider shows one card per company, placed where its best offer ranked. By configuration lists every offer.
What you can rent comes first: in stock, then waitlist, then not reported, then out of stock. Priced offers rank ahead of quote-only, and on-demand ahead of other billing types.
Within a group, five factors set the order:
We match the provider, its country, GPU model, form factor, billing type, availability and instance name. Matching is partial and case-insensitive. Everyday words work too, so "interruptible" finds spot and "sold out" finds out of stock.
No providers match .
Heads up: A provider's own page may quote a different figure, for example on tax, region or a promotion. A monthly-only plan shows a derived hourly rate. Verify before provisioning. More on how we price.
One L40S has 48 GB of VRAM. In practice, that's enough memory for roughly 68B parameters at 4-bit or 17B at 16-bit, assuming a 32K context. Below are some open-weight LLMs, with the estimated memory and GPUs each one needs.
| Model | Memory (INT4 / FP4) | L40Ss needed | Cost /hr | Cost /mo |
|---|---|---|---|---|
|
|
18 GB
|
1
|
$1.50
|
$1,080
|
|
|
22 GB
|
1
|
$1.50
|
$1,080
|
|
|
67 GB
|
2
|
$3.00
|
$2,160
|
|
|
178 GB
|
8
|
$12.00
|
$8,640
|
|
|
239 GB
|
8
|
$12.00
|
$8,640
|
|
|
306 GB
|
8
|
$12.00
|
$8,640
|
|
|
1,544 GB
|
36
|
–
|
–
|
Estimates based on the median on-demand rate. Memory is weights plus FP8 KV cache at 32K context per request (FP16/BF16 in the 16-bit column). GPU counts assume 90% of advertised VRAM is usable. No guarantee of runtime support, usable performance, or that a matching quantized build exists. Pricing methodology.
Nvidia L40S · Per GPU
| Compute · dense | |
|---|---|
| FP8 | 733 TFLOPS 1,466 with sparsity |
| FP16 / BF16 | 362.1 TFLOPS 733 with sparsity |
| INT8 | 733 TOPS 1,466 with sparsity |
| FP32 | 91.6 TFLOPS |
| Precision support | FP8FP16BF16TF32FP32INT4INT8 |
| Memory | |
|---|---|
| Capacity | 48 GB GDDR6 |
| Bandwidth | 864 GB/s |
| Bus width | 384-bit |
| ECC | Yes |
| Silicon | |
|---|---|
| Architecture | Ada Lovelace |
| Process | TSMC 4N |
| Transistors | 76.3 billion |
| Shader cores | 18,176 CUDA cores |
| Matrix cores | 568 Tensor cores |
| Compute units | 142 SMs |
| Fabric and host | |
|---|---|
| Host interface | PCIe 4.0 x16 |
| Power | |
|---|---|
| Board power | 350 W |
| Cooling | Passive |
Source: official Nvidia L40S datasheet.
At 720 hours per month, one L40S costs an estimated $1,080 at the median on-demand price. Cheapest verified in stock: $389 per month on-demand, $353 reserved (6 mo), $245 spot.
GetDeploying currently tracks L40S configs from 38 providers. The cheapest verified in-stock on-demand configs come from Vast.ai, Novita, GPU.ai and Runpod. See the full price comparison above for every provider and config.
As of October 2, 2026, the median on-demand price has been flat over the past 90 days, at about $1.50 per GPU per hour, though it is about 9% above where it was a year ago.
One L40S runs models up to roughly 68B parameters at 4-bit quantization or 17B at 16-bit, assuming a 32K context. Larger models run across multiple GPUs: the model table above shows the estimated memory and GPU count for popular open-weight LLMs.
PCIe-only with no NVLink, limiting multi-GPU training performance. GDDR6 memory bandwidth (864 GB/s) is significantly lower than HBM-based data center GPUs.
Last updated