Runpod Our sponsor
USA 1 config 1x
Low-power professional card for small-form-factor workstations and dense inference nodes.
With a spread of only 12%, price optimization yields diminishing returns for the RTX 4000 SFF Ada. The market is relatively flat, with the lowest rate sitting at $0.44/hr per GPU (on-demand).
Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.
20GB GDDR6 with ECC inside a 70W board power envelope, so it runs in small-form-factor workstations and dense chassis that cannot supply or cool a full-height card. Ada Lovelace FP8 Tensor Cores cover small-model inference, AI development, rendering and viewport work.
The 70W envelope costs throughput: single-precision performance is roughly 30% below the full-height RTX 4000 Ada, which shares the same 20GB and pin count. There is no NVLink to pool memory across cards. For continuous cloud inference at a similar power draw, the L4 offers 24GB and higher memory bandwidth.
With 20GB of VRAM, the RTX 4000 SFF Ada can typically run 7B to 13B models in FP16, or larger models in 4-bit quantized form.
The RTX 4000 SFF Ada has 20GB of VRAM. Multi-GPU setups increase total memory, but that memory is not automatically pooled across GPUs.
The RTX 4000 SFF Ada has 280 GB/s of memory bandwidth. Higher bandwidth helps with faster data transfer between GPU memory and compute cores.
The RTX 4000 SFF Ada supports 7 precision formats. Training: BF16, FP16, TF32, FP32. Inference: FP8, INT4, INT8.
No. The RTX 4000 SFF Ada is a PCIe-only GPU with no NVLink, so it is better suited to single-GPU inference and smaller-scale workloads than large distributed training jobs.
RTX 4000 SFF Ada pricing currently ranges from $0.44/hr to $0.50/hr per GPU, depending on the provider, instance type, and billing model.
At 720 hours per month, one RTX 4000 SFF Ada can cost between $316.80 to $360.00 per month, depending on the provider. Reserved and spot pricing can lower that further.
The RTX 4000 SFF Ada is available from 3 cloud providers: Hetzner, Koyeb, Runpod.
Yes. We currently track 3 RTX 4000 SFF Ada listings across 3 cloud providers:
| Billing type | Listings | Avg $/GPU/hr |
|---|---|---|
| On-demand | 3 | $0.37/hr |
| GPU Memory | 20GB GDDR6 |
| Memory Interface | 160-bit |
| Memory Bandwidth | 280 GB/s |
| Error Correcting Code | Yes |
| Architecture | NVIDIA Ada Lovelace |
| CUDA Cores | 6,144 |
| Tensor Cores (Fourth-generation) | 192 |
| RT Cores (Third-generation) | 48 |
| Single-Precision Performance | 19.2 TFLOPS |
| RT Core Performance | 44.3 TFLOPS |
| Tensor Performance | 306.8 TFLOPS |
| System Interface | PCIe 4.0 x16 |
| Power Consumption | Total board power: 70W |
| Thermal Solution | Active |
| Form Factor | 2.7" H x 6.6" L, dual slot |
| Display Connectors | 4x mini DisplayPort 1.4a |
| Encode/Decode Engines | 2x encode, 2x decode (+AV1 encode and decode) |
| VR Ready | Yes |
| Graphics APIs | DirectX 12 Ultimate, Shader Model 6.7, OpenGL 4.6, Vulkan 1.3 |
| Compute APIs | CUDA 12.0, OpenCL 3.0, DirectCompute |
| NVIDIA NVLink® | No |
Full-height Ada professional card with the same memory capacity at a higher power envelope. From $0.08/hr per GPU across 4 providers.
Ada Lovelace cloud inference GPU in a comparable power envelope. From $0.11/hr per GPU across 16 providers.
Smaller Ada professional card at the same power envelope.. 1 provider, pricing on request.
Previous-generation compact professional GPU. From $0.03/hr per GPU across 10 providers.
Last updated