Vast.ai Our sponsor
USA 5 configs 1x-2x
Low-cost inference GPU with wide cloud availability.
Of 122 listings across 10 cloud providers, 94 are verified in stock. Cheapest of those: on-demand from $0.13 per GPU per hour, spot from $0.06, reservations from $0.21 (36 mo). The median on-demand price is $0.81, down 6% over the past 90 days.
USA 5 configs 1x-2x
USA 2 configs 1x-2x
USA 2 configs 1x-2x
USA 16 configs 1x-4x
USA 21 configs 1x-8x
USA 48 configs 1x-4x
USA 4 configs 1x-8x
USA 1 config 1x
Singapore 12 configs 1x-4x
Netherlands 11 configs 1x-2x
Heads up: Prices are estimates from published rates for common setups and vary by region and usage; verify with the provider before provisioning.
16GB GDDR6 in a 70W low-power form factor. Very low cost per hour across cloud providers. INT8 Tensor Cores for efficient inference on smaller models.
Turing architecture lacks BF16 and FP8. 16GB VRAM limits model sizes. Slow for training. For production inference, the L4 offers significantly better throughput.
The median on-demand price across providers has fallen about 6% over the past 90 days, from about $0.86 to $0.81 per GPU per hour.
At 720 hours per month, one T4 costs about $583 at the median on-demand price. Cheapest verified in stock: $97 per month on-demand, $154 reserved (36 mo), $40 spot.
The T4 is listed by 10 providers. Cheapest verified in stock: Vast.ai, Theta EdgeCloud, Amazon Web Services and Microsoft Azure.
The T4 supports 4 precision formats. Training: FP16, FP32. Inference: INT4, INT8.
No. Multi-GPU setups communicate over PCIe.
One T4 has 16 GB GDDR6. The table shows the smallest T4 node that fits each model, assuming moderate context, and what that node costs at the cheapest in-stock on-demand rate for one T4, $0.13 per hour.
| Model | 4-bit | 8-bit | BF16 |
|---|---|---|---|
| 2x (16 GB), from $0.26/hr | 4x (32 GB), from $0.52/hr | 8x (65 GB), from $1.04/hr | |
| 2x (18 GB), from $0.26/hr | 4x (37 GB), from $0.52/hr | 8x (74 GB), from $1.04/hr | |
| 4x (42 GB), from $0.52/hr | 8x (84 GB), from $1.04/hr | Over 8x (168 GB) | |
| 8x (70 GB), from $1.04/hr | Over 8x (140 GB) | Over 8x (281 GB) | |
| Over 8x (141 GB) | Over 8x (282 GB) | Over 8x (564 GB) | |
| Over 8x (452 GB) | Over 8x (904 GB) | Over 8x (1,807 GB) |
Estimated memory required is parameters × bytes per weight at each precision, plus 20% for KV cache and runtime overhead. Anything larger than 8x needs more than one node.
As of August 27, 2026, the median on-demand price is $0.81 per GPU per hour across 9 providers.
| Billing type | Listings | Median / GPU / hr | Cheapest in stock |
|---|---|---|---|
| On-demand | 41 | $0.81 | $0.13 (Vast.ai) |
| Reserved | 63 | $0.51 | $0.21 (36 mo, Google Cloud) |
| Spot | 18 | $0.27 | $0.06 (Vast.ai) |
| GPU Architecture | NVIDIA Turing |
| NVIDIA Turing Tensor Cores | 320 |
| NVIDIA CUDA® Cores | 2,560 |
| Single-Precision | 8.1 TFLOPS |
| Mixed-Precision (FP16/FP32) | 65 TFLOPS |
| INT8 | 130 TOPS |
| INT4 | 260 TOPS |
| GPU Memory | 16 GB GDDR6 |
| Memory Bandwidth | 300 GB/sec |
| ECC | Yes |
| Interconnect Bandwidth | 32 GB/sec |
| System Interface | x16 PCIe Gen3 |
| Form Factor | Low-Profile PCIe |
| Thermal Solution | Passive |
| Compute APIs | CUDA, NVIDIA TensorRT™, ONNX |
Source: official Nvidia T4 datasheet.
Last updated