Intel Gaudi 3
Intel's third-generation AI accelerator with 128GB of HBM2e and on-chip 200GbE networking for Ethernet scale-out.
- Arch
- Gaudi (3rd Gen)
- Memory
- 128 GB HBM2e
- Bandwidth
- 3.70 TB/s
- Released
- Q2 2024
At a glance
We don't currently track any Gaudi 3 offers. See alternative GPUs.
Technical Specifications
Intel Gaudi 3 · Per GPU
| Compute · dense | |
|---|---|
| FP8 | 1,678 TFLOPS |
| FP16 / BF16 | 1,678 TFLOPS |
| Precision support | FP8FP16BF16TF32FP32 |
| Memory | |
|---|---|
| Capacity | 128 GB HBM2e |
| Bandwidth | 3,700 GB/s |
| Silicon | |
|---|---|
| Architecture | Gaudi (3rd Gen) |
| Process | TSMC 5nm |
| Fabric and host | |
|---|---|
| GPU interconnect | RoCE 1,200 GB/s |
| Host interface | PCIe 5.0 x16 |
| Power | |
|---|---|
| Board power | 900 W |
Source: official Intel Gaudi 3 datasheet.
Frequently Asked Questions
Why choose the Gaudi 3?
Comes with 128GB HBM2e per card and 24x 200GbE RoCEv2 ports on-chip, so an 8-card node reaches 1TB of accelerator memory and scales out over standard Ethernet without separate NICs. BF16 and FP8 run at the same peak rate, and a 70B-parameter model in FP8 fits on a single card.
When is the Gaudi 3 not a good fit?
The Gaudi software stack supports fewer frameworks, kernels and models than CUDA, so workloads relying on custom CUDA kernels or niche libraries need porting. For broad ecosystem compatibility, H100 or H200 are easier to adopt; Gaudi 3 remains a fit for PyTorch and Hugging Face workloads that run on Gaudi as is.
Alternatives to Intel Gaudi 3
-
Nvidia H100
80 GB HBM3 · released Q3 2022
From $1.30 /GPU/hr Compare -
Nvidia H200
141 GB HBM3e · 30% more bandwidth than the Gaudi 3
From $2.09 /GPU/hr Compare -
AMD MI300X
192 GB HBM3 · 43% more bandwidth than the Gaudi 3
From $2.59 /GPU/hr Compare -
Intel Gaudi 2
96 GB HBM2e · 34% less bandwidth than the Gaudi 3
From $1.02 /GPU/hr Compare
Last updated