Vultr In stock
USA 1 config 1x
Grace Hopper superchip combining ARM CPU with H100 GPU.
The GH200 is listed by 3 cloud providers. Available capacity currently sits around $6.50/hr per GPU (on-demand). Spot instances start lower, at $1.99/hr per GPU.
USA 1 config 1x
USA 1 config 1x
USA 1 config 1x
Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.
Grace Hopper superchip combines ARM CPU with H100 GPU and 96GB or 144GB HBM3e. Unified memory architecture with 900 GB/s NVLink. Efficient for AI workloads that benefit from tight CPU-GPU integration.
ARM CPU ecosystem may have compatibility issues with some software stacks. For standard GPU cloud workloads, a regular H100 is simpler to deploy and more widely available.
The median on-demand price across providers has fallen about 35% since August 2025, from $3.90 to $2.52/hr per GPU. This reflects the market-wide median, which moves both when providers change prices and when lower or higher priced offerings enter the market.
With 96GB of VRAM, the GH200 is well suited to 30B-class models in FP16, and 70B-class models in 4-bit or 8-bit quantized form.
The GH200 has 96GB of VRAM. Multi-GPU setups increase total memory, but that memory is not automatically pooled across GPUs.
The GH200 has 4,000 GB/s of memory bandwidth. Higher bandwidth helps with faster data transfer between GPU memory and compute cores.
The GH200 supports 7 precision formats. Training: BF16, FP16, TF32, FP32. Inference: FP8, INT8. Scientific: FP64.
Yes. The GH200 supports NVLink with 900 GB/s of bidirectional bandwidth. This helps accelerate multi-GPU communication.
GH200 pricing currently ranges from $1.99/hr to $6.50/hr per GPU, depending on the provider, instance type, and billing model.
At 720 hours per month, one GH200 can cost between $1,432.80 to $4,680.00 per month, depending on the provider. Reserved and spot pricing can lower that further.
The GH200 is available from 3 cloud providers: CoreWeave, Vultr, Lambda Labs.
Yes. We currently track 3 GH200 listings across 3 cloud providers:
| Billing type | Listings | Avg $/GPU/hr |
|---|---|---|
| On-demand | 2 | $4.39/hr |
| Spot | 1 | $1.99/hr |
| Feature | GH200 | GH200 NVL2 |
|---|---|---|
| CPU core count | 72 Arm Neoverse V2 cores | 144 Arm Neoverse V2 cores |
| L1 cache | 64KB i-cache + 64KB d-cache | 64KB i-cache + 64KB d-cache |
| L2 cache | 1MB per core | 1MB per core |
| L3 cache | 114MB | 228MB |
| Base frequency | all-core single instruction, multiple data (SIMD) frequency | 3.1GHz | 3.0GHz |
| LPDDR5X size | 480GB, 120GB, 240GB | 960GB, 240GB, 480GB |
| Memory bandwidth | Up to 384GB/s, Up to 512GB/s | Up to 768GB/s, Up to 1024GB/s |
| PCIe links | Up to 4x PCIe x16 (Gen5) | Up to 8x PCIe x16 (Gen5) |
| FP64 | 34 teraFLOPS | 68 teraFLOPS |
| FP64 Tensor Core | 67 teraFLOPS | 134 teraFLOPS |
| FP32 | 67 teraFLOPS | 134 teraFLOPS |
| TF32 Tensor Core | 989 teraFLOPS*, 494 teraFLOPS | 1,979 teraFLOPS*, 990 teraFLOPS |
| BFLOAT16 Tensor Core | 1,979 teraFLOPS*, 990 teraFLOPS | 3,958 teraFLOPS*, 1,979 teraFLOPS |
| FP16 Tensor Core | 1,979 teraFLOPS*, 990 teraFLOPS | 3,958 teraFLOPS*, 1,979 teraFLOPS |
| FP8 Tensor Core | 3,958 teraFLOPS*, 1,979 teraFLOPS | 7,916 teraFLOPS*, 3,958 teraFLOPS |
| INT8 Tensor Core | 3,958 TOPS*, 1,979 TOPS | 7,916 TOPS*, 3,958 TOPS |
| High-bandwidth memory (HBM) size | 96GB HBM3 | 144GB HBM3e, Up to 288GB HBM3e |
| Memory bandwidth | Up to 4TB/s, Up to 4.9TB/s | Up to 9.8TB/s |
| NVIDIA NVLink-C2C CPU-to-GPU bandwidth | 900 GB/s | 900 GB/s |
| Power | Configurable 450 to 1000W (Memory + CPU + GPU) | Configurable 900W to 2000W (Memory + CPU + GPU) |
| Thermal solution | Air cooled or liquid cooled | Air cooled or liquid cooled |
* with sparsity
Source: official Nvidia GH200 datasheet.
Last updated