Nemotron 3 Super 120B A12B

View website

Released March 2026 under the NVIDIA Nemotron Open Model License, a 120B LatentMoE hybrid activating 12B parameters per token, combining Mamba-2, MoE, and attention layers with Multi-Token Prediction for faster generation. Trained with NVFP4 quantization, covers seven languages, and needs at least 8x H100 80GB to serve.

At a glance

Context window
1M tokens
Knowledge cutoff
Jun 2025
Modalities
Text Text Text Text

Capabilities

Function calling

Function calling

Connect to external tools, APIs, and systems.

Structured output

Structured output

Return responses in structured formats like JSON.

Pricing by provider

Provider Input / 1M tokens Output / 1M tokens
Geodd logo Geodd $0.09 $0.50 View

Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.

Compare with other models

Estimated prices shown. Actual costs may vary based on context length, batch size, caching, and provider-specific pricing tiers.