Nvidia logo

Nemotron 3 Ultra 550B A55B

Open weights · OpenMDW-1.1

The largest of NVIDIA's Nemotron 3 line, a 550B LatentMoE hybrid activating 55B per token across Mamba-2, MoE and attention layers with Multi-Token Prediction, under the OpenMDW-1.1 license.

Price via Together
Input
$0.60 / 1M tokens
Output
$3.60 / 1M tokens

About $13.20 for 10M input and 2M output tokens. Estimate yours

Together Nvidia 2 providers

Key Specifications

Context window
1M tokens
Knowledge cutoff
Inputs
Text Outputs: Text

Nemotron 3 Ultra 550B A55B pricing by provider

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
Together logo Together $0.60 $3.60 $13.20 View
Nvidia logo Nvidia Creator Open weights, no hosted price View

Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning. LLM rates are base-tier, on-demand prices per 1M tokens; cached-input, batch and long-context tiers are not included.

Compare every model at this volume in the LLM cost calculator.

Capabilities

Function calling

Function calling

Connect to external tools, APIs, and systems.

Structured output

Structured output

Return responses in structured formats like JSON.

Estimated cost to self-host Nemotron 3 Ultra 550B A55B

Nemotron 3 Ultra 550B A55B has 550B parameters, about 55B active per token. At 4-bit it needs about 330 GB of GPU memory, at 8-bit 660 GB and at BF16 1,320 GB, counting 20% on top of the weights for KV cache and runtime overhead.

Precision Memory needed Cheapest rentals that fit Per month, 24/7
4-bit 330 GB Nvidia logo 8× RTX A6000 ($4.48/hr) Amd logo 2× MI300X ($5.80/hr) $3,226
8-bit 660 GB Amd logo 4× MI300X ($11.60/hr) Amd logo 4× MI325X ($12.28/hr) $8,352
BF16 1,320 GB Amd logo 8× MI300X ($23.20/hr) Amd logo 8× MI325X ($24.56/hr) $16,704

$4.48 an hour (8x RTX A6000) buys about 4.1M tokens an hour at the cheapest hosted rate we track ($0.60 in / $3.60 out per 1M tokens on Together, 5:1 input to output); below that volume the API is cheaper, before idle time. Compare hosted costs.

Memory is parameters × bytes per weight at each precision, plus 20% for KV cache and runtime overhead. Rentals are the cheapest cards that hold it, at provider-weighted median on-demand prices for the week of August 17, 2026, consumer cards included, in nodes of up to 8 GPUs; a fit with under 15% headroom is marked tight. A month is 720 hours. See the best-value GPUs guide for the same table across models, and cloud GPU pricing for every card.

More from Nvidia

Model Context Input / 1M Output / 1M
Nvidia logo Nemotron 3 Super 120B A12B Open weights 1M $0.09 $0.50
Nvidia logo Nemotron 3 Nano 30B A3B Open weights 262K $0.05 $0.20

Models near this price

The models nearest this one by input rate, five each way, at each one's cheapest listed provider.

Frequently Asked Questions

How much does Nemotron 3 Ultra 550B A55B cost?

Nemotron 3 Ultra 550B A55B costs $0.60 per 1M input tokens and $3.60 per 1M output tokens via Together. Nvidia publishes the weights, so it can also be run on your own hardware.

What does Nemotron 3 Ultra 550B A55B cost for 10M input and 2M output tokens?

At Together's rates, 10M input tokens and 2M output tokens cost about $13.20: $6.00 for input and $7.20 for output. Input is prompts and context, output is what the model writes back; a workload that generates more than it reads shifts the cost toward the output rate of $3.60 per 1M tokens.

Which providers offer Nemotron 3 Ultra 550B A55B?

2 providers list Nemotron 3 Ultra 550B A55B: Together ($0.60 in / $3.60 out) and Nvidia (open weights, no hosted price).

What is Nemotron 3 Ultra 550B A55B's context window?

Nemotron 3 Ultra 550B A55B accepts up to 1M tokens of input per request. The context window is the prompt plus any documents, conversation history and tool results sent with it; every token in it is billed at the input rate.

What is Nemotron 3 Ultra 550B A55B's knowledge cutoff?

Nemotron 3 Ultra 550B A55B's knowledge cutoff is September 2025: its training data runs up to that month and it has no built-in knowledge of later events.

What inputs and outputs does Nemotron 3 Ultra 550B A55B support?

Nemotron 3 Ultra 550B A55B accepts text as input and produces text. Its listed capabilities are function calling and structured output.

Can I self-host Nemotron 3 Ultra 550B A55B?

Yes. Nvidia publishes Nemotron 3 Ultra 550B A55B's weights under the OpenMDW-1.1 license. At 4-bit it needs about 330 GB of GPU memory; the cheapest rental that holds it is 8x RTX A6000 at $4.48 per hour, about $3,226 a month. $4.48 an hour (8x RTX A6000) buys about 4.1M tokens an hour at the cheapest hosted rate we track ($0.60 in / $3.60 out per 1M tokens on Together, 5:1 input to output); below that volume the API is cheaper, before idle time. The estimate above prices 8-bit and BF16 too.

How does Nemotron 3 Ultra 550B A55B compare with Nemotron 3 Super 120B A12B?

Nemotron 3 Ultra 550B A55B costs $0.60 per 1M input tokens against Nemotron 3 Super 120B A12B's $0.09, 6.7x more expensive ($3.60 vs $0.50 per 1M output tokens). Both have a 1M-token context window.

What are cheaper alternatives to Nemotron 3 Ultra 550B A55B?

Models from other creators with a lower input rate and at least Nemotron 3 Ultra 550B A55B's 1M-token context window: Gemini 3 Flash Preview at $0.50 per 1M input tokens (1M context), Qwen3.6-Plus at $0.50 per 1M input tokens (1M context) and DeepSeek V4 Pro at $0.45 per 1M input tokens (1M context). Rates are the cheapest listed provider for each; whether the quality holds for a given task is a separate question.

Cheaper alternatives to Nemotron 3 Ultra 550B A55B