Nvidia logo

Nemotron 3 Ultra 550B A55B

The largest of NVIDIA's Nemotron 3 line, a 550B LatentMoE hybrid activating 55B per token across Mamba-2, MoE and attention layers with Multi-Token Prediction, under the OpenMDW-1.1 license.

Params
550B, 55B active
Context
1M tokens
Cutoff
Released
License
OpenMDW-1.1

Hosted API price

Cheapest via Azure · per 1M tokens

USD
Input
$0.66 /1M
Output
$2.64 /1M
Hosted by
Azure DigitalOcean Nvidia 3 providers

Hosted API pricing

How we price

Order

Providers are listed cheapest input rate first, then by output rate, a tie going to the creator's own listing. A provider that lists the model without a published rate sits at the end.

Prices

Base-tier, on-demand rates in USD per 1M tokens, converted at ECB reference rates where a provider publishes in another currency. Cached-input and batch rates show only where a provider publishes them. "Cheapest" marks the one provider with the lowest input and output pair; a tie wears no badge.

Cost

Input rate times the input tokens plus output rate times the output tokens, at the monthly volume set above the table. Cached-input, batch, long-context and reasoning-token billing are not modelled.

Transparency and funding

  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.

Every provider serving Nemotron 3 Ultra 550B A55B, cheapest input first. Set your monthly volume to see what each would bill.

Every provider serving Nemotron 3 Ultra 550B A55B, with its price per 1M input and output tokens and what the volume above would cost
Provider Input /1M Output /1M Cached input /1M Cost at 10M in + 2M out Link
Microsoft Azure logo Azure Cheapest $0.66 $2.64 $0.13 $11.88 Visit website
DigitalOcean logo DigitalOcean $0.90 $1.70 $12.40 Visit website
Nvidia logo Nvidia Creator Open weights, no hosted price Visit website

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

Nemotron 3 Ultra 550B A55B needs about 313 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 8× A40 at about $2,822 a month.

That costs the same as roughly 3B tokens a month on Azure's API. Self-hosting is more expensive below that volume.

Single node ≤ 8
What it costs to self-host Nemotron 3 Ultra 550B A55B at each precision: the memory it needs, the cheapest rental that holds it, that rental's month and the API volume it pays off at
Precision Memory Cheapest fit (32K context) Cost /mo Break-even vs API
4-bit INT4 / FP4 313 GB Nvidia logo 8× A40 $2,822 3B tokens /mo
8-bit FP8 / INT8 560 GB Nvidia logo 8× A100 $10,138 10B tokens /mo
16-bit FP16 / BF16 1,110 GB Nvidia logo 8× B200 $36,000 36B tokens /mo

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. KV cache assumed at 250 KB per token, a conservative default. Break-even assumes a 5:1 input-to-output ratio. No guarantee of runtime support, usable performance, or that a matching quantized build exists. How we estimate costs.

Similarly priced models

The models nearest Nemotron 3 Ultra 550B A55B by blended rate, each at its own cheapest provider.

Models priced nearest Nemotron 3 Ultra 550B A55B per 1M tokens, each at its cheapest provider
Model Blended /1M Input /1M Output /1M Context Cutoff vs Nemotron 3 Ultra 550B A55B
DeepSeek DeepSeek R1 Distill Llama 70B DeepSeek $0.80 $0.80 $0.80 131K −19%
Z.AI GLM-4.7 Z.AI $0.8667 $0.60 $2.20 205K −12%
Google Cloud Gemini 3 Flash Preview Google Cloud $0.9167 $0.50 $3.00 1M Jan 2025 −7%
Alibaba Cloud Qwen3.6-Plus Alibaba Cloud $0.9167 $0.50 $3.00 1M −7%
Z.AI GLM-5.2 Z.AI $0.95 $0.70 $2.20 1M −4%
Nvidia Nemotron 3 Ultra 550B A55B This model Nvidia $0.99 $0.66 $2.64 1M Sep 2025
DeepSeek DeepSeek R1 DeepSeek $1.00 $0.70 $2.50 128K +1%
DeepSeek DeepSeek R1 Turbo DeepSeek $1.00 $0.70 $2.50 64K +1%
Meta Llama 3 70B Meta $1.00 $0.65 $2.75 8K Dec 2023 +1%
Alibaba Cloud Qwen3-VL-235B-A22B Thinking Alibaba Cloud $1.00 $0.40 $4.00 262K +1%
Moonshot AI Kimi K2.6 Moonshot AI $1.0333 $0.60 $3.20 262K +4%

USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.

Capabilities

What Nemotron 3 Ultra 550B A55B accepts and can do, as published by Nvidia.

  • Text Accepts and generates natural-language text.
  • Reasoning Thinks before it answers, always on or as a switchable mode.
  • Function calling Connect to external tools, APIs, and systems.
  • Structured output Return responses in structured formats like JSON.

Common questions

What is Nemotron 3 Ultra 550B A55B good for?

Complex multi-step agents, analysis and RAG over a very long context, at the heavy end of the Nemotron 3 line. Open weights under OpenMDW-1.1, on a multi-GPU node.

When is Nemotron 3 Ultra 550B A55B not a good fit?

Cost-sensitive or modest self-hosted use: it sits at the top of NVIDIA's price range and takes a multi-GPU node. Text only, with no published maximum output.

What is the cheapest way to run Nemotron 3 Ultra 550B A55B?

Hosted, unless you push serious volume. Azure charges $0.66 in / $2.64 out per 1M tokens. The cheapest rental that fits is 8x A40 at $2,822 a month, which costs the same as about 3B tokens a month on that API.

Can I self-host Nemotron 3 Ultra 550B A55B?

Yes. The weights are OpenMDW-1.1 licensed. At 4-bit it needs about 313 GB of GPU memory, which starts at roughly $2,822 a month on the cheapest rental that fits.

More from Nvidia

  • Nemotron 3 Super 120B A12B Open weights 1M context · cutoff Jun 2025
    From $0.09 / $0.50 /1M in / out at the cheapest provider
  • Nemotron 3 Nano 30B A3B Open weights 262K context · cutoff Jun 2025
    From $0.05 / $0.20 /1M in / out at the cheapest provider