DeepSeek V4 Pro
DeepSeek's flagship 1.6T-parameter MoE (49B active) with hybrid sparse attention, previewed April 2026 and replaced by the 0813 release in August 2026, with weights under the MIT license.
- Input
- $1.32 / 1M tokens
- Output
- $3.96 / 1M tokens
Cheapest: $0.45 / $0.89 per 1M tokens via GPUhub
Key Specifications
- Context window
- 1M tokens
- Max output
- 384K tokens
- Released
- Parameters
- 1.6T, 49B active
- Inputs
- Text
- Capabilities Show details
- Reasoning Thinks before it answers, always on or as a switchable mode. Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON. Web search Search the web for up-to-date information.
Hosted API pricing
| Provider | Input / 1M tokens | Output / 1M tokens | Cached input / 1M | Cost at 10M in + 2M out | |
|---|---|---|---|---|---|
|
|
$0.45 | $0.89 | $6.28 | View | |
|
|
$1.32 | $3.96 | $21.12 | View | |
|
|
$1.32 | $3.96 | $0.13 | $21.12 | View |
|
|
$1.48 | $3.40 | $21.60 | View | |
|
|
$1.60 | $3.20 | $22.40 | View | |
|
|
$1.75 | $3.50 | $24.50 | View |
Heads up: Prices are estimates using base-tier, on-demand rates per 1M tokens from published pages; cached, batch and long-context tiers are shown only where listed. Providers may serve a shorter context or a quantized build than the creator's release. Verify with the provider before provisioning. How we estimate costs.
Estimated cost to self-host
DeepSeek V4 Pro needs about 890 GB of GPU memory at 4-bit. The cheapest rental that fits is 4× B300 at about $22,723 a month.
That costs the same as roughly 13B tokens a month on DeepSeek's API. Self-hosting is more expensive below that volume.
| Precision | Cheapest, 1x concurrency | Cheapest, 8x concurrency |
|---|---|---|
| 4-bit |
|
|
| 8-bit |
|
|
| 16-bit |
Needs 3,210 GB
|
Needs 3,266 GB
|
Estimates, not quotes. Nvidia cards at median on-demand rates, 32K context per request. Break-even assumes 5:1 input to output. How we estimate costs.
Similarly priced models
The models nearest DeepSeek V4 Pro by blended rate, five each way, each at its own cheapest provider.
| Model | Blended / 1M | Input / 1M | Output / 1M | Context | Cutoff | vs DeepSeek V4 Pro |
|---|---|---|---|---|---|---|
|
|
$0.46 | $0.25 | $1.49 | 262K | −13% | |
|
|
$0.46 | $0.25 | $1.50 | 1M | Jan 2025 | −12% |
|
|
$0.46 | $0.25 | $1.50 | 1M | −12% | |
|
|
$0.48 | $0.32 | $1.28 | 1M | −8% | |
|
|
$0.50 | $0.30 | $1.50 | 262K | −4% | |
|
|
$0.52 | $0.45 | $0.89 | 1M | ||
|
|
$0.54 | $0.25 | $2.00 | 400K | May 2024 | +4% |
|
|
$0.54 | $0.35 | $1.50 | 131K | Jan 2026 | +4% |
|
|
$0.54 | $0.25 | $2.00 | 262K | +4% | |
|
|
$0.55 | $0.40 | $1.30 | 64K | +5% | |
|
|
$0.58 | $0.38 | $1.55 | 262K | +10% |
Prices are USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.
Frequently Asked Questions
What is DeepSeek V4 Pro good for?
DeepSeek's flagship for the widest range of text tasks, with MIT weights.
When is DeepSeek V4 Pro not a good fit?
Cost-sensitive or self-hosted use. It is the most expensive in DeepSeek's lineup and needs about 960 GB to self-host at 4-bit. Text only.
What is the cheapest way to run DeepSeek V4 Pro?
Hosted, unless you push serious volume. DeepSeek charges $1.32 in / $3.96 out per 1M tokens. The cheapest rental that fits is 4x B300 at $22,723 a month, which costs the same as about 13B tokens a month on that API.
Can I self-host DeepSeek V4 Pro?
Yes. The weights are MIT licensed. At 4-bit it needs about 890 GB of GPU memory, which starts at roughly $22,723 a month on the cheapest rental that fits. See the table above for 8-bit and 16-bit.
More from DeepSeek
| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
|
|
64K | $0.40 | $1.30 |
|
|
1M | $0.44 | $1.32 |
|
|
128K | $0.27 | $1.12 |
|
|
131K | $0.27 | $1.00 |
|
|
131K | $0.27 | $1.00 |
|
|
131K | $0.80 | $0.80 |
|
|
131K | $0.27 | $0.40 |
|
|
64K | $0.70 | $2.50 |