DeepSeek logo Open weights · MIT

DeepSeek V4 Flash

DeepSeek's lightweight 284B MoE (13B active) under the MIT license, previewed April 2026 and released as DeepSeek-V4-Flash-0731 in July 2026 with gains in agentic and coding work.

List price via DeepSeek
Input
$0.44 / 1M tokens
Output
$1.32 / 1M tokens

Cheapest: $0.14 / $0.28 per 1M tokens via Novita

Novita Together AI Geodd GPUhub Lyceum 6 providers

Key Specifications

Context window
1M tokens
Max output
384K tokens
Released
Parameters
284B, 13B active
Inputs
Text
Capabilities Show details
Reasoning Thinks before it answers, always on or as a switchable mode. Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON. Web search Search the web for up-to-date information.

Hosted API pricing

Provider Input / 1M tokens Output / 1M tokens Cached input / 1M Cost at 10M in + 2M out
Novita logo Novita $0.14 $0.28 $1.96 View
Together AI logo Together AI $0.14 $0.28 $0.03 $1.96 View
Geodd logo Geodd $0.14 $0.30 $2.00 View
GPUhub logo GPUhub $0.15 $0.30 $2.10 View
Lyceum logo Lyceum $0.15 $0.30 $2.10 View
DeepSeek logo DeepSeek Creator $0.44 $1.32 $7.04 View

Heads up: Prices are estimates using base-tier, on-demand rates per 1M tokens from published pages; cached, batch and long-context tiers are shown only where listed. Providers may serve a shorter context or a quantized build than the creator's release. Verify with the provider before provisioning. How we estimate costs.

Estimated cost to self-host

DeepSeek V4 Flash needs about 166 GB of GPU memory at 4-bit. The cheapest rental that fits is 8× RTX 3090 at about $749 a month.

That costs the same as roughly 1B tokens a month on DeepSeek's API. Self-hosting is more expensive below that volume.

Precision Cheapest, 1x concurrency Cheapest, 8x concurrency
4-bit
Nvidia logo 8× RTX 3090 166 GB · $749/mo
Nvidia logo 8× RTX 5090 222 GB · $2,650/mo
8-bit
Nvidia logo 8× RTX A6000 294 GB · $3,226/mo
Nvidia logo 8× A100 350 GB · $10,138/mo
16-bit
Nvidia logo 8× RTX PRO 6000 578 GB · $12,614/mo
Nvidia logo 8× RTX PRO 6000 634 GB · $12,614/mo

Estimates, not quotes. Nvidia cards at median on-demand rates, 32K context per request. Break-even assumes 5:1 input to output. How we estimate costs.

Similarly priced models

The models nearest DeepSeek V4 Flash by blended rate, five each way, each at its own cheapest provider.

Model Blended / 1M vs DeepSeek V4 Flash
Google Cloud Gemini 2.5 Flash-Lite Google Cloud $0.15 −8%
OpenAI GPT-4.1 Nano OpenAI $0.15 −8%
Mistral Ministral 3 8B Mistral $0.15 −8%
Alibaba Cloud Qwen3.5-Flash Alibaba Cloud $0.15 −8%
Nvidia Nemotron 3 Super 120B A12B Nvidia $0.16 −3%
DeepSeek DeepSeek V4 Flash This model DeepSeek $0.16
Google Cloud Gemma 4 31B Google Cloud $0.17 +4%
Alibaba Cloud Qwen3-235B-A22B Instruct (2507) Alibaba Cloud $0.17 +5%
Google Cloud Gemma 4 26B A4B Google Cloud $0.18 +7%
Meta Llama 3.3 70B Meta $0.18 +12%
Mistral Ministral 3 14B Mistral $0.20 +22%

Prices are USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.

Frequently Asked Questions

What is DeepSeek V4 Flash good for?

Long-context agent and coding work. It is the cheapest of DeepSeek's chat models, and its MIT weights need about 170 GB at 4-bit to self-host, far less than V4 Pro.

When is DeepSeek V4 Flash not a good fit?

Image input, which DeepSeek offers only through the separate V4 Flash Vision Exp variant. DeepSeek's own API charges the list rate at peak hours, so the price depends on time of day.

What is the cheapest way to run DeepSeek V4 Flash?

Hosted, unless you push serious volume. DeepSeek charges $0.44 in / $1.32 out per 1M tokens. The cheapest rental that fits is 8x RTX 3090 at $749 a month, which costs the same as about 1B tokens a month on that API.

Can I self-host DeepSeek V4 Flash?

Yes. The weights are MIT licensed. At 4-bit it needs about 166 GB of GPU memory, which starts at roughly $749 a month on the cheapest rental that fits. See the table above for 8-bit and 16-bit.

More from DeepSeek

Model Context Input / 1M Output / 1M
DeepSeek DeepSeek V3.2 Open weights 131K $0.27 $0.40
DeepSeek DeepSeek V3.1 Open weights 131K $0.27 $1.00
DeepSeek DeepSeek V3.1 Terminus Open weights 131K $0.27 $1.00
DeepSeek DeepSeek V3 Open weights 128K $0.27 $1.12
DeepSeek DeepSeek V4 Pro Open weights 1M $0.45 $0.89
DeepSeek DeepSeek V3 Turbo Open weights 64K $0.40 $1.30
DeepSeek DeepSeek V4 Flash Vision Exp 1M $0.44 $1.32
DeepSeek DeepSeek R1 Distill Llama 70B Open weights 131K $0.80 $0.80