Mistral logo

Ministral 3 14B

Open weights · Apache 2.0

The largest of the Ministral 3 family, a 14B dense model (13.5B language model plus a 0.4B vision encoder) under Apache 2.0, built for local and edge deployment.

Price via Mistral
Input
$0.20 / 1M tokens
Output
$0.20 / 1M tokens

About $2.40 for 10M input and 2M output tokens. Estimate yours

Mistral 1 provider

Key Specifications

Context window
256K tokens
Inputs
Text, Image Outputs: Text

Ministral 3 14B pricing by provider

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
Mistral logo Mistral Creator $0.20 $0.20 $2.40 View

Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning. LLM rates are base-tier, on-demand prices per 1M tokens; cached-input, batch and long-context tiers are not included.

Compare every model at this volume in the LLM cost calculator.

Capabilities

Function calling

Function calling

Connect to external tools, APIs, and systems.

Structured output

Structured output

Return responses in structured formats like JSON.

Estimated cost to self-host Ministral 3 14B

Ministral 3 14B has 13.9B parameters. At 4-bit it needs about 8 GB of GPU memory, at 8-bit 17 GB and at BF16 33 GB, counting 20% on top of the weights for KV cache and runtime overhead.

Precision Memory needed Cheapest rentals that fit Per month, 24/7
4-bit 8 GB Nvidia logo RTX 3060 ($0.07/hr) Nvidia logo RTX 2080 Ti ($0.08/hr) $50
8-bit 17 GB Nvidia logo RTX 3090 ($0.13/hr) Nvidia logo RTX 3090 Ti ($0.19/hr) $94
BF16 33 GB Nvidia logo RTX A6000 ($0.56/hr) Nvidia logo A40 ($0.74/hr) $403

$0.07 an hour (RTX 3060) buys about 350K tokens an hour at the cheapest hosted rate we track ($0.20 in / $0.20 out per 1M tokens on Mistral, 5:1 input to output); below that volume the API is cheaper, before idle time. Compare hosted costs.

Memory is parameters × bytes per weight at each precision, plus 20% for KV cache and runtime overhead. Rentals are the cheapest cards that hold it, at provider-weighted median on-demand prices for the week of August 17, 2026, consumer cards included, in nodes of up to 8 GPUs; a fit with under 15% headroom is marked tight. A month is 720 hours. See the best-value GPUs guide for the same table across models, and cloud GPU pricing for every card.

More from Mistral

Model Context Input / 1M Output / 1M
Mistral logo Ministral 3 8B Open weights 256K $0.15 $0.15
Mistral logo Mistral Small 4 Open weights 256K $0.15 $0.60
Mistral logo Codestral 128K $0.30 $0.90
Mistral logo Ministral 3 3B Open weights 256K $0.10 $0.10
Mistral logo Mistral Large 3 Open weights 256K $0.50 $1.50
Mistral logo Mistral NeMo Open weights 128K $0.04 $0.17
Mistral logo Mistral Medium 3.5 Open weights 256K $1.50 $7.50
Mistral logo Devstral Medium 128K

Models near this price

The models nearest this one by input rate, five each way, at each one's cheapest listed provider.

Frequently Asked Questions

How much does Ministral 3 14B cost?

Ministral 3 14B costs $0.20 per 1M input tokens and $0.20 per 1M output tokens via Mistral. Mistral publishes the weights, so it can also be run on your own hardware.

What does Ministral 3 14B cost for 10M input and 2M output tokens?

At Mistral's rates, 10M input tokens and 2M output tokens cost about $2.40: $2.00 for input and $0.40 for output. Input is prompts and context, output is what the model writes back; a workload that generates more than it reads shifts the cost toward the output rate of $0.20 per 1M tokens.

Which providers offer Ministral 3 14B?

One provider lists Ministral 3 14B: Mistral ($0.20 in / $0.20 out).

What is Ministral 3 14B's context window?

Ministral 3 14B accepts up to 256K tokens of input per request. The context window is the prompt plus any documents, conversation history and tool results sent with it; every token in it is billed at the input rate.

What inputs and outputs does Ministral 3 14B support?

Ministral 3 14B accepts text and images as input and produces text. Its listed capabilities are function calling and structured output.

Can I self-host Ministral 3 14B?

Yes. Mistral publishes Ministral 3 14B's weights under the Apache 2.0 license. At 4-bit it needs about 8 GB of GPU memory; the cheapest rental that holds it is RTX 3060 at $0.07 per hour, about $50 a month. $0.07 an hour (RTX 3060) buys about 350K tokens an hour at the cheapest hosted rate we track ($0.20 in / $0.20 out per 1M tokens on Mistral, 5:1 input to output); below that volume the API is cheaper, before idle time. The estimate above prices 8-bit and BF16 too.

How does Ministral 3 14B compare with Ministral 3 8B?

Ministral 3 14B costs $0.20 per 1M input tokens against Ministral 3 8B's $0.15, 1.3x more expensive ($0.20 vs $0.15 per 1M output tokens). Both have a 256K-token context window.

What are cheaper alternatives to Ministral 3 14B?

Models from other creators with a lower input rate and at least Ministral 3 14B's 256K-token context window: Llama 4 Scout at $0.17 per 1M input tokens (10M context), GLM-5.3-Flash at $0.15 per 1M input tokens (1M context) and Qwen3-Next-80B-A3B Instruct at $0.15 per 1M input tokens (262K context). Rates are the cheapest listed provider for each; whether the quality holds for a given task is a separate question.

Cheaper alternatives to Ministral 3 14B