Mistral logo Retired Jul 2026 Open weights · Apache 2.0

Magistral Small

A 24B open-weight reasoning model from Mistral AI under Apache 2.0, released June 2025, with traceable chain-of-thought.

Key Specifications

Context window
128K tokens
Max output
33K tokens
Released
Parameters
24B
Inputs
Text, image
Capabilities Show details
Reasoning Thinks before it answers, always on or as a switchable mode. Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON.

Hosted API pricing

Provider Input /1M tokens Output /1M tokens Cost at 10M in + 2M out
Mistral logo Mistral Creator Open weights, no hosted price View

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

Magistral Small needs about 20 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 2× RTX 3060 at about $86 a month.

Single node ≤ 8
Precision Memory Cheapest fit (32K context) Cost /mo
4-bitINT4 / FP4
20 GB
$86
8-bitFP8 / INT8
29 GB
$158
16-bitFP16 / BF16
55 GB
$317

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. No guarantee of runtime support, usable performance, or that a matching quantized build exists. How we estimate costs.

Frequently Asked Questions

What is Magistral Small good for?

Self-hosted reasoning with a traceable chain of thought, taking text and images. Apache 2.0 weights, small enough for a single consumer GPU.

When is Magistral Small not a good fit?

Hosted use: Mistral retired it from its API in July 2026, which leaves self-hosting. Its maximum output limits long reasoning traces in one response.

What is the cheapest way to run Magistral Small?

No provider we track hosts it; self-hosting on the cheapest rental that fits (2x RTX 3060, $0.12 per hour) costs about $86 a month.

Can I self-host Magistral Small?

Yes. The weights are Apache 2.0 licensed. At 4-bit it needs about 20 GB of GPU memory, which starts at roughly $86 a month on the cheapest rental that fits.

More from Mistral

Model Context Input /1M Output /1M
Mistral Ministral 3 14B Open weights 256K $0.20 $0.20
Mistral Ministral 3 3B Open weights 256K $0.10 $0.10
Mistral Ministral 3 8B Open weights 256K $0.15 $0.15
Mistral Mistral Large 3 Open weights 256K $0.50 $1.50
Mistral Mistral Medium 3.5 Open weights 256K $1.50 $7.50
Mistral Mistral Small 4 Open weights 256K $0.15 $0.60
Mistral Codestral 128K $0.30 $0.90
Mistral Devstral Medium 128K