Devstral Small
Mistral AI's 24B agentic coding model, updated in July 2025, with weights under Apache 2.0.
Key Specifications
- Context window
- 128K tokens
- Max output
- 33K tokens
- Released
- Parameters
- 24B
- Inputs
- Text
- Capabilities Show details
- Function calling Connect to external tools, APIs, and systems. Code execution Write and execute code in a sandboxed environment.
Hosted API pricing
| Provider | Input / 1M tokens | Output / 1M tokens | Cost at 10M in + 2M out | |
|---|---|---|---|---|
|
|
Open weights, no hosted price | View | ||
Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.
Estimated cost to self-host
Devstral Small needs about 20 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 2× RTX 3060 at about $86 a month.
| Precision | Memory | Cheapest fit (32K context) | Cost / mo |
|---|---|---|---|
| 4-bitINT4 / FP4 |
20 GB
|
$86
|
|
| 8-bitFP8 / INT8 |
29 GB
|
$158
|
|
| 16-bitFP16 / BF16 |
55 GB
|
$317
|
Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. No guarantee of runtime support, usable performance, or that a matching quantized build exists. How we estimate costs.
Frequently Asked Questions
What is Devstral Small good for?
Agent-style coding on your own hardware. Apache 2.0 weights, small enough for a single consumer GPU.
When is Devstral Small not a good fit?
Hosted use: Mistral retired it from its API in May 2026, which leaves self-hosting. Text only, and its maximum output caps very long single responses.
What is the cheapest way to run Devstral Small?
No provider we track hosts it; self-hosting on the cheapest rental that fits (2x RTX 3060, $0.12 per hour) costs about $86 a month.
Can I self-host Devstral Small?
Yes. The weights are Apache 2.0 licensed. At 4-bit it needs about 20 GB of GPU memory, which starts at roughly $86 a month on the cheapest rental that fits.
More from Mistral
| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
|
|
256K | $0.20 | $0.20 |
|
|
256K | $0.10 | $0.10 |
|
|
256K | $0.15 | $0.15 |
|
|
256K | $0.50 | $1.50 |
|
|
256K | $1.50 | $7.50 |
|
|
256K | $0.15 | $0.60 |
|
|
128K | $0.30 | $0.90 |
|
|
128K | ||