MiniMax-M2.5
Open weights · Modified MIT
A 229B-parameter open-weight MoE from MiniMax activating roughly 10B per token under a modified MIT license, the predecessor of MiniMax-M2.7, aimed at software engineering and browsing tasks.
- Input
- $0.30 / 1M tokens
- Output
- $1.20 / 1M tokens
About $5.40 for 10M input and 2M output tokens. Estimate yours
Key Specifications
- Context window
- 205K tokens
- Inputs
- Text Outputs: Text
MiniMax-M2.5 pricing by provider
| Provider | Input / 1M tokens | Output / 1M tokens | Cost at 10M in + 2M out | |
|---|---|---|---|---|
|
|
$0.30 | $1.20 | $5.40 | View |
|
|
$0.30 | $1.20 | $5.40 | View |
|
|
$0.30 | $1.20 | $5.40 | View |
Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning. LLM rates are base-tier, on-demand prices per 1M tokens; cached-input, batch and long-context tiers are not included.
Compare every model at this volume in the LLM cost calculator.
Capabilities
Function calling
Connect to external tools, APIs, and systems.
Structured output
Return responses in structured formats like JSON.
Estimated cost to self-host MiniMax-M2.5
MiniMax-M2.5 has 229B parameters, about 10B active per token. At 4-bit it needs about 137 GB of GPU memory, at 8-bit 275 GB and at BF16 550 GB, counting 20% on top of the weights for KV cache and runtime overhead.
| Precision | Memory needed | Cheapest rentals that fit | Per month, 24/7 |
|---|---|---|---|
| 4-bit | 137 GB |
|
$2,088 |
| 8-bit | 275 GB |
|
$5,681 |
| BF16 | 550 GB |
|
$8,352 |
$2.90 an hour (MI300X) buys about 6.4M tokens an hour at the cheapest hosted rate we track ($0.30 in / $1.20 out per 1M tokens on MiniMax, 5:1 input to output); below that volume the API is cheaper, before idle time. Compare hosted costs.
Memory is parameters × bytes per weight at each precision, plus 20% for KV cache and runtime overhead. Rentals are the cheapest cards that hold it, at provider-weighted median on-demand prices for the week of August 17, 2026, consumer cards included, in nodes of up to 8 GPUs; a fit with under 15% headroom is marked tight. A month is 720 hours. See the best-value GPUs guide for the same table across models, and cloud GPU pricing for every card.
More from MiniMax
| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
|
|
1M | $0.30 | $1.20 |
|
|
205K | $0.30 | $1.20 |
Models near this price
The models nearest this one by input rate, five each way, at each one's cheapest listed provider.
Frequently Asked Questions
How much does MiniMax-M2.5 cost?
MiniMax-M2.5 costs $0.30 per 1M input tokens and $1.20 per 1M output tokens via MiniMax; all 3 providers listing it charge the same rate. MiniMax publishes the weights, so it can also be run on your own hardware.
What does MiniMax-M2.5 cost for 10M input and 2M output tokens?
At MiniMax's rates, 10M input tokens and 2M output tokens cost about $5.40: $3.00 for input and $2.40 for output. Input is prompts and context, output is what the model writes back; a workload that generates more than it reads shifts the cost toward the output rate of $1.20 per 1M tokens.
Which providers offer MiniMax-M2.5?
3 providers list MiniMax-M2.5: MiniMax ($0.30 in / $1.20 out), Lyceum ($0.30 in / $1.20 out) and Novita ($0.30 in / $1.20 out). Rates are per 1M tokens in USD, cheapest input rate first.
What is MiniMax-M2.5's context window?
MiniMax-M2.5 accepts up to 205K tokens of input per request. The context window is the prompt plus any documents, conversation history and tool results sent with it; every token in it is billed at the input rate.
What inputs and outputs does MiniMax-M2.5 support?
MiniMax-M2.5 accepts text as input and produces text. Its listed capabilities are function calling and structured output.
Can I self-host MiniMax-M2.5?
Yes. MiniMax publishes MiniMax-M2.5's weights under the Modified MIT license. At 4-bit it needs about 137 GB of GPU memory; the cheapest rental that holds it is MI300X at $2.90 per hour, about $2,088 a month. $2.90 an hour (MI300X) buys about 6.4M tokens an hour at the cheapest hosted rate we track ($0.30 in / $1.20 out per 1M tokens on MiniMax, 5:1 input to output); below that volume the API is cheaper, before idle time. The estimate above prices 8-bit and BF16 too.
How does MiniMax-M2.5 compare with MiniMax-M3?
MiniMax-M2.5 and MiniMax-M3 both cost $0.30 per 1M input tokens ($1.20 vs $1.20 per 1M output tokens). The context window is 205K tokens against 1M.
What are cheaper alternatives to MiniMax-M2.5?
Models from other creators with a lower input rate and at least MiniMax-M2.5's 205K-token context window: Gemini 3.1 Flash-Lite at $0.25 per 1M input tokens (1M context), Llama 4 Maverick at $0.25 per 1M input tokens (1M context) and GPT-5 Mini at $0.25 per 1M input tokens (400K context). Rates are the cheapest listed provider for each; whether the quality holds for a given task is a separate question.
Cheaper alternatives to MiniMax-M2.5
Gemini 3.1 Flash-Lite
$0.25 per 1M input tokens and $1.50 per 1M output tokens via Google Cloud, 1M-token context, by Google Cloud.
Llama 4 Maverick
$0.25 per 1M input tokens and $0.95 per 1M output tokens via Replicate, 1M-token context, by Meta.
GPT-5 Mini
$0.25 per 1M input tokens and $2.00 per 1M output tokens via OpenAI, 400K-token context, by OpenAI.