MiniMax logo

MiniMax-M3

MiniMax's top model, succeeding M2.7: a 428B open-weight MoE activating roughly 23B per token across 256 experts with sparse attention, trained natively on text, image and video.

Cheapest of 5 providers, via MiniMax
Input
$0.30 / 1M tokens
Output
$1.20 / 1M tokens

About $5.40 for 10M input and 2M output tokens. Estimate yours

MiniMax GPUhub Novita Together Lyceum 5 providers

Key Specifications

Context window
1M tokens
Inputs
Text, Image, Video Outputs: Text

MiniMax-M3 pricing by provider

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
MiniMax logo MiniMax Creator Cheapest $0.30 $1.20 $5.40 View
GPUhub logo GPUhub $0.30 $1.20 $5.40 View
Novita logo Novita $0.30 $1.20 $5.40 View
Together logo Together $0.30 $1.20 $5.40 View
Lyceum logo Lyceum $0.40 $2.00 $8.00 View

Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning. LLM rates are base-tier, on-demand prices per 1M tokens; cached-input, batch and long-context tiers are not included.

Compare every model at this volume in the LLM cost calculator.

Capabilities

Function calling

Function calling

Connect to external tools, APIs, and systems.

Structured output

Structured output

Return responses in structured formats like JSON.

More from MiniMax

Model Context Input / 1M Output / 1M
MiniMax logo MiniMax-M2.5 205K $0.30 $1.20
MiniMax logo MiniMax-M2.7 205K $0.30 $1.20

Models near this price

The models nearest this one by input rate, five each way, at each one's cheapest listed provider.

Frequently Asked Questions

How much does MiniMax-M3 cost?

MiniMax-M3 costs $0.30 per 1M input tokens and $1.20 per 1M output tokens via MiniMax, the cheapest of 5 providers listing it. The highest rate is Lyceum at $0.40 in / $2.00 out.

What does MiniMax-M3 cost for 10M input and 2M output tokens?

At MiniMax's rates, 10M input tokens and 2M output tokens cost about $5.40: $3.00 for input and $2.40 for output. Input is prompts and context, output is what the model writes back; a workload that generates more than it reads shifts the cost toward the output rate of $1.20 per 1M tokens.

Which providers offer MiniMax-M3?

5 providers list MiniMax-M3: MiniMax ($0.30 in / $1.20 out), GPUhub ($0.30 in / $1.20 out), Novita ($0.30 in / $1.20 out), Together ($0.30 in / $1.20 out) and Lyceum ($0.40 in / $2.00 out). Rates are per 1M tokens in USD, cheapest input rate first.

What is MiniMax-M3's context window?

MiniMax-M3 accepts up to 1M tokens of input per request. The context window is the prompt plus any documents, conversation history and tool results sent with it; every token in it is billed at the input rate.

What inputs and outputs does MiniMax-M3 support?

MiniMax-M3 accepts text, images and video as input and produces text. Its listed capabilities are function calling and structured output.

How does MiniMax-M3 compare with MiniMax-M2.5?

MiniMax-M3 and MiniMax-M2.5 both cost $0.30 per 1M input tokens ($1.20 vs $1.20 per 1M output tokens). The context window is 1M tokens against 205K.

What are cheaper alternatives to MiniMax-M3?

Models from other creators with a lower input rate and at least MiniMax-M3's 1M-token context window: Gemini 3.1 Flash-Lite at $0.25 per 1M input tokens (1M context), Llama 4 Maverick at $0.25 per 1M input tokens (1M context) and Qwen3.6-Flash at $0.25 per 1M input tokens (1M context). Rates are the cheapest listed provider for each; whether the quality holds for a given task is a separate question.

Cheaper alternatives to MiniMax-M3