Z.AI logo

GLM-5.3-Flash

Open weights · MIT

Released August 2026, Z.AI's natively multimodal 320B-parameter MoE (18B active) under MIT license, with hybrid sparse and linear attention, built for coding and long-horizon agent tasks.

Cheapest of 2 providers, via Z.AI
Input
$0.15 / 1M tokens
Output
$0.50 / 1M tokens

About $2.50 for 10M input and 2M output tokens. Estimate yours

Z.AI Novita 2 providers

Key Specifications

Context window
1M tokens
Max output
131K tokens
Inputs
Text, Image, Video Outputs: Text

GLM-5.3-Flash pricing by provider

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
Z.AI logo Z.AI Creator Cheapest $0.15 $0.50 $2.50 View
Novita logo Novita $0.15 $0.50 $2.50 View

Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning. LLM rates are base-tier, on-demand prices per 1M tokens; cached-input, batch and long-context tiers are not included.

Compare every model at this volume in the LLM cost calculator.

Capabilities

Function calling

Function calling

Connect to external tools, APIs, and systems.

Structured output

Structured output

Return responses in structured formats like JSON.

Estimated cost to self-host GLM-5.3-Flash

GLM-5.3-Flash has 320B parameters, about 18B active per token. At 4-bit it needs about 192 GB of GPU memory, at 8-bit 384 GB and at BF16 768 GB, counting 20% on top of the weights for KV cache and runtime overhead.

Precision Memory needed Cheapest rentals that fit Per month, 24/7
4-bit 192 GB Amd logo MI300X ($2.90/hr, tight) Amd logo MI325X ($3.07/hr) $2,088
8-bit 384 GB Nvidia logo 8× RTX A6000 ($4.48/hr, tight) Amd logo 2× MI300X ($5.80/hr, tight) $3,226
BF16 768 GB Amd logo 4× MI300X ($11.60/hr, tight) Amd logo 4× MI325X ($12.28/hr) $8,352

$2.90 an hour (MI300X) buys about 14M tokens an hour at the cheapest hosted rate we track ($0.15 in / $0.50 out per 1M tokens on Z.AI, 5:1 input to output); below that volume the API is cheaper, before idle time. Compare hosted costs.

Memory is parameters × bytes per weight at each precision, plus 20% for KV cache and runtime overhead. Rentals are the cheapest cards that hold it, at provider-weighted median on-demand prices for the week of August 17, 2026, consumer cards included, in nodes of up to 8 GPUs; a fit with under 15% headroom is marked tight. A month is 720 hours. See the best-value GPUs guide for the same table across models, and cloud GPU pricing for every card.

More from Z.AI

Model Context Input / 1M Output / 1M
Z.AI logo GLM-4.7 Open weights 205K $0.60 $2.20
Z.AI logo GLM-5 Open weights 200K $0.60 $1.60
Z.AI logo GLM-5.2 Open weights 1M $0.90 $2.80
Z.AI logo GLM-5-Turbo 205K $1.20 $4.00
Z.AI logo GLM-5V-Turbo 205K $1.20 $4.00
Z.AI logo GLM-5.1 Open weights 205K $1.38 $4.40
Z.AI logo GLM-5.3 1M $1.40 $4.40

Models near this price

The models nearest this one by input rate, five each way, at each one's cheapest listed provider.

Frequently Asked Questions

How much does GLM-5.3-Flash cost?

GLM-5.3-Flash costs $0.15 per 1M input tokens and $0.50 per 1M output tokens via Z.AI; all 2 providers listing it charge the same rate. Z.AI publishes the weights, so it can also be run on your own hardware.

What does GLM-5.3-Flash cost for 10M input and 2M output tokens?

At Z.AI's rates, 10M input tokens and 2M output tokens cost about $2.50: $1.50 for input and $1.00 for output. Input is prompts and context, output is what the model writes back; a workload that generates more than it reads shifts the cost toward the output rate of $0.50 per 1M tokens.

Which providers offer GLM-5.3-Flash?

2 providers list GLM-5.3-Flash: Z.AI ($0.15 in / $0.50 out) and Novita ($0.15 in / $0.50 out). Rates are per 1M tokens in USD, cheapest input rate first.

What is GLM-5.3-Flash's context window?

GLM-5.3-Flash accepts up to 1M tokens of input per request and returns up to 131K tokens per response. The context window is the prompt plus any documents, conversation history and tool results sent with it; every token in it is billed at the input rate.

What inputs and outputs does GLM-5.3-Flash support?

GLM-5.3-Flash accepts text, images and video as input and produces text. Its listed capabilities are function calling and structured output.

Can I self-host GLM-5.3-Flash?

Yes. Z.AI publishes GLM-5.3-Flash's weights under the MIT license. At 4-bit it needs about 192 GB of GPU memory; the cheapest rental that holds it is MI300X at $2.90 per hour, about $2,088 a month. $2.90 an hour (MI300X) buys about 14M tokens an hour at the cheapest hosted rate we track ($0.15 in / $0.50 out per 1M tokens on Z.AI, 5:1 input to output); below that volume the API is cheaper, before idle time. The estimate above prices 8-bit and BF16 too.

How does GLM-5.3-Flash compare with GLM-4.7?

GLM-5.3-Flash costs $0.15 per 1M input tokens against GLM-4.7's $0.60, 4x cheaper ($0.50 vs $2.20 per 1M output tokens). The context window is 1M tokens against 205K.

What are cheaper alternatives to GLM-5.3-Flash?

Models from other creators with a lower input rate and at least GLM-5.3-Flash's 1M-token context window: Gemini 2.5 Flash-Lite at $0.10 per 1M input tokens (1M context). Rates are the cheapest listed provider for each; whether the quality holds for a given task is a separate question.

Cheaper alternatives to GLM-5.3-Flash