Z.AI logo

GLM-5.3-Flash

Z.AI's natively multimodal 320B-parameter MoE (18B active) under MIT license, with hybrid sparse and linear attention, built for coding and long-horizon agent tasks.

Params
320B, 18B active
Context
1M tokens
Max output
131K
Released
License
MIT

Hosted API price

List price via Z.AI · per 1M tokens

USD
Input
$0.08 /1M
Output
$0.25 /1M
Hosted by
Z.AI EmpirioLabs AI Novita Cloudflare DigitalOcean 7 providers

Hosted API pricing

How we price

Order

Providers are listed cheapest input rate first, then by output rate, a tie going to the creator's own listing. A provider that lists the model without a published rate sits at the end.

Prices

Base-tier, on-demand rates in USD per 1M tokens, converted at ECB reference rates where a provider publishes in another currency. Cached-input and batch rates show only where a provider publishes them. "Cheapest" marks the one provider with the lowest input and output pair; a tie wears no badge.

Cost

Input rate times the input tokens plus output rate times the output tokens, at the monthly volume set above the table. Cached-input, batch, long-context and reasoning-token billing are not modelled.

Transparency and funding

  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.

Every provider serving GLM-5.3-Flash, cheapest input first. Set your monthly volume to see what each would bill.

Every provider serving GLM-5.3-Flash, with its price per 1M input and output tokens and what the volume above would cost
Provider Input /1M Output /1M Cached input /1M Cost at 10M in + 2M out Link
Z.AI logo Z.AI Creator $0.08 $0.25 $0.02 $1.30 Visit website
EmpirioLabs AI logo EmpirioLabs AI $0.08 $0.25 $1.30 Visit website
Novita logo Novita $0.08 $0.25 $0.02 $1.30 Visit website
Cloudflare logo Cloudflare $0.15 $0.50 $0.03 $2.50 Visit website
DigitalOcean logo DigitalOcean $0.15 $0.50 $0.03 $2.50 Visit website
Together AI logo Together AI $0.15 $0.50 $0.03 $2.50 Visit website
Lyceum logo Lyceum $0.20 $0.50 $3.00 Visit website

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

GLM-5.3-Flash needs about 178 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 8× A40 at about $2,822 a month.

That costs the same as roughly 26B tokens a month on Z.AI's API. Self-hosting is more expensive below that volume.

Single node ≤ 8
What it costs to self-host GLM-5.3-Flash at each precision: the memory it needs, the cheapest rental that holds it, that rental's month and the API volume it pays off at
Precision Memory Cheapest fit (32K context) Cost /mo Break-even vs API
4-bit INT4 / FP4 178 GB Nvidia logo 8× A40 $2,822 26B tokens /mo
8-bit FP8 / INT8 322 GB Nvidia logo 8× A40 $2,822 26B tokens /mo
16-bit FP16 / BF16 642 GB Nvidia logo 8× RTX PRO 6000 $12,730 118B tokens /mo

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. Break-even assumes a 5:1 input-to-output ratio. No guarantee of runtime support, usable performance, or that a matching quantized build exists. How we estimate costs.

Similarly priced models

The models nearest GLM-5.3-Flash by blended rate, each at its own cheapest provider.

Models priced nearest GLM-5.3-Flash per 1M tokens, each at its cheapest provider
Model Blended /1M Input /1M Output /1M Context Cutoff vs GLM-5.3-Flash
Meta Llama 3 8B Meta $0.0833 $0.05 $0.25 8K Mar 2023 −23%
Alibaba Cloud Qwen3-Coder-30B-A3B Alibaba Cloud $0.0917 $0.06 $0.25 262K −15%
Meta Llama 3.2 3B Meta $0.0983 $0.05 $0.34 128K Dec 2023 −9%
Alibaba Cloud Qwen3-30B-A3B FP8 Alibaba Cloud $0.0983 $0.05 $0.34 33K −9%
Mistral Ministral 3 3B Mistral $0.10 $0.10 $0.10 256K −8%
Z.AI GLM-5.3-Flash This model Z.AI $0.1083 $0.08 $0.25 1M
OpenAI GPT-5 Nano OpenAI $0.1083 $0.05 $0.40 400K May 2024 0%
Alibaba Cloud Qwen3.5-35B-A3B Alibaba Cloud $0.1267 $0.06 $0.46 262K +17%
Alibaba Cloud Qwen3.6-35B-A3B Alibaba Cloud $0.1283 $0.07 $0.42 262K +18%
Google Cloud Gemma 3 27B Instruct Google Cloud $0.1333 $0.10 $0.30 128K Aug 2024 +23%
Google Cloud Gemma 4 26B A4B Google Cloud $0.1333 $0.10 $0.30 256K Jan 2025 +23%

USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.

Capabilities

What GLM-5.3-Flash accepts and can do, as published by Z.AI.

  • Text Accepts and generates natural-language text.
  • Image Accepts images as input alongside text.
  • Video Accepts video as input.
  • Reasoning Thinks before it answers, always on or as a switchable mode.
  • Function calling Connect to external tools, APIs, and systems.
  • Structured output Return responses in structured formats like JSON.

Common questions

What is GLM-5.3-Flash good for?

Visual coding, office and document work, and long agent tasks, with images and video as input. MIT weights allow commercial self-hosting, at the lower-cost end of Z.AI's lineup.

When is GLM-5.3-Flash not a good fit?

No audio input, and no built-in web search, though a host or your own tool loop can add one. Self-hosting still takes a multi-GPU node: memory follows the full 320B parameters, not the 18B active.

What is the cheapest way to run GLM-5.3-Flash?

Hosted, unless you push serious volume. Z.AI charges $0.08 in / $0.25 out per 1M tokens. The cheapest rental that fits is 8x A40 at $2,822 a month, which costs the same as about 26B tokens a month on that API.

Can I self-host GLM-5.3-Flash?

Yes. The weights are MIT licensed. At 4-bit it needs about 178 GB of GPU memory, which starts at roughly $2,822 a month on the cheapest rental that fits.

More from Z.AI

  • GLM-5 Open weights 200K context
    From $0.60 / $1.60 /1M in / out at the cheapest provider
  • GLM-4.7 Open weights 205K context
    From $0.60 / $2.20 /1M in / out at the cheapest provider
  • GLM-5.2 Open weights 1M context
    From $0.70 / $2.20 /1M in / out at the cheapest provider
  • GLM-5.1 Open weights 205K context
    From $0.83 / $3.30 /1M in / out at the cheapest provider
  • GLM-5-Turbo 205K context
    From $1.20 / $4.00 /1M in / out at the cheapest provider
  • GLM-5V-Turbo 205K context
    From $1.20 / $4.00 /1M in / out at the cheapest provider
  • GLM-5.3 1M context
    From $1.40 / $4.40 /1M in / out at the cheapest provider