Z.AI logo

GLM-5

Z.AI's fifth-generation flagship, a 744B-parameter MoE activating 40B per token with DeepSeek Sparse Attention, aimed at agentic engineering and long-horizon work.

Params
744B, 40B active
Context
200K tokens
Max output
131K
Released
License
MIT

Hosted API price

List price via Z.AI · per 1M tokens

USD
Input
$1.00 /1M
Output
$3.20 /1M
Cheapest
$0.60 / $1.60 via Geodd
Hosted by
Geodd Z.AI DigitalOcean Lyceum Novita 6 providers

Hosted API pricing

How we price

Order

Providers are listed cheapest input rate first, then by output rate, a tie going to the creator's own listing. A provider that lists the model without a published rate sits at the end.

Prices

Base-tier, on-demand rates in USD per 1M tokens, converted at ECB reference rates where a provider publishes in another currency. Cached-input and batch rates show only where a provider publishes them. "Cheapest" marks the one provider with the lowest input and output pair; a tie wears no badge.

Cost

Input rate times the input tokens plus output rate times the output tokens, at the monthly volume set above the table. Cached-input, batch, long-context and reasoning-token billing are not modelled.

Transparency and funding

  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.

Every provider serving GLM-5, cheapest input first. Set your monthly volume to see what each would bill.

Every provider serving GLM-5, with its price per 1M input and output tokens and what the volume above would cost
Provider Input /1M Output /1M Cached input /1M Cost at 10M in + 2M out Link
Geodd logo Geodd Cheapest $0.60 $1.60 $9.20 Visit website
Z.AI logo Z.AI Creator $1.00 $3.20 $0.20 $16.40 Visit website
DigitalOcean logo DigitalOcean $1.00 $3.20 $0.20 $16.40 Visit website
Lyceum logo Lyceum $1.00 $3.20 $16.40 Visit website
Novita logo Novita $1.00 $3.20 $0.20 $16.40 Visit website
Microsoft Azure logo Azure $1.10 $3.52 $0.22 $18.04 Visit website

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

GLM-5 needs about 414 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 8× A100 at about $10,080 a month.

That costs the same as roughly 7B tokens a month on Z.AI's API. Self-hosting is more expensive below that volume.

Single node ≤ 8
What it costs to self-host GLM-5 at each precision: the memory it needs, the cheapest rental that holds it, that rental's month and the API volume it pays off at
Precision Memory Cheapest fit (32K context) Cost /mo Break-even vs API
4-bit INT4 / FP4 414 GB Nvidia logo 8× A100 $10,080 7B tokens /mo
8-bit FP8 / INT8 747 GB Nvidia logo 4× B300 $22,666 17B tokens /mo
16-bit FP16 / BF16 1,493 GB Nvidia logo 8× B300 $45,331 33B tokens /mo

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. Break-even assumes a 5:1 input-to-output ratio. No guarantee of runtime support, usable performance, or that a matching quantized build exists. Pricing methodology.

Similarly priced models

The models nearest GLM-5 by blended rate, each at its own cheapest provider.

Models priced nearest GLM-5 per 1M tokens, each at its cheapest provider
Model Blended /1M Input /1M Output /1M Context Cutoff vs GLM-5
Google Cloud Gemini 2.5 Flash Google Cloud $0.6667 $0.30 $2.50 1M Jan 2025 −13%
Google Cloud Gemini 3.5 Flash-Lite Google Cloud $0.6667 $0.30 $2.50 1M Mar 2026 −13%
OpenAI GPT-3.5 Turbo OpenAI $0.6667 $0.50 $1.50 16K Sep 2021 −13%
Mistral Mistral Large 3 Mistral $0.6667 $0.50 $1.50 256K −13%
Alibaba Cloud Qwen3.5-Plus Alibaba Cloud $0.7333 $0.40 $2.40 1M −4%
Z.AI GLM-5 This model Z.AI $0.7667 $0.60 $1.60 200K
DeepSeek DeepSeek R1 Distill Llama 70B DeepSeek $0.80 $0.80 $0.80 131K +4%
Z.AI GLM-4.7 Z.AI $0.8667 $0.60 $2.20 205K +13%
Google Cloud Gemini 3 Flash Preview Google Cloud $0.9167 $0.50 $3.00 1M Jan 2025 +20%
Alibaba Cloud Qwen3.6-Plus Alibaba Cloud $0.9167 $0.50 $3.00 1M +20%
Z.AI GLM-5.2 Z.AI $0.95 $0.70 $2.20 1M +24%

USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.

Capabilities

What GLM-5 accepts and can do, as published by Z.AI.

  • Text Accepts and generates natural-language text.
  • Reasoning Thinks before it answers, always on or as a switchable mode.
  • Function calling Connect to external tools, APIs, and systems.
  • Structured output Return responses in structured formats like JSON.

Common questions

What is GLM-5 good for?

Complex system engineering and long-range agent work. MIT weights allow commercial self-hosting, on a multi-GPU node.

When is GLM-5 not a good fit?

Text only, and no built-in web search, though a host or your own tool loop can add one. GLM-5.1 has since replaced it, and self-hosting takes a large multi-GPU node.

What is the cheapest way to run GLM-5?

Hosted, unless you push serious volume. Z.AI charges $1.00 in / $3.20 out per 1M tokens. The cheapest rental that fits is 8x A100 at $10,080 a month, which costs the same as about 7B tokens a month on that API.

How does GLM-5 compare with GLM-5.1?

Z.AI says GLM-5.1 codes much better than GLM-5 and, unlike earlier models that run out of ideas early, keeps improving its answer over many rounds of tool calls on long agent tasks. Same text input. GLM-5 lists at 29% less per 1M input tokens and 27% less per 1M output tokens than GLM-5.1. It came out 2 months earlier and has a smaller context (200K against 205K tokens).

Can I self-host GLM-5?

Yes. The weights are MIT licensed. At 4-bit it needs about 414 GB of GPU memory, which starts at roughly $10,080 a month on the cheapest rental that fits.

More from Z.AI

  • GLM-4.7 Open weights 205K context
    From $0.60 / $2.20 /1M in / out at the cheapest provider
  • GLM-5.2 Open weights 1M context
    From $0.70 / $2.20 /1M in / out at the cheapest provider
  • GLM-5.1 Open weights 205K context
    From $0.825 / $3.301 /1M in / out at the cheapest provider
  • GLM-5-Turbo 205K context
    From $1.20 / $4.00 /1M in / out at the cheapest provider
  • GLM-5V-Turbo 205K context
    From $1.20 / $4.00 /1M in / out at the cheapest provider
  • GLM-5.3 1M context
    From $1.40 / $4.40 /1M in / out at the cheapest provider
  • GLM-5.3-Flash Open weights 1M context
    From $0.075 / $0.25 /1M in / out at the cheapest provider