Z.AI logo Open weights · MIT

GLM-5

Z.AI's fifth-generation flagship, a 744B-parameter MoE activating 40B per token with DeepSeek Sparse Attention, aimed at agentic engineering and long-horizon work.

List price via Z.AI
Input
$1.00 / 1M tokens
Output
$3.20 / 1M tokens

Cheapest: $0.60 / $1.60 per 1M tokens via Geodd

Geodd Z.AI DigitalOcean Lyceum Novita 6 providers

Key Specifications

Context window
200K tokens
Max output
131K tokens
Released
Parameters
744B, 40B active
Inputs
Text
Capabilities Show details
Reasoning Thinks before it answers, always on or as a switchable mode. Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON.

Hosted API pricing

Provider Input / 1M tokens Output / 1M tokens Cached input / 1M Cost at 10M in + 2M out
Geodd logo Geodd Cheapest $0.60 $1.60 $9.20 View
Z.AI logo Z.AI Creator $1.00 $3.20 $0.20 $16.40 View
DigitalOcean logo DigitalOcean $1.00 $3.20 $0.20 $16.40 View
Lyceum logo Lyceum $1.00 $3.20 $16.40 View
Novita logo Novita $1.00 $3.20 $0.20 $16.40 View
Microsoft Azure logo Azure $1.10 $3.52 $0.22 $18.04 View

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

GLM-5 needs about 414 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 8× A100 at about $10,138 a month.

That costs the same as roughly 7B tokens a month on Z.AI's API. Self-hosting is more expensive below that volume.

Precision Cheapest, 32K context Cheapest, full 200K context
4-bitINT4 / FP4
Nvidia logo 8× A100 414 GB VRAM · $10,138/mo
Nvidia logo 8× A100 429 GB VRAM · $10,138/mo
8-bitFP8 / INT8
Nvidia logo 4× B300 747 GB VRAM · $22,666/mo
Nvidia logo 4× B300 755 GB VRAM · $22,666/mo
16-bitFP16 / BF16
Nvidia logo 8× B300 1,493 GB VRAM · $45,331/mo
Nvidia logo 8× B300 1,508 GB VRAM · $45,331/mo

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. Break-even assumes a 5:1 input-to-output ratio. No guarantee of runtime support, usable performance, or that a matching quantized build exists. How we estimate costs.

Similarly priced models

The models nearest GLM-5 by blended rate, each at its own cheapest provider.

Model Blended / 1M vs GLM-5
Google Cloud Gemini 2.5 Flash Google Cloud $0.6667 −13%
Google Cloud Gemini 3.5 Flash-Lite Google Cloud $0.6667 −13%
OpenAI GPT-3.5 Turbo OpenAI $0.6667 −13%
Mistral Mistral Large 3 Mistral $0.6667 −13%
Alibaba Cloud Qwen3.5-Plus Alibaba Cloud $0.7333 −4%
Z.AI GLM-5 This model Z.AI $0.7667
DeepSeek DeepSeek R1 Distill Llama 70B DeepSeek $0.80 +4%
Z.AI GLM-4.7 Z.AI $0.8667 +13%
Alibaba Cloud Qwen3.5-122B-A10B Alibaba Cloud $0.8667 +13%
Google Cloud Gemini 3 Flash Preview Google Cloud $0.9167 +20%
Alibaba Cloud Qwen3.6-Plus Alibaba Cloud $0.9167 +20%

Prices are USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.

Frequently Asked Questions

What is GLM-5 good for?

Complex system engineering and long-range agent work. MIT weights allow commercial self-hosting, on a multi-GPU node.

When is GLM-5 not a good fit?

Text only, and no built-in web search, though a host or your own tool loop can add one. GLM-5.1 has since replaced it, and self-hosting takes a large multi-GPU node.

What is the cheapest way to run GLM-5?

Hosted, unless you push serious volume. Z.AI charges $1.00 in / $3.20 out per 1M tokens. The cheapest rental that fits is 8x A100 at $10,138 a month, which costs the same as about 7B tokens a month on that API.

How does GLM-5 compare with GLM-5.1?

Z.AI says GLM-5.1 codes much better than GLM-5 and, unlike earlier models that run out of ideas early, keeps improving its answer over many rounds of tool calls on long agent tasks. Same text input. GLM-5 lists at 29% less per 1M input tokens and 27% less per 1M output tokens than GLM-5.1. It came out 2 months earlier and has a smaller context (200K against 205K tokens).

Can I self-host GLM-5?

Yes. The weights are MIT licensed. At 4-bit it needs about 414 GB of GPU memory, which starts at roughly $10,138 a month on the cheapest rental that fits.

More from Z.AI

Model Context Input / 1M Output / 1M
Z.AI GLM-4.7 Open weights 205K $0.60 $2.20
Z.AI GLM-5.2 Open weights 1M $0.70 $2.20
Z.AI GLM-5-Turbo 205K $1.20 $4.00
Z.AI GLM-5V-Turbo 205K $1.20 $4.00
Z.AI GLM-5.1 Open weights 205K $1.30 $4.30
Z.AI GLM-5.3 1M $1.40 $4.40
Z.AI GLM-5.3-Flash Open weights 1M $0.08 $0.25