Model · stat sheet
GLM-5
Zhipu AI's flagship open model; strong bilingual reasoning and dependable agentic tool use.
What it is
GLM-5 is the flagship of Zhipu AI's open-weights line — the lab also trades as Z.ai internationally — and it is one of the small group of Chinese open models that ship weights, a cheap first-party API and listings on the Western gateways at the same time. On this board it appears on OpenRouter, OpenCode Zen and OpenCode Go, which is the same three-gateway reach as DeepSeek V4.
The GLM series has consistently been positioned around agentic and bilingual Chinese–English use. Verify the specific licence, parameter count and context window for GLM-5 in Z.ai's own documentation rather than from any leaderboard, including this one.
How the composite is built
Overall 82 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. GLM-5's profile is unusually even for an open model, with no axis below 74 and none above 86.
- Coding 85
- Seven behind DeepSeek V4's 92 and level with Mistral Large 3. Competent on ordinary work; not the open model to pick if code quality is the only criterion.
- Terminal 79
- The weakest axis. Sixteen behind Opus 5 and five behind V4. Long unattended shell sessions are not where this model is strongest.
- Reasoning 86
- The best of its six. Ahead of Qwen 4 Coder's 84 and level with Llama 5. Solid planning for the price.
- Tool use 80
- Worth reading carefully against the reputation. GLM is marketed on agentic tool use, and 80 is respectable — level with GPT-5.6 Luna and Kimi k3 — but it is sixteen behind Opus 5 and five behind Claude Haiku 4.5, which is a small model by any measure. Dependable, not exceptional.
- Context 85
- Mid-field, level with DeepSeek R2. Fine for module-scale work.
- Speed 74
- Level with DeepSeek V4 on hosted inference. As with any open model, self-hosting makes this number a property of your hardware rather than the model.
Where it fits
Cost-controlled agent work where you want open weights as an exit option: supervised coding loops, bilingual applications, internal tooling where the data cannot leave your infrastructure. It is a reasonable default for teams who want DeepSeek-class economics with a second vendor in the mix.
Limits
- No value rating on this board — the column reads Unrated because no Artificial Analysis index is published for this model, not because the price is bad. At $0.60/$1.92 the price is among the better ones here.
- Terminal 79 caps autonomy. Keep a human or a stricter model between GLM-5 and anything destructive.
- The agentic reputation runs ahead of the tool-use number. Test on your own tool schemas before believing the marketing or the leaderboard.
- First-party hosting is in China. For regulated workloads, use the open weights or a Western gateway rather than the direct endpoint.
Price and access
$0.60 per million input tokens and $1.92 per million output, blending to $0.93/M. Listed on OpenRouter, OpenCode Zen and OpenCode Go as well as Z.ai's own platform, with weights published for self-hosting. Note that Z.ai has since shipped GLM 5.3 and GLM 5.3 Flash, both listed unscored on this board — the Flash variant at $0.04/$0.14, which is the cheapest entry anywhere on the page. Treat GLM-5's position as a point-in-time read. Last scored 15 Sep 2026.
Alternatives on this board
- DeepSeek V4 — 86 at $0.78/$1.57 with Coding 92. Four points up overall and clearly better at code for a similar rate.
- Xiaomi MiMo 2.5 Pro — 82 at $0.43/$0.87, the same overall score for a third less, with Speed 83 against GLM-5's 74.
- Qwen 4 Max — 85 overall from Alibaba, three points up, but with no published price on this board.
- Mistral Large 3 — 83 at $0.50/$1.50, if EU hosting and jurisdiction matter to the decision.
Sources
- Z.ai's platform documentation — model identifiers, context limits and current API pricing.
- The Z.ai / Zhipu weights on Hugging Face — architecture and licence terms, which decide whether self-hosting is permitted for your use.
- GLM-5 on OpenRouter for third-party hosting and live rates.
- On this site: local LLM inference, September 2026, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.