fullauto.online

← Model leaderboard

Model · stat sheet

GLM-5

Zhipu AI · open weights·open
Overall
82
Rank
#19 / 34
Coding85
Terminal79
Reasoning86
Tool use80
Context85
Speed74
CodingTerminalReasoningTool useContextSpeed

Zhipu AI's flagship open model; strong bilingual reasoning and dependable agentic tool use.

What it is

GLM-5 is the flagship of Zhipu AI's open-weights line — the lab also trades as Z.ai internationally — and it is one of the small group of Chinese open models that ship weights, a cheap first-party API and listings on the Western gateways at the same time. On this board it appears on OpenRouter, OpenCode Zen and OpenCode Go, which is the same three-gateway reach as DeepSeek V4.

The GLM series has consistently been positioned around agentic and bilingual Chinese–English use. Verify the specific licence, parameter count and context window for GLM-5 in Z.ai's own documentation rather than from any leaderboard, including this one.

How the composite is built

Overall 82 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. GLM-5's profile is unusually even for an open model, with no axis below 74 and none above 86.

Coding 85
Seven behind DeepSeek V4's 92 and level with Mistral Large 3. Competent on ordinary work; not the open model to pick if code quality is the only criterion.
Terminal 79
The weakest axis. Sixteen behind Opus 5 and five behind V4. Long unattended shell sessions are not where this model is strongest.
Reasoning 86
The best of its six. Ahead of Qwen 4 Coder's 84 and level with Llama 5. Solid planning for the price.
Tool use 80
Worth reading carefully against the reputation. GLM is marketed on agentic tool use, and 80 is respectable — level with GPT-5.6 Luna and Kimi k3 — but it is sixteen behind Opus 5 and five behind Claude Haiku 4.5, which is a small model by any measure. Dependable, not exceptional.
Context 85
Mid-field, level with DeepSeek R2. Fine for module-scale work.
Speed 74
Level with DeepSeek V4 on hosted inference. As with any open model, self-hosting makes this number a property of your hardware rather than the model.

Where it fits

Cost-controlled agent work where you want open weights as an exit option: supervised coding loops, bilingual applications, internal tooling where the data cannot leave your infrastructure. It is a reasonable default for teams who want DeepSeek-class economics with a second vendor in the mix.

Limits

  • No value rating on this board — the column reads Unrated because no Artificial Analysis index is published for this model, not because the price is bad. At $0.60/$1.92 the price is among the better ones here.
  • Terminal 79 caps autonomy. Keep a human or a stricter model between GLM-5 and anything destructive.
  • The agentic reputation runs ahead of the tool-use number. Test on your own tool schemas before believing the marketing or the leaderboard.
  • First-party hosting is in China. For regulated workloads, use the open weights or a Western gateway rather than the direct endpoint.

Price and access

$0.60 per million input tokens and $1.92 per million output, blending to $0.93/M. Listed on OpenRouter, OpenCode Zen and OpenCode Go as well as Z.ai's own platform, with weights published for self-hosting. Note that Z.ai has since shipped GLM 5.3 and GLM 5.3 Flash, both listed unscored on this board — the Flash variant at $0.04/$0.14, which is the cheapest entry anywhere on the page. Treat GLM-5's position as a point-in-time read. Last scored 15 Sep 2026.

Alternatives on this board

  • DeepSeek V4 — 86 at $0.78/$1.57 with Coding 92. Four points up overall and clearly better at code for a similar rate.
  • Xiaomi MiMo 2.5 Pro — 82 at $0.43/$0.87, the same overall score for a third less, with Speed 83 against GLM-5's 74.
  • Qwen 4 Max — 85 overall from Alibaba, three points up, but with no published price on this board.
  • Mistral Large 3 — 83 at $0.50/$1.50, if EU hosting and jurisdiction matter to the decision.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.