fullauto.online

← Model leaderboard

Model · stat sheet

Qwen 4 Coder

Alibaba · open weights·open
Overall
84
Rank
#16 / 34
Coding90
Terminal80
Reasoning84
Tool use80
Context84
Speed78
CodingTerminalReasoningTool useContextSpeed

Alibaba's code-specialist open weights; among the best self-hostable models for repository-scale coding.

What it is

Qwen 4 Coder is Alibaba's code-specialist open weights — the note on this page calls it among the best self-hostable models for repository-scale coding, and the Coding score of 90 backs the claim. The Qwen line has been the main Chinese open-weights counterweight to Llama since 2023, and the Coder branch is where Alibaba concentrates the code data.

Weights and licence live on Alibaba's Qwen channels; the licence terms decide whether commercial self-hosting is straightforward, so read them rather than assuming from the label "open".

How the composite is built

Overall 84 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. A specialist shape: Coding 90 leads, and the other five axes fall to a flat 78 to 84.

Coding 90
Level with Sonnet 5 and Grok 5, and the strongest self-hostable coding score here after DeepSeek V4's 92. This number is the case for the model.
Terminal 80
Fifteen behind Opus 5's 95 and three behind Llama 5's 83. It writes the patch better than it drives the loop that applies it.
Reasoning 84
Level with Mistral Large 3 and MiMo 2.5 Pro. Planning is adequate for task-scale work.
Tool use 80
Level with GLM-5 and Kimi k3, ten behind Sonnet 5's 90. Budget for retries and schema validation in long chains.
Context 84
Level with Llama 5 and Nemo 3 Ultra. Module-scale work; "repository-scale" in the note means a fair amount of retrieval, not a whole tree in one prompt.
Speed 78
Two above Qwen 4 Max's 76 and behind the small tiers by a wide margin. Self-hosted, the number is your hardware.

Where it fits

Code generation where you supervise the result: patch writing, test authoring, translation between languages, bulk refactoring checked by humans or stricter models. And the usual open-weights case — when the code cannot leave your infrastructure and you still want near-frontier coding quality.

Limits

  • Terminal 80 and Tool use 80 cap autonomy. The strongest cheap code writer here is not the strongest cheap loop runner.
  • No price or index is published here, so Value is Unrated — the cost of running it is GPU time. Compare that honestly against DeepSeek V4's $0.78/$1.57 hosted rate before committing to hardware.
  • It is a specialist. General reasoning work is better served by Qwen 4 Max or the generalist tiers.
  • Alibaba's Qwen line releases frequently, and the board carries unscored Qwen listings — Qwen3.8 Flash at $0.15/$0.47 and others. Check what is current.

Price and access

No published price or gateway listing on this board — access is the weights from Alibaba's channels, self-hosted or through a third-party host. Last scored 15 Sep 2026.

Alternatives on this board

  • DeepSeek V4 — 86 at $0.78/$1.57 with Coding 92 and open weights. The one open model that beats it on code, with a hosted rate attached.
  • Qwen 4 Max — 85 with Coding 88, the same vendor's generalist.
  • GLM-5 — 82 at $0.60/$1.92 with Coding 85 and three gateways.
  • Claude Sonnet 5 — 88 at $2/$10 with Tool use 90, when the failure mode is orchestration rather than code.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.