Model · stat sheet
Qwen 4 Coder
Alibaba's code-specialist open weights; among the best self-hostable models for repository-scale coding.
What it is
Qwen 4 Coder is Alibaba's code-specialist open weights — the note on this page calls it among the best self-hostable models for repository-scale coding, and the Coding score of 90 backs the claim. The Qwen line has been the main Chinese open-weights counterweight to Llama since 2023, and the Coder branch is where Alibaba concentrates the code data.
Weights and licence live on Alibaba's Qwen channels; the licence terms decide whether commercial self-hosting is straightforward, so read them rather than assuming from the label "open".
How the composite is built
Overall 84 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. A specialist shape: Coding 90 leads, and the other five axes fall to a flat 78 to 84.
- Coding 90
- Level with Sonnet 5 and Grok 5, and the strongest self-hostable coding score here after DeepSeek V4's 92. This number is the case for the model.
- Terminal 80
- Fifteen behind Opus 5's 95 and three behind Llama 5's 83. It writes the patch better than it drives the loop that applies it.
- Reasoning 84
- Level with Mistral Large 3 and MiMo 2.5 Pro. Planning is adequate for task-scale work.
- Tool use 80
- Level with GLM-5 and Kimi k3, ten behind Sonnet 5's 90. Budget for retries and schema validation in long chains.
- Context 84
- Level with Llama 5 and Nemo 3 Ultra. Module-scale work; "repository-scale" in the note means a fair amount of retrieval, not a whole tree in one prompt.
- Speed 78
- Two above Qwen 4 Max's 76 and behind the small tiers by a wide margin. Self-hosted, the number is your hardware.
Where it fits
Code generation where you supervise the result: patch writing, test authoring, translation between languages, bulk refactoring checked by humans or stricter models. And the usual open-weights case — when the code cannot leave your infrastructure and you still want near-frontier coding quality.
Limits
- Terminal 80 and Tool use 80 cap autonomy. The strongest cheap code writer here is not the strongest cheap loop runner.
- No price or index is published here, so Value is Unrated — the cost of running it is GPU time. Compare that honestly against DeepSeek V4's $0.78/$1.57 hosted rate before committing to hardware.
- It is a specialist. General reasoning work is better served by Qwen 4 Max or the generalist tiers.
- Alibaba's Qwen line releases frequently, and the board carries unscored Qwen listings — Qwen3.8 Flash at $0.15/$0.47 and others. Check what is current.
Price and access
No published price or gateway listing on this board — access is the weights from Alibaba's channels, self-hosted or through a third-party host. Last scored 15 Sep 2026.
Alternatives on this board
- DeepSeek V4 — 86 at $0.78/$1.57 with Coding 92 and open weights. The one open model that beats it on code, with a hosted rate attached.
- Qwen 4 Max — 85 with Coding 88, the same vendor's generalist.
- GLM-5 — 82 at $0.60/$1.92 with Coding 85 and three gateways.
- Claude Sonnet 5 — 88 at $2/$10 with Tool use 90, when the failure mode is orchestration rather than code.
Sources
- The Qwen blog and documentation — release notes, model cards and context limits.
- The Qwen organisation on Hugging Face — weights and the licence that governs self-hosting.
- On this site: local LLM inference, September 2026, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.