fullauto.online

← Model leaderboard

Model · stat sheet

DeepSeek V4

DeepSeek · open weights·open
Overall
86
Rank
#8 / 34
Coding92
Terminal84
Reasoning90
Tool use83
Context86
Speed74
CodingTerminalReasoningTool useContextSpeed

The open-weights model that closed the gap — frontier-class coding you can self-host.

What it is

DeepSeek V4 is the open-weights model that made the gap argument stop working. It scores 92 on Coding — fifth on this board, behind only the three Anthropic frontier tiers and GPT-5.6 Sol, and ahead of Gemini 3.6 Pro — at $0.78/$1.57, with weights you can download and run yourself.

DeepSeek's practice through the V2 and V3 generations was a large mixture-of-experts architecture published on Hugging Face under a permissive licence, with a cheap hosted API alongside. We are not asserting V4's exact parameter count, expert routing or licence text from memory — check the repository, because those details decide whether self-hosting is realistic for you.

How the composite is built

Overall 86 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. V4 is a coding model that scores like one: strong where code is written, weaker where an agent has to orchestrate.

Coding 92
Four behind Opus 5 at roughly a tenth of the blended price, and the highest coding score of any open-weights model here. This single number is the case for the model.
Terminal 84
Eleven behind Opus 5. It writes the patch better than it drives the loop that applies it — a common shape in open-weights models and the main thing to design around.
Reasoning 90
Eighth on the board and ahead of Claude Sonnet 5's 88. Good enough that planning is not usually the failure mode.
Tool use 83
The weakest capability axis and thirteen behind Opus 5. Expect more malformed calls and more recovery logic in long chains. Budget for retries and schema validation.
Context 86
Mid-field, four behind Opus 5 and twelve behind Gemini 3.6 Pro. Adequate for file-level and module-level work, not for whole-repository reads.
Speed 74
Twelve points quicker than Opus 5 on the hosted API — but if you self-host, your speed is a function of your hardware, not this number.

Where it fits

Code generation at volume where you can supervise the loop: patch writing, test authoring, translation between languages, bulk refactoring with a human or a stricter model checking the result. It is also the default answer when the requirement is that weights and data stay inside your own boundary — see local inference as of September 2026 for what that actually takes.

Limits

  • Tool use 83 is the real constraint in an agent harness. The model that writes the best cheap code is not the best cheap model at running a toolchain.
  • Self-hosting is not free. A frontier-class MoE needs serious GPU memory. The $0.78/$1.57 hosted rate is often cheaper than running it yourself at low volume.
  • Data residency cuts both ways. The hosted API is operated from China; if that matters for your compliance posture, use the open weights or a third-party host, not the first-party endpoint.
  • Value is rated Good, not Excellent — 30.4 index points per $0.98/M blended is 31.1, behind Xiaomi MiMo 2.5 Pro and GPT-5.6 Luna.

Price and access

$0.78 per million input tokens and $1.57 per million output on the hosted API, blending to $0.98/M — note how flat that is compared with the closed models, where output typically costs five times input. Listed on OpenRouter, OpenCode Zen and OpenCode Go, which is the broadest gateway availability of any open model here. Weights are published for self-hosting; verify the licence before commercial deployment. Last scored 15 Sep 2026.

Alternatives on this board

  • Qwen 4 Coder — 84 with Coding 90, open weights from Alibaba. The closest rival on the coding axis specifically.
  • GLM-5 — 82 at $0.60/$1.92 with open weights and the same three-gateway availability. Cheaper, three points down, slightly weaker on code.
  • Kimi k3 — 84 with Context 92 at $3/$15. Better long-context handling, much higher hosted price.
  • Claude Sonnet 5 — 88 at $2/$10 with Tool use 90. The upgrade to buy if your agent loop is failing on orchestration rather than on code.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.