Model · stat sheet
DeepSeek R2
DeepSeek's reasoning line; deliberate chain-of-thought that punches above its size on maths and proofs.
What it is
DeepSeek R2 is the reasoning branch of DeepSeek's line — the successor to the R1 approach that made explicit, visible chain-of-thought training a mainstream technique rather than a research curiosity. Where DeepSeek V4 is a general model that codes well, R2 spends tokens thinking before it answers, and the score sheet shows exactly that trade.
R2 carries the board's reasoning tag rather than the open-weights tag that V4 has. DeepSeek has published weights for its reasoning models before, but we are not going to assert R2's licence or availability from memory — check the repository, since that single fact changes what you can do with it.
How the composite is built
Overall 83 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. R2 is a specialist being scored by a generalist index: it wins on one axis worth a fifth of the total and pays for it on the rest.
- Coding 84
- Eight behind V4's 92. A reasoning model is not automatically a better programmer; thinking longer about the wrong abstraction still produces the wrong abstraction.
- Terminal 79
- Weak, sixteen behind Opus 5. Deliberation and shell loops are a bad combination — each command costs a thinking pass.
- Reasoning 93
- Joint fifth on the board with Gemini 3.6 Pro, ahead of Grok 5 (91) and Claude Sonnet 5 (88). For maths, proofs and problems with a checkable answer, this is a genuinely frontier-adjacent number from a much smaller operation.
- Tool use 79
- The weakest axis, level with its Terminal score. Reasoning training tends not to help with schema adherence, and it sometimes hurts — the model reasons about the call instead of making it.
- Context 85
- Mid-field. Enough for a long problem statement and its supporting material, not enough for whole-codebase work.
- Speed 64
- Slow, and honestly so — two points behind Gemini 3.6 Pro's 66 and two ahead of Opus 5's 62. Visible chain-of-thought costs output tokens and wall-clock time by construction.
Where it fits
Bounded problems with a verifiable answer: maths, formal reasoning, algorithmic design, test-case derivation, root-cause analysis on a well-described failure. It is a component, not a harness — put it behind a step in a pipeline where a better tool-caller collects its output and acts on it.
Limits
- Tool use 79 rules it out as an agent driver. Use V4 or a Claude tier for the loop and call R2 for the hard step.
- No published price on this board, so no value rating. The dash in the value column is missing data, not a poor score.
- Reasoning tokens are billed tokens. Wherever you run it, a visible chain of thought inflates output volume, so a low per-token rate can still produce a high per-task cost.
- Longer thinking is not always better thinking. On simple questions a reasoning model can talk itself out of the right answer. Route by difficulty rather than defaulting here.
Access
Through DeepSeek's own API. This board records no OpenRouter or OpenCode listing and no published price for R2, which is unusual next to V4's three-gateway availability — check DeepSeek's API docs for current model identifiers and rates before planning around it. Last scored 15 Sep 2026.
Alternatives on this board
- Claude Fable 5 — Reasoning 97, always-on adaptive thinking, at $10/$50. The frontier version of the same idea, at roughly twenty times the likely cost.
- GPT-5.6 Sol — Reasoning 95 at $2/$10, with Tool use 91. Better thinking and far better orchestration, at a published price.
- DeepSeek V4 — 86 overall, Reasoning 90, $0.78/$1.57 and open weights. Three points of reasoning down for a much more usable general model.
- Phi-5 — 78 overall but Reasoning 84 from a small open model, if the constraint is hardware rather than capability.
Sources
- DeepSeek's API documentation — model identifiers, reasoning-token behaviour and pricing.
- The DeepSeek repositories on GitHub — papers, weights and licence terms, which is where the open-weights question is actually settled.
- On this site: DeepSeek V4 for the general-purpose sibling, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.