fullauto.online

← Model leaderboard

Model · stat sheet

Claude Sonnet 5

Anthropic · Jun 2026·balanced
Overall
88
Rank
#6 / 34
Coding90
Terminal90
Reasoning88
Tool use90
Context88
Speed78
CodingTerminalReasoningTool useContextSpeed

The workhorse most agent loops should start on — near-frontier quality at a fraction of the cost and latency.

What it is

Claude Sonnet 5 is Anthropic's middle tier, dated June 2026 here, and the model most agent loops should be built on before anything more expensive is considered. It carries the same 1M-token context window and 128k maximum output as Opus 5 and Fable 5, with adaptive thinking available rather than always on, at a fifth of Fable's per-token price.

The interesting thing about its score is how flat it is. Five of six axes land between 88 and 90. There is no hole to design around.

How the composite is built

Overall 88 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Four points behind Opus 5 across the board rather than badly behind on anything.

Coding 90
Six behind Opus 5's 96. In practice that gap shows up on long, multi-file changes rather than on ordinary feature work, which is the case for starting here and escalating on measured failures.
Terminal 90
Level with Fable 5 and five behind Opus 5, despite costing a fifth of Fable's rate. The best terminal score per pound on the Anthropic line.
Reasoning 88
The lowest of the six and the axis where the tier gap is real. On genuinely ambiguous problems Fable 5's 97 is nine points clear; that is what escalation buys.
Tool use 90
Dependable across long chains. Six behind Opus 5, but ahead of every open-weights model on the board.
Context 88
Two behind Opus 5 despite the same nominal window, which is a read on effective recall rather than capacity. The Opus 4.7-generation tokeniser applies here too: roughly 30% more tokens for the same text than pre-4.7 Claude models.
Speed 78
Sixteen points clear of Opus 5 and thirty clear of Fable 5. This is the number that makes Sonnet 5 usable in interactive tools, and it is the axis the composite weights least.

Where it fits

The default. Interactive coding assistants, review passes, ticket-sized changes, anything a person is waiting on. It is also the sensible base layer in a tiered system: run the loop on Sonnet 5, route the cases that fail to Opus 5, and reserve Fable 5 for the handful of problems that defeat both.

Limits

  • Reasoning 88 is the ceiling to watch. If your failures are "it picked the wrong approach" rather than "it wrote a bug", a bigger model is the fix, not a better prompt.
  • Value is rated Poor. 38.2 index points per $4.00/M blended is 9.6 — better than Opus 5 and Fable 5, far worse than the open-weights field. Cheap relative to its siblings is not the same as cheap.
  • Same tokeniser trap as the rest of the line. Budgets ported from pre-4.7 Claude models will run over.
  • The effort parameter defaults to high on the Claude API and in Claude Code. On tolerant, high-volume work, lowering it is a cheaper first move than switching model.

Price and access

$2 per million input tokens and $10 per million output, blending to $4.00/M. Sold through the Claude API and listed on OpenRouter and OpenCode Zen — the same blended rate as GPT-5.6 Sol, which scores two points higher. Last scored 15 Sep 2026.

Alternatives on this board

  • GPT-5.6 Sol — 90 overall at the same $2/$10, with Reasoning 95. The head-to-head worth running on your own workload.
  • GPT-5.6 Terra — 86 at $2/$12, OpenAI's balanced tier and the closest match in intent.
  • DeepSeek V4 — 86 overall with Coding 92 at $0.78/$1.57 and open weights. Better at code, cheaper, weaker on tool use.
  • Claude Haiku 4.5 — 80 at $1/$5 with Speed 94, for the small calls that should never reach Sonnet.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.