Model · stat sheet
Claude Sonnet 5
The workhorse most agent loops should start on — near-frontier quality at a fraction of the cost and latency.
What it is
Claude Sonnet 5 is Anthropic's middle tier, dated June 2026 here, and the model most agent loops should be built on before anything more expensive is considered. It carries the same 1M-token context window and 128k maximum output as Opus 5 and Fable 5, with adaptive thinking available rather than always on, at a fifth of Fable's per-token price.
The interesting thing about its score is how flat it is. Five of six axes land between 88 and 90. There is no hole to design around.
How the composite is built
Overall 88 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Four points behind Opus 5 across the board rather than badly behind on anything.
- Coding 90
- Six behind Opus 5's 96. In practice that gap shows up on long, multi-file changes rather than on ordinary feature work, which is the case for starting here and escalating on measured failures.
- Terminal 90
- Level with Fable 5 and five behind Opus 5, despite costing a fifth of Fable's rate. The best terminal score per pound on the Anthropic line.
- Reasoning 88
- The lowest of the six and the axis where the tier gap is real. On genuinely ambiguous problems Fable 5's 97 is nine points clear; that is what escalation buys.
- Tool use 90
- Dependable across long chains. Six behind Opus 5, but ahead of every open-weights model on the board.
- Context 88
- Two behind Opus 5 despite the same nominal window, which is a read on effective recall rather than capacity. The Opus 4.7-generation tokeniser applies here too: roughly 30% more tokens for the same text than pre-4.7 Claude models.
- Speed 78
- Sixteen points clear of Opus 5 and thirty clear of Fable 5. This is the number that makes Sonnet 5 usable in interactive tools, and it is the axis the composite weights least.
Where it fits
The default. Interactive coding assistants, review passes, ticket-sized changes, anything a person is waiting on. It is also the sensible base layer in a tiered system: run the loop on Sonnet 5, route the cases that fail to Opus 5, and reserve Fable 5 for the handful of problems that defeat both.
Limits
- Reasoning 88 is the ceiling to watch. If your failures are "it picked the wrong approach" rather than "it wrote a bug", a bigger model is the fix, not a better prompt.
- Value is rated Poor. 38.2 index points per $4.00/M blended is 9.6 — better than Opus 5 and Fable 5, far worse than the open-weights field. Cheap relative to its siblings is not the same as cheap.
- Same tokeniser trap as the rest of the line. Budgets ported from pre-4.7 Claude models will run over.
- The effort parameter defaults to high on the Claude API and in Claude Code. On tolerant, high-volume work, lowering it is a cheaper first move than switching model.
Price and access
$2 per million input tokens and $10 per million output, blending to $4.00/M. Sold through the Claude API and listed on OpenRouter and OpenCode Zen — the same blended rate as GPT-5.6 Sol, which scores two points higher. Last scored 15 Sep 2026.
Alternatives on this board
- GPT-5.6 Sol — 90 overall at the same $2/$10, with Reasoning 95. The head-to-head worth running on your own workload.
- GPT-5.6 Terra — 86 at $2/$12, OpenAI's balanced tier and the closest match in intent.
- DeepSeek V4 — 86 overall with Coding 92 at $0.78/$1.57 and open weights. Better at code, cheaper, weaker on tool use.
- Claude Haiku 4.5 — 80 at $1/$5 with Speed 94, for the small calls that should never reach Sonnet.
Sources
- Anthropic's model overview — context, output caps, thinking behaviour and cutoffs.
- Sonnet 5 on OpenRouter for live pricing and providers.
- On this site: how the Claude 5 tiers differ, Claude Code, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.