Model · stat sheet
Claude Opus 5
Anthropic's agentic flagship and the default for long, hands-off coding and terminal work. Tops the board on tool use and autonomous operations.
What it is
Claude Opus 5 is Anthropic's agentic tier, dated July 2026 here. Anthropic's own model documentation points at Opus 5 — not the nominally stronger Fable 5 — as the model to reach for on complex agentic coding and enterprise work. That distinction is worth holding onto: Opus 5 is not a cut-down Fable, it is the recommended default, tuned to finish long jobs rather than to think hardest about short ones.
It shares the 1M-token context window and 128k maximum output of the rest of the Claude 5 line, and adaptive thinking is available rather than compulsory — the lever Fable 5 takes away.
How the composite is built
Overall 92 is a weighted sum: Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Opus 5 leads the board because the heaviest axes are also its best.
- Coding 96
- The board high, on the axis that carries the most weight. Nothing else scored here reaches 96 — Fable 5 sits at 94, DeepSeek V4 at 92.
- Terminal 95
- Also the board high. This is what separates a model that writes a good patch from one that can drive a shell, read its own failures and carry on without a human.
- Reasoning 94
- Strong, not top. Fable 5's 97, Mythos 5's 96 and GPT-5.6 Sol's 95 all beat it. Opus 5 is the pragmatist of the line rather than its deepest thinker.
- Tool use 96
- The board high, four clear of Mythos 5 and Fable 5. Long tool chains are where the overall 92 actually comes from, and this is the axis to weigh if your agent lives inside MCP servers and shell calls.
- Context 90
- Good rather than class-leading — Gemini 3.6 Pro takes 98. Note the tokeniser introduced with Opus 4.7: the same text yields roughly 30% more tokens than pre-4.7 Claude models, so a 1M window holds less prose than the number implies.
- Speed 62
- The weakest figure on the sheet and the honest price of the rest. Among scored models only Fable 5 and Mythos 5 (48 each) are slower. At 8% weight it barely dents the composite — a scoring choice you should second-guess if a person is waiting on the output.
Where it fits
Unattended work with a long horizon: multi-file refactors, framework migrations, test-and-fix loops, anything where the agent runs for minutes and the correctness of the last step matters more than the latency of the first. It is the default pairing for Claude Code and the Claude Agent SDK.
Limits
- It is slow, and the weighting hides it. A composite built around autonomous work under-prices latency at 8%. Read the Speed bar directly before putting Opus 5 in an interactive path.
- Reasoning is not the top tier. For genuinely ambiguous problems Fable 5 is the escalation, at double the price.
- Token budgets ported from older Claude models will be wrong, thanks to the 4.7-generation tokeniser. Recheck chunk sizes and cost models rather than trusting the nominal window.
- Value is poor by design. 50.8 index points per $10.00/M blended gives 5.1 — near the bottom of the rated field. You are buying reliability, not efficiency.
Price and access
$5 per million input tokens and $25 per million output, blending to $10.00/M for the board's value column. Sold through the Claude API and listed on OpenRouter and OpenCode Zen. The effort parameter defaults to high on the Claude API and in Claude Code, and lowering it is the first cost lever to try before changing model. Anthropic has since shipped Claude Opus 5.5 at $4/$20 — cheaper than Opus 5 — which appears on the board unscored; treat this sheet as a point-in-time read, last scored 15 Sep 2026.
Alternatives on this board
- Claude Sonnet 5 — 88 overall at $2/$10, Speed 78. The sensible place to start a loop; escalate only where you have measured a failure.
- Claude Fable 5 — 90 overall, Reasoning 97, $10/$50. Better thinking, worse coding and terminal work, far slower.
- GPT-5.6 Sol — 90 overall at $2/$10. Two points behind for a quarter of the blended price; the serious cost-side comparison.
- DeepSeek V4 — 86 overall, Coding 92, $0.78/$1.57 with open weights. Six points down for roughly a tenth of the spend.
Sources
- Anthropic's model overview — the place to verify context, output limits, thinking behaviour and knowledge cutoffs.
- Opus 5 on OpenRouter for current per-token pricing and provider availability.
- On this site: what actually changes across the Claude 5 tiers, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.