fullauto.online

← Model leaderboard

Model · stat sheet

GPT-5.6 Sol

OpenAI·frontier
Overall
90
Rank
#4 / 34
Coding93
Terminal89
Reasoning95
Tool use91
Context92
Speed60
CodingTerminalReasoningTool useContextSpeed

OpenAI's frontier tier for complex professional work — broad, reliable and strong across every category.

What it is

GPT-5.6 Sol is OpenAI's frontier tier in the 5.6 generation, sitting above Terra (balanced) and Luna (small). It is the one to reach for on complex professional work, and its distinguishing feature on this board is breadth: no axis below 89 except Speed, and a Reasoning score beaten only by the two slowest Anthropic models.

At $2/$10 it is priced level with Claude Sonnet 5, a tier below it in Anthropic's line-up, which makes this the most interesting price comparison on the board.

How the composite is built

Overall 90 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Sol's rank comes from having no weak axis rather than from winning any of them.

Coding 93
Fourth on the board, behind Opus 5's 96 and the Fable 5 / Mythos 5 pair at 94. A real gap at the top, but Sol is doing it at a fifth of Fable's blended price.
Terminal 89
Six behind Opus 5 and the softest of Sol's strong axes. Long unattended shell sessions are still Anthropic's home ground on this board.
Reasoning 95
Third overall, behind Fable 5's 97 and Mythos 5's 96 — and ahead of Opus 5's 94. For hard thinking per pound spent, this is the best number on the board.
Tool use 91
Strong. Five behind Opus 5 and one behind Fable 5, one ahead of Sonnet 5. Reliable through long function-calling chains, which is what most agent frameworks actually exercise.
Context 92
Behind only Gemini 3.6 Pro (98), Fable 5 and Mythos 5 (95 each) and Gemini 3.6 Flash (93), and level with Kimi k3. We are not quoting a token figure for 5.6 Sol here — check OpenAI's model page, since the published window has moved between generations.
Speed 60
The weak spot, and slower than Opus 5's 62. Sol is a deliberating model. The 8% weight means the composite forgives this far more readily than a user will.

Where it fits

Hard work with a deadline measured in minutes rather than seconds: architectural decisions, tricky debugging, analysis over long documents, review passes on another model's output. It is a credible default for a whole agent loop if you are already in the OpenAI ecosystem, and the natural model behind the OpenAI Agents SDK or Codex CLI.

Limits

  • Speed 60 rules out interactive use without a faster model in front of it. Pair it with Luna or Terra for the fast path.
  • Terminal 89 is its softest heavy-weighted axis. If your workload is long, unattended shell loops rather than reasoning, Opus 5 is measurably better at the thing you are buying.
  • Value is rated Pricey at 11.8 — 47.0 index points per $4.00/M blended. Better than every Anthropic tier on this measure, far behind the open-weights field.
  • The 5.6 line is being superseded. GPT-6 Astra and GPT-6 Sol appear on this board unscored, dated September 2026 with 1M context. Sol's position is an early read that will move.

Price and access

$2 per million input tokens and $10 per million output, blending to $4.00/M. Sold through the OpenAI API and listed on OpenRouter and OpenCode Zen. Check OpenAI's own model page for current context limits, reasoning-effort options and any cached-input discount before budgeting — those terms change more often than the price does. Last scored 15 Sep 2026.

Alternatives on this board

  • Claude Sonnet 5 — 88 at the identical $2/$10, with Speed 78 against Sol's 60. Two points down, much quicker.
  • Claude Opus 5 — 92 at $5/$25. Better coding, terminal and tool use; worse reasoning per pound.
  • GPT-5.6 Terra — 86 at $2/$12, the balanced tier in the same family and a quicker sibling at 75 on Speed.
  • DeepSeek V4 — 86 with Coding 92 at $0.78/$1.57 and open weights. Four points down for a quarter of the spend.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.