fullauto.online

← Model leaderboard

Model · stat sheet

GPT-5.6 Luna

OpenAI·small
Overall
81
Rank
#22 / 34
Coding80
Terminal76
Reasoning80
Tool use80
Context89
Speed90
CodingTerminalReasoningTool useContextSpeed

OpenAI's cost tier for high-volume, latency-sensitive workloads at the GPT-5.6 context size.

What it is

GPT-5.6 Luna is OpenAI's cost tier in the 5.6 generation, below Terra and Sol. The thing that makes it interesting is not its rank — 22nd of 34 — but that it keeps the family's context handling while giving up capability everywhere else, and prices accordingly.

On this board's value measure it is the single best model listed: 37.3 index points per $0.45/M blended works out at 82.9, rated Excellent, against 47.8 for the next best and 5.1 for Claude Opus 5.

How the composite is built

Overall 81 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Luna is penalised by that shape: its two best axes carry 21% between them, its two weakest carry 44%.

Coding 80
Mid-table, and eight to thirteen points behind the frontier tiers. Adequate for bounded, well-specified edits; not a model to hand an open-ended refactor.
Terminal 76
The weakest axis. Long unattended shell loops are where a small model's mistakes compound, and this score is an honest warning rather than a rounding error.
Reasoning 80
Fifteen behind Sol's 95. The tier gap is real and it shows up as wrong plans rather than wrong syntax.
Tool use 80
Solid for the price — level with GLM-5 and Kimi k3, five behind Claude Haiku 4.5's 85. Good enough for structured function calls in a supervised loop.
Context 89
The surprise on the sheet. Higher than Claude Sonnet 5's 88 and far above Haiku 4.5's 72, from a model at a fifth of the price. Long inputs are cheap here.
Speed 90
Joint third among scored models with Phi-5, behind Haiku 4.5 (94) and Grok 5 Mini (91). Combined with the price, this is what Luna is for.

Where it fits

High-volume, latency-sensitive work with long inputs: retrieval-augmented answering over big documents, log triage, bulk classification and extraction, first-pass summarisation before a bigger model sees the text. It is also a good fast path in front of Sol — Luna answers, Sol handles the escalations.

Limits

  • Terminal 76 and Reasoning 80 set a hard ceiling. Do not give it autonomy over a shell. The failures are quiet and the cheap tokens make it tempting to run it unsupervised.
  • Excellent value is a ratio, not a verdict. 82.9 means good points per pound; it does not mean the points you need are there.
  • Output tokens cost six times input ($1.20 against $0.20). Verbose responses erase the saving faster than people expect — cap the output length.
  • The 5.6 line is being superseded. GPT-6 models appear unscored on this board dated September 2026; this is an early read.

Price and access

$0.20 per million input tokens and $1.20 per million output, blending to $0.45/M — the cheapest non-open model on the scored board. Sold through the OpenAI API and listed on OpenRouter, OpenCode Zen and OpenCode Go, which is unusually broad availability for a small model. Check OpenAI's model page for the current context limit rather than assuming it matches Sol. Last scored 15 Sep 2026.

Alternatives on this board

  • Gemini 3.6 Flash — 84 at $0.75/$3.75, Context 93, Speed 88. Three points better for roughly three times the blended price.
  • Xiaomi MiMo 2.5 Pro — 82 at $0.43/$0.87 with open weights and value 47.8. Cheaper on output, better overall, weaker on context.
  • Claude Haiku 4.5 — 80 at $1/$5 with Speed 94 and Tool use 85. Faster and better at tool calls; much worse context and four times the price.
  • Mistral Small 3.2 — 79 at $0.09/$0.25 and open weights, if self-hosting matters more than the last two points.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.