fullauto.online

← Model leaderboard

Model · stat sheet

Gemini 3.6 Flash

Google·balanced
Overall
84
Rank
#13 / 34
Coding84
Terminal80
Reasoning83
Tool use82
Context93
Speed88
CodingTerminalReasoningTool useContextSpeed

Google's speed tier; cheap and quick with a big context, ideal for high-volume calls.

What it is

Gemini 3.6 Flash is Google's speed tier: the distilled sibling of Gemini 3.6 Pro, built for high call volumes where latency and unit cost decide the architecture. Its defining trait on this board is that it keeps most of Pro's context handling — 93 against 98 — while gaining twenty-two points of Speed and arriving with a published price.

It is the only Google model on this board listed on OpenRouter and OpenCode Zen as well as Google's own surfaces, which makes it the easy one to drop into an existing gateway-based harness.

How the composite is built

Overall 84 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Flash is a lopsided model and the index flattens that — its two best axes carry 21% of the weight between them.

Coding 84
Seven behind Pro. Competent on contained changes, unreliable on anything that needs a plan held across several files.
Terminal 80
The weakest axis, and six behind Pro's already-soft 86. Not a model to leave alone with a shell.
Reasoning 83
Ten behind Pro. That is the distillation showing. Give it well-posed questions, not open ones.
Tool use 82
Reasonable for the tier — level with Qwen 4 Max and ahead of DeepSeek R2's 79. Fine for schema-bound calls under supervision.
Context 93
Fourth on the board behind Pro (98) and the Fable/Mythos pair (95), and the reason to pick Flash over other cheap models. Long inputs at a low rate is a rare combination.
Speed 88
Behind Claude Haiku 4.5 (94), Grok 5 Mini (91) and the GPT-5.6 Luna / Phi-5 pair (90) — and none of those come close to its context score.

Where it fits

Retrieval-augmented answering over large corpora, document and transcript processing at volume, classification and extraction where the input is long, and as the fast front end to a slower model. If your bottleneck is "we need to read a lot of text cheaply and quickly", this is the model on the board that does it best.

Limits

  • Reasoning 83 and Terminal 80 are the ceiling. Flash reads well and decides poorly. Keep the decisions with a bigger model.
  • Output is five times the input rate ($3.75 against $0.75). The economics only work while responses stay short; a chatty prompt template erases the advantage.
  • Value is rated Fair, not Excellent — 34.0 index points per $1.50/M blended gives 22.7, behind Xiaomi MiMo 2.5 Pro (47.8) and well behind GPT-5.6 Luna (82.9).
  • Context 93 is a handling score, not a capacity claim. Check Google's published window for the tier and surface you are actually calling.

Price and access

$0.75 per million input tokens and $3.75 per million output, blending to $1.50/M. Available through the Gemini API, Google AI Studio and Vertex AI, and listed on OpenRouter and OpenCode Zen. Google's free AI Studio quotas make this a cheap model to evaluate before committing, though those quotas change frequently enough that it is not worth quoting a figure. Last scored 15 Sep 2026.

Alternatives on this board

  • GPT-5.6 Luna — 81 at $0.20/$1.20 with Context 89 and Speed 90. Three points down for a third of the blended price; the direct rival.
  • Gemini 3.6 Pro — 89 with Context 98, when the reasoning over the long input matters more than the latency.
  • Kimi k3 — 84 overall with Context 92 at $3/$15 and open weights. Same overall score, four times the price, self-hostable.
  • Gemma 3 27B — 79 at $0.08/$0.45 and open weights, if the work can run on your own single GPU.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.