Model · stat sheet
Gemini 3.6 Flash
Google's speed tier; cheap and quick with a big context, ideal for high-volume calls.
What it is
Gemini 3.6 Flash is Google's speed tier: the distilled sibling of Gemini 3.6 Pro, built for high call volumes where latency and unit cost decide the architecture. Its defining trait on this board is that it keeps most of Pro's context handling — 93 against 98 — while gaining twenty-two points of Speed and arriving with a published price.
It is the only Google model on this board listed on OpenRouter and OpenCode Zen as well as Google's own surfaces, which makes it the easy one to drop into an existing gateway-based harness.
How the composite is built
Overall 84 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Flash is a lopsided model and the index flattens that — its two best axes carry 21% of the weight between them.
- Coding 84
- Seven behind Pro. Competent on contained changes, unreliable on anything that needs a plan held across several files.
- Terminal 80
- The weakest axis, and six behind Pro's already-soft 86. Not a model to leave alone with a shell.
- Reasoning 83
- Ten behind Pro. That is the distillation showing. Give it well-posed questions, not open ones.
- Tool use 82
- Reasonable for the tier — level with Qwen 4 Max and ahead of DeepSeek R2's 79. Fine for schema-bound calls under supervision.
- Context 93
- Fourth on the board behind Pro (98) and the Fable/Mythos pair (95), and the reason to pick Flash over other cheap models. Long inputs at a low rate is a rare combination.
- Speed 88
- Behind Claude Haiku 4.5 (94), Grok 5 Mini (91) and the GPT-5.6 Luna / Phi-5 pair (90) — and none of those come close to its context score.
Where it fits
Retrieval-augmented answering over large corpora, document and transcript processing at volume, classification and extraction where the input is long, and as the fast front end to a slower model. If your bottleneck is "we need to read a lot of text cheaply and quickly", this is the model on the board that does it best.
Limits
- Reasoning 83 and Terminal 80 are the ceiling. Flash reads well and decides poorly. Keep the decisions with a bigger model.
- Output is five times the input rate ($3.75 against $0.75). The economics only work while responses stay short; a chatty prompt template erases the advantage.
- Value is rated Fair, not Excellent — 34.0 index points per $1.50/M blended gives 22.7, behind Xiaomi MiMo 2.5 Pro (47.8) and well behind GPT-5.6 Luna (82.9).
- Context 93 is a handling score, not a capacity claim. Check Google's published window for the tier and surface you are actually calling.
Price and access
$0.75 per million input tokens and $3.75 per million output, blending to $1.50/M. Available through the Gemini API, Google AI Studio and Vertex AI, and listed on OpenRouter and OpenCode Zen. Google's free AI Studio quotas make this a cheap model to evaluate before committing, though those quotas change frequently enough that it is not worth quoting a figure. Last scored 15 Sep 2026.
Alternatives on this board
- GPT-5.6 Luna — 81 at $0.20/$1.20 with Context 89 and Speed 90. Three points down for a third of the blended price; the direct rival.
- Gemini 3.6 Pro — 89 with Context 98, when the reasoning over the long input matters more than the latency.
- Kimi k3 — 84 overall with Context 92 at $3/$15 and open weights. Same overall score, four times the price, self-hostable.
- Gemma 3 27B — 79 at $0.08/$0.45 and open weights, if the work can run on your own single GPU.
Sources
- Google's Gemini API model documentation — context windows, modalities and rate limits per tier.
- Gemini 3.6 Flash on OpenRouter for live pricing and provider availability.
- On this site: Gemini CLI, using OpenRouter as a gateway, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.