fullauto.online

← Model leaderboard

Model · stat sheet

Gemini 3.6 Pro

Google·frontier
Overall
89
Rank
#5 / 34
Coding91
Terminal86
Reasoning93
Tool use88
Context98
Speed66
CodingTerminalReasoningTool useContextSpeed

Google's frontier tier; unmatched long-context handling and multimodal breadth.

What it is

Gemini 3.6 Pro is Google's frontier tier, and on this board it owns one axis outright: Context, at 98. Google's line has carried the longest usable windows in the market since the 1.5 generation, and that lead is still the reason to pick this model rather than a competitor two points higher overall.

It is served through the Gemini API, Google AI Studio and Vertex AI. It does not appear on OpenRouter or the OpenCode gateways in this board's source data, which is worth knowing if your harness talks to one of those rather than directly to Google.

How the composite is built

Overall 89 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Context is worth only 13%, so the axis Gemini wins is the one the index rewards least — read the bars, not just the total.

Coding 91
Fifth on the board, behind Opus 5 (96), Fable 5 and Mythos 5 (94) and GPT-5.6 Sol (93). Genuinely frontier, not quite the top.
Terminal 86
The weakest of its capability axes and nine behind Opus 5. Unattended shell work is the clearest case against Gemini 3.6 Pro on this board.
Reasoning 93
Joint fifth with DeepSeek R2, behind the three Anthropic frontier tiers and GPT-5.6 Sol. Strong enough that reasoning is rarely the reason to look elsewhere.
Tool use 88
Good, eight behind Opus 5. Adequate for structured function calling; the gap widens on very long chains.
Context 98
The highest score on the board on any axis, by any model. Whole-repository reading, long transcripts, large document sets. We have not printed a token figure because Google's published window differs by surface and tier — check the model page.
Speed 66
Slow for a non-Anthropic frontier model, though six points ahead of GPT-5.6 Sol's 60. If you want Google and you want speed, that is what Flash is for.

Where it fits

Work where the input is enormous and the reasoning over it has to hold together: reading an unfamiliar codebase end to end, analysis across long document sets, multimodal inputs where Google's breadth is an advantage. It is a poor fit for a long-running terminal agent and a good fit as the analysis stage feeding one.

Limits

  • No published price on this board, so no value rating — the column shows a dash, not a bad score. Google's pricing varies by tier, context length and surface; you have to price your own workload.
  • Terminal 86 is the real weakness. On a board where the top model scores 95, that is a meaningful gap on 20% of the index.
  • Context 98 measures handling, not free capacity. Long windows still cost per token and still degrade retrieval; a big window is not a substitute for retrieval design.
  • Availability is narrower than the competition's. No OpenRouter or OpenCode listing here means an extra integration if your harness is gateway-based.

Access

Via the Gemini API, Google AI Studio and Vertex AI. Vertex adds the enterprise controls — regional deployment, VPC, IAM — that are usually the reason a team picks Google in the first place. Google publishes free-tier quotas on AI Studio that change often; treat any figure you read elsewhere as stale. Last scored 15 Sep 2026.

Alternatives on this board

  • Gemini 3.6 Flash — 84 at $0.75/$3.75 with Context 93 and Speed 88. Most of the context advantage, a fraction of the latency, a published price.
  • GPT-5.6 Sol — 90 at $2/$10, with Context 92. One point up overall and priced in the open.
  • Claude Opus 5 — 92 at $5/$25. Better at everything except context and worse on none of the heavy axes.
  • Kimi k3 — 84 with Context 92 at $3/$15, open weights, if long context inside your own boundary is the requirement.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.