Model · stat sheet
Gemma 3 27B
Google's open family; a clean, permissively-licensed base that runs on a single GPU.
What it is
Gemma 3 27B is the largest member of Google's open-weights family, and the oldest model on this board by some distance — Gemma 3 shipped in March 2025, alongside 1B, 4B and 12B siblings. It is here because it is still the reference point for "a capable model on one GPU": the 27B checkpoint, particularly in Google's quantisation-aware-training int4 form, fits on a single consumer card.
The family is multimodal from 4B upwards, handles a 128K context at those sizes, and covers a very wide range of languages. It is not, however, open source in the OSI sense — Gemma ships under Google's own terms of use with usage restrictions attached.
How the composite is built
Overall 79 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. For a model of this age and size, placing 28th of 34 on a board full of 2026 frontier releases is a respectable result.
- Coding 80
- Level with Yi-2 Large and ahead of Mistral Small 3.2 (79). Fine for boilerplate, small functions and translation between languages; twelve behind DeepSeek V4.
- Terminal 74
- The weakest axis, and the expected one. A 27B dense model does not hold a long shell session together; it loses the thread after a handful of commands.
- Reasoning 82
- The best of its six and genuinely good for the size — ahead of Mistral Small 3.2 (78) and level with Llama 5 Scout. This is what Google's distillation work buys.
- Tool use 78
- Workable under supervision, eight behind Command R+ 2. Expect to validate every structured call rather than trusting the schema.
- Context 82
- Solid for a single-GPU model, reflecting the 128K window on the larger Gemma 3 sizes. Note that serving a long context locally costs KV-cache memory you may not have spare after the weights.
- Speed 84
- Fast on hosted inference, level with Amazon Nova 2 Pro. Locally, this number is entirely a property of your hardware and quantisation.
Where it fits
Local and edge deployment where the data cannot leave the machine, and cheap high-volume work where 80-ish quality is sufficient: classification, extraction, summarisation, drafting, offline assistants. It is also the standard starting point for fine-tuning, because the ecosystem support is unusually complete — Ollama, llama.cpp, vLLM, Hugging Face and Google's own tooling all handle it without argument.
Limits
- It is a 2025 model on a 2026 board. The gap to current frontier releases is real and growing. Check whether a newer Gemma generation exists before committing.
- The licence is not OSI-approved. Gemma's terms include use restrictions and a prohibited-use policy that follow the weights downstream. Read them if you plan to redistribute or build a product on top.
- Terminal 74 and Tool use 78 make it a poor agent. Use it as a component behind a stronger orchestrator, not as the loop.
- Value is rated Fair despite the price. 4.9 index points per $0.17/M blended gives 28.4 — cheap tokens, but not many points per token. Cheapness alone does not make a model good value.
Price and access
$0.08 per million input tokens and $0.45 per million output through hosted providers, blending to $0.17/M — the cheapest published rate on the scored board after Mistral Small 3.2. Listed on OpenRouter, and free to download and run yourself under the Gemma terms. Last scored 15 Sep 2026.
Alternatives on this board
- Mistral Small 3.2 — 79 at $0.09/$0.25 with Speed 88 and a genuinely permissive Apache licence. The closest rival and the better licence.
- Phi-5 — 78 with Reasoning 84 and Speed 90 from Microsoft, if reasoning per parameter is the priority.
- Xiaomi MiMo 2.5 Pro — 82 at $0.43/$0.87, three points better for a higher but still small price.
- DeepSeek V4 — 86 with Coding 92, when the single-GPU constraint stops being the binding one.
Sources
- Google's Gemma documentation — sizes, context windows, modalities and the quantised checkpoints.
- The Gemma terms of use — read these before shipping anything built on the weights.
- Gemma 3 27B on OpenRouter for hosted pricing and providers.
- On this site: local LLM inference, September 2026, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.