Model · stat sheet
Llama 5 Scout
The smaller, faster Llama 5 for edge deployment and high-throughput self-hosting.
What it is
Llama 5 Scout is the smaller, faster member of Meta's Llama 5 line, positioned by the note on this page for edge deployment and high-throughput self-hosting. It gives up a consistent three to five points on every capability axis against Llama 5 and returns Speed 85 and a lighter serving footprint in exchange.
One caution on naming: in the earlier Llama generation the Scout label came with an unusually long context window. Do not carry that assumption across — check Llama 5 Scout's actual window in Meta's model card.
How the composite is built
Overall 81 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. The capability axes sit in a narrow 78 to 82 band; Speed 85 is the only axis that steps out of it.
- Coding 82
- One behind Xiaomi MiMo 2.5 Pro's 83 and four behind Llama 5's 86. Everyday code work with a human reviewing the diffs.
- Terminal 78
- Five behind Llama 5's 83 and seventeen behind Opus 5's 95. The weakest axis; keep shell autonomy narrow.
- Reasoning 82
- Level with Gemma 3 27B and four behind Llama 5's 86. Fine for planning at task scale.
- Tool use 80
- Level with GLM-5 and Kimi k3, ten behind Sonnet 5's 90. Structured calls work; long chains accumulate errors faster than the 85-plus models.
- Context 82
- Level with Grok 5 Mini and Gemma 3 27B. Mid-field — the model trades context headroom along with everything else.
- Speed 85
- Eight above Llama 5's 77 and the axis the model is sold on. As with all self-hosted numbers, throughput is finally a function of your GPUs and serving stack.
Where it fits
High-throughput self-hosted inference: batch classification, extraction, code assistance at scale, edge or on-prem deployments where the full Llama 5 footprint does not fit. It is also the pragmatic fine-tuning base when the domain task is narrow and a larger model would be wasted on it.
Limits
- Everything is a few points behind Llama 5. The saving is in hosting cost and latency, not capability — price the GPUs before assuming the small model is cheaper.
- No price or Artificial Analysis index is published here, so Value is Unrated and comparisons against hosted small models like Haiku 4.5 or Luna are apples to oranges.
- Terminal 78 and Tool use 80 make it a poor autonomous agent. It is a throughput model, not an operator.
- The Llama licence is custom, with acceptable-use terms. Verify before commercial redistribution.
Price and access
No published price on this board — access is the weights from Meta's channels, then your own hardware or a third-party host. Hosted rates for Llama-family models vary widely between providers and none are measured on this board. Last scored 15 Sep 2026.
Alternatives on this board
- Mistral Small 3.2 — 79 at $0.09/$0.25 with open weights and Speed 88. The comparable small open model, with a published hosted rate.
- Xiaomi MiMo 2.5 Pro — 82 at $0.43/$0.87 with Speed 83, if hosted inference is acceptable.
- GPT-5.6 Luna — 81 at $0.20/$1.20 with Speed 90 and Context 89, the hosted cost-tier comparison.
- Llama 5 — 84, when the workload is not throughput-bound.
Sources
- Meta AI's model pages — Scout model card, context window and licence.
- The Meta Llama organisation on Hugging Face — weights and licence terms.
- On this site: local LLM inference, September 2026, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.