Model · stat sheet
Phi-5
Microsoft's small model trained on textbook-quality data; reasoning well beyond its parameter count.
What it is
Phi-5 is Microsoft's small open-weights model, the current line of a family that has always been sold on training data rather than parameter count — the original Phi work was famously built on textbook-quality synthetic data, and the note on this page claims reasoning well beyond the model's size. The six scores mostly support that: Reasoning 84 and Speed 90 sit against Tool use 74 and Context 72.
As with the rest of the Phi line, the interesting engineering is in the data curation. Details of the 5 generation — parameters, data recipe, licence — should be verified on Microsoft's channels.
How the composite is built
Overall 78 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Phi-5 is the most lopsided model in its band: a genuine reasoning score paired with the joint-lowest context read here.
- Coding 79
- Level with Mistral Small 3.2 and Grok 5 Mini. Exercises and small modules, with review.
- Terminal 72
- Two above OLMo 3's 70 and nine behind Haiku 4.5's 81. Not an operator model.
- Reasoning 84
- The claim of the model, and the strongest reasoning score of any small tier on this board — level with Qwen 4 Coder and Mistral Large 3, both much larger propositions.
- Tool use 74
- Eleven behind Haiku 4.5's 85. The axis that suffers in agentic use; a small model reasoning well is not the same as a small model calling tools well.
- Context 72
- Level with Claude Haiku 4.5's 72 — the joint-lowest context read here. Small working sets are the hard constraint.
- Speed 90
- Level with GPT-5.6 Luna's 90 and behind only Haiku 4.5's 94 and Grok 5 Mini's 91. Quick on hosted inference, quicker self-hosted on modest hardware.
Where it fits
Local and edge reasoning: constrained hardware that nonetheless needs judgement rather than lookup, latency-sensitive classification with a reasoning bent, offline tooling. If the task is "think briefly, answer once" and the context is small, Phi-5 is the best value proposition in its class.
Limits
- Tool use 74 and Terminal 72 rule out agentic loops. It is a reasoner, not an operator.
- Context 72 caps the size of any problem it can hold. Chunking is mandatory.
- No price or index is published here, so Value is Unrated — for orientation, the previous Phi release on this board lists at $0.07/$0.14.
- The reasoning score is a small-model score. It beats its class, not the frontier tiers above it.
Price and access
No published price or gateway listing on this board. Access is the weights from Microsoft's channels under their published licence — verify the terms before redistribution — or third-party hosting where Phi-family models are commonly served cheaply. Last scored 15 Sep 2026.
Alternatives on this board
- Grok 5 Mini — 80 with Speed 91, the closed cost-tier comparison.
- Mistral Small 3.2 — 79 at $0.09/$0.25 with open weights and Context 80.
- Claude Haiku 4.5 — 80 at $1/$5 with Tool use 85 — the small model to pick when tools matter.
- Llama 5 Scout — 81 with a flatter profile and no reasoning spike.
Sources
- Microsoft Research's Phi project page — papers and the data-quality research behind the line.
- The Microsoft organisation on Hugging Face — weights and licence terms.
- On this site: local LLM inference, September 2026, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.