fullauto.online

← Model leaderboard

Model · stat sheet

Phi-5

Microsoft · open weights·open
Overall
78
Rank
#30 / 34
Coding79
Terminal72
Reasoning84
Tool use74
Context72
Speed90
CodingTerminalReasoningTool useContextSpeed

Microsoft's small model trained on textbook-quality data; reasoning well beyond its parameter count.

What it is

Phi-5 is Microsoft's small open-weights model, the current line of a family that has always been sold on training data rather than parameter count — the original Phi work was famously built on textbook-quality synthetic data, and the note on this page claims reasoning well beyond the model's size. The six scores mostly support that: Reasoning 84 and Speed 90 sit against Tool use 74 and Context 72.

As with the rest of the Phi line, the interesting engineering is in the data curation. Details of the 5 generation — parameters, data recipe, licence — should be verified on Microsoft's channels.

How the composite is built

Overall 78 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Phi-5 is the most lopsided model in its band: a genuine reasoning score paired with the joint-lowest context read here.

Coding 79
Level with Mistral Small 3.2 and Grok 5 Mini. Exercises and small modules, with review.
Terminal 72
Two above OLMo 3's 70 and nine behind Haiku 4.5's 81. Not an operator model.
Reasoning 84
The claim of the model, and the strongest reasoning score of any small tier on this board — level with Qwen 4 Coder and Mistral Large 3, both much larger propositions.
Tool use 74
Eleven behind Haiku 4.5's 85. The axis that suffers in agentic use; a small model reasoning well is not the same as a small model calling tools well.
Context 72
Level with Claude Haiku 4.5's 72 — the joint-lowest context read here. Small working sets are the hard constraint.
Speed 90
Level with GPT-5.6 Luna's 90 and behind only Haiku 4.5's 94 and Grok 5 Mini's 91. Quick on hosted inference, quicker self-hosted on modest hardware.

Where it fits

Local and edge reasoning: constrained hardware that nonetheless needs judgement rather than lookup, latency-sensitive classification with a reasoning bent, offline tooling. If the task is "think briefly, answer once" and the context is small, Phi-5 is the best value proposition in its class.

Limits

  • Tool use 74 and Terminal 72 rule out agentic loops. It is a reasoner, not an operator.
  • Context 72 caps the size of any problem it can hold. Chunking is mandatory.
  • No price or index is published here, so Value is Unrated — for orientation, the previous Phi release on this board lists at $0.07/$0.14.
  • The reasoning score is a small-model score. It beats its class, not the frontier tiers above it.

Price and access

No published price or gateway listing on this board. Access is the weights from Microsoft's channels under their published licence — verify the terms before redistribution — or third-party hosting where Phi-family models are commonly served cheaply. Last scored 15 Sep 2026.

Alternatives on this board

  • Grok 5 Mini — 80 with Speed 91, the closed cost-tier comparison.
  • Mistral Small 3.2 — 79 at $0.09/$0.25 with open weights and Context 80.
  • Claude Haiku 4.5 — 80 at $1/$5 with Tool use 85 — the small model to pick when tools matter.
  • Llama 5 Scout — 81 with a flatter profile and no reasoning spike.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.