fullauto.online

← Model leaderboard

Model · stat sheet

Mistral Large 3

Mistral AI·balanced
Overall
83
Rank
#18 / 34
Coding85
Terminal81
Reasoning84
Tool use83
Context82
Speed80
CodingTerminalReasoningTool useContextSpeed

Mistral's flagship; fast, efficient and strong on European languages and tool use.

What it is

Mistral Large 3 is Mistral AI's flagship: the note on this page calls it fast and efficient, strong on European languages and tool use. Mistral is the European frontier lab, and EU jurisdiction and data residency are the reasons organisations shortlist it when the American and Chinese alternatives are both awkward.

Unlike Mistral's smaller releases, the Large tier is a hosted product first; whether weights are published for any part of the line should be verified on Mistral's channels rather than assumed from the brand.

How the composite is built

Overall 83 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Large 3 sits in a five-point band from 80 to 85 — the profile of a balanced tier that intends to do nothing badly.

Coding 85
Level with GLM-5 and three behind Terra's 88. Solid everyday engineering output.
Terminal 81
Its joint-weakest relative axis: two behind Llama 5's 83 and fourteen behind Opus 5's 95. Plan for supervision on anything destructive.
Reasoning 84
Level with Qwen 4 Coder and MiMo 2.5 Pro, two behind Llama 5's 86. Competent planning without the deliberation overhead of the reasoning-specialist lines.
Tool use 83
Level with DeepSeek V4's 83 — consistent with the note's claim, though still seven behind Sonnet 5's 90. Test against your own tool schemas.
Context 82
Level with Gemma 3 27B and Yi-2 Large. The mid-field read; not the axis to choose this model for.
Speed 80
Level with the efficient balanced tiers and well ahead of the frontier cluster. This is where the "fast and efficient" reputation is measurable.

Where it fits

European deployments with data-residency or jurisdiction requirements, multilingual work across EU languages, and general agent loops where the pricing is predictable and the quality ceiling is not the deciding factor. It is a defensible default for a European business that must answer questions about where inference happens.

Limits

  • Value is rated Pricey — the Artificial Analysis index is 9.3, and per $0.75/M blended that is 12.4 index points per dollar: better than Terra's 9.4, a fraction of MiMo 2.5 Pro's 48.1.
  • Nothing here is class-leading. Every axis is mid-board; the frontier tiers beat it across the board at higher prices.
  • Terminal 81 caps autonomous use. Keep a human between the model and production systems.
  • Context 82 is the weakest fit for long-document work — Kimi k3 and the 1M-context tiers are the better answer there.

Price and access

$0.50 per million input tokens and $1.50 per million output, blending to $0.75/M — one of the cheapest flagship-tier rates here, though the low index makes the value column read Pricey. Listed on OpenRouter as well as Mistral's own platform. Mistral's line moves quickly; the board carries unscored Mistral listings such as Medium 3.5 at $1.50/$7.50 and Small 4 at $0.15/$0.60. Last scored 15 Sep 2026.

Alternatives on this board

  • GPT-5.6 Terra — 86 at $2/$12. Three points up at six times the blended rate.
  • DeepSeek V4 — 86 at $0.78/$1.57 with open weights and Coding 92, if jurisdiction is not the constraint.
  • GLM-5 — 82 at $0.60/$1.92 with open weights and three gateways.
  • Xiaomi MiMo 2.5 Pro — 82 at $0.43/$0.87 with Speed 83, the cheaper balanced option.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.