fullauto.online

← Model leaderboard

Model · stat sheet

OLMo 3

Allen AI · open weights·open
Overall
74
Rank
#32 / 34
Coding74
Terminal70
Reasoning75
Tool use72
Context74
Speed82
CodingTerminalReasoningTool useContextSpeed

Allen AI's fully open model — open weights, data and training code end to end.

What it is

OLMo 3 is Allen AI's fully open model — the note on this page is precise about what that means: open weights, open data and open training code, end to end. The Allen Institute for AI publishes the whole recipe, including the data mix and training logs, which makes OLMo the model on this board you can actually study rather than merely run.

That is the case for the model, and the scores are not. OLMo 3 sits at the foot of the board on capability, and it is important to be honest about that: this is a research and audit artefact that happens to be servable, not a competitor to the commercial tiers.

How the composite is built

Overall 74 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. The capability axes run 70 to 75 — the lowest group anywhere here — while Speed 82 reflects a small, efficient model rather than any engineering edge.

Coding 74
One above Falcon 3's 73 at the bottom of the board, and eighteen behind DeepSeek V4's 92. Simple scripts and exercises, not production work.
Terminal 70
The lowest terminal score here alongside Falcon 3's 69. No autonomous shell use.
Reasoning 75
One above Falcon 3's 74 and nine behind Phi-5's 84 — the small-model reasoning outlier it is most often compared against.
Tool use 72
One above Falcon 3's 71 and thirteen behind Haiku 4.5's 85. Expect frequent malformed calls.
Context 74
The joint-lowest context read on this board. Small working sets are a hard limit here.
Speed 82
Its best axis and a real advantage: a small model served on modest hardware is quick by construction. Self-hosted throughput depends on your stack.

Where it fits

Research and education: fine-tuning experiments where you need the full training recipe, ablation studies, teaching materials, audit-friendly baselines for regulated research. If the requirement is that every component of a model — weights, data, code — be inspectable, OLMo is the honest answer on this board and the capability scores are the price of that.

Limits

  • Every capability axis is at or near the board's floor. Do not put it in a production agent loop; it is not built for that.
  • Tool use 72 and Terminal 70 mean even lightweight automation will need heavy guardrails.
  • No price or index is published here, so Value is Unrated — but the budget question for OLMo is GPU time, not tokens.
  • "Fully open" still has a licence. Check the terms attached to the weights and the data before redistribution.

Price and access

No published price or gateway listing on this board — access is the weights, data and training code from Allen AI's channels, run on your own hardware. That is the point of the model rather than a gap: the whole pipeline is downloadable. Last scored 15 Sep 2026.

Alternatives on this board

  • Phi-5 — 78 with Reasoning 84 and Speed 90. If openness is negotiable but hardware is tight.
  • Gemma 3 27B — 79 at $0.08/$0.45 with open weights and a published hosted rate.
  • Mistral Small 3.2 — 79 at $0.09/$0.25, the small open model with the cheapest hosted rate here.
  • Llama 5 Scout — 81, when the deployment needs capability rather than inspectability.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.