fullauto.online

← Model leaderboard

Model · stat sheet

Command R+ 2

Cohere · open weights·open
Overall
80
Rank
#25 / 34
Coding78
Terminal74
Reasoning80
Tool use86
Context88
Speed76
CodingTerminalReasoningTool useContextSpeed

Cohere's RAG- and tool-tuned open weights; built for retrieval and long grounded context.

What it is

Command R+ 2 is Cohere's retrieval- and tool-tuned flagship. Cohere has spent the whole Command R line optimising for a narrow set of things — grounded generation with inline citations, multi-step tool calling, multilingual retrieval — rather than chasing general capability, and this score sheet is the clearest illustration of that strategy on the board.

Its two best axes are Context (88) and Tool use (86). Its two worst are Terminal (74) and Coding (78). No other model here has a profile that lopsided in that direction.

How the composite is built

Overall 80 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Command R+ 2 is penalised badly by that shape: 44% of the index goes to the two things Cohere deliberately did not optimise.

Coding 78
Level with Amazon Nova 2 Pro and fourteen behind DeepSeek V4. Not a coding model and not sold as one.
Terminal 74
Joint fourth lowest on the scored board, above only Falcon 3 (69), OLMo 3 (70) and Phi-5 (72). Do not put this model in a shell loop.
Reasoning 80
Mid-table. Enough to decide which document answers a question; not enough to reason its way out of a contradiction in the documents.
Tool use 86
The highest of any open-weights model on this board, ahead of Nemo 3 Ultra (85), Llama 5 (84) and DeepSeek V4 (83) — and ahead of Claude Haiku 4.5's 85. This is what Cohere built and it shows.
Context 88
Level with Claude Sonnet 5 and Grok 5, from a model seven or eight points below them overall. Long grounded context is the other half of the retrieval story.
Speed 76
Middling. Fast enough for a document-answering service, slow enough that you would not put it on a per-keystroke path.

Where it fits

Retrieval-augmented generation done properly: answering over an internal corpus with citations you can show a user, multilingual knowledge bases, tool-calling agents that query APIs rather than write code. Cohere's surrounding models — embeddings and reranking — are part of the case; the generation model is the last stage of a pipeline the same vendor supplies end to end.

Limits

  • The composite understates it for RAG and overstates it for everything else. If your workload is retrieval, weight Tool use and Context and ignore the 80.
  • Check the licence before you deploy. Cohere's open-weights releases have historically been non-commercial, with commercial use running through their paid API or an enterprise agreement. Confirm the terms for this release rather than assuming the open-weights tag on this board means freely usable.
  • No published price here, so no value rating — the dash is missing data. Cohere prices per token on their own platform and is also available through cloud marketplaces.
  • Version naming is worth verifying. Cohere's Command releases are dated rather than neatly numbered; check which checkpoint you are actually calling.

Access

Cohere's own platform and enterprise deployments, with weights published for research use. This board records no OpenRouter or OpenCode listing. Cohere's differentiator on the enterprise side has been private and on-premises deployment — running the whole retrieval stack inside a customer's boundary — which is usually the reason a regulated buyer ends up here rather than with a larger vendor. Last scored 15 Sep 2026.

Alternatives on this board

  • Gemini 3.6 Flash — 84 at $0.75/$3.75 with Context 93 and Speed 88. Better at the same job if you do not need on-premises deployment.
  • Claude Haiku 4.5 — Tool use 85 and Speed 94 at $1/$5, for tool-calling at volume where the context stays under 200k.
  • Mistral Large 3 — 83 at $0.50/$1.50, the other European vendor with a serious private-deployment story.
  • Llama 5 — 84 with Tool use 84 and open weights, if you want a general model rather than a retrieval specialist.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.