Model · stat sheet
Command R+ 2
Cohere's RAG- and tool-tuned open weights; built for retrieval and long grounded context.
What it is
Command R+ 2 is Cohere's retrieval- and tool-tuned flagship. Cohere has spent the whole Command R line optimising for a narrow set of things — grounded generation with inline citations, multi-step tool calling, multilingual retrieval — rather than chasing general capability, and this score sheet is the clearest illustration of that strategy on the board.
Its two best axes are Context (88) and Tool use (86). Its two worst are Terminal (74) and Coding (78). No other model here has a profile that lopsided in that direction.
How the composite is built
Overall 80 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Command R+ 2 is penalised badly by that shape: 44% of the index goes to the two things Cohere deliberately did not optimise.
- Coding 78
- Level with Amazon Nova 2 Pro and fourteen behind DeepSeek V4. Not a coding model and not sold as one.
- Terminal 74
- Joint fourth lowest on the scored board, above only Falcon 3 (69), OLMo 3 (70) and Phi-5 (72). Do not put this model in a shell loop.
- Reasoning 80
- Mid-table. Enough to decide which document answers a question; not enough to reason its way out of a contradiction in the documents.
- Tool use 86
- The highest of any open-weights model on this board, ahead of Nemo 3 Ultra (85), Llama 5 (84) and DeepSeek V4 (83) — and ahead of Claude Haiku 4.5's 85. This is what Cohere built and it shows.
- Context 88
- Level with Claude Sonnet 5 and Grok 5, from a model seven or eight points below them overall. Long grounded context is the other half of the retrieval story.
- Speed 76
- Middling. Fast enough for a document-answering service, slow enough that you would not put it on a per-keystroke path.
Where it fits
Retrieval-augmented generation done properly: answering over an internal corpus with citations you can show a user, multilingual knowledge bases, tool-calling agents that query APIs rather than write code. Cohere's surrounding models — embeddings and reranking — are part of the case; the generation model is the last stage of a pipeline the same vendor supplies end to end.
Limits
- The composite understates it for RAG and overstates it for everything else. If your workload is retrieval, weight Tool use and Context and ignore the 80.
- Check the licence before you deploy. Cohere's open-weights releases have historically been non-commercial, with commercial use running through their paid API or an enterprise agreement. Confirm the terms for this release rather than assuming the open-weights tag on this board means freely usable.
- No published price here, so no value rating — the dash is missing data. Cohere prices per token on their own platform and is also available through cloud marketplaces.
- Version naming is worth verifying. Cohere's Command releases are dated rather than neatly numbered; check which checkpoint you are actually calling.
Access
Cohere's own platform and enterprise deployments, with weights published for research use. This board records no OpenRouter or OpenCode listing. Cohere's differentiator on the enterprise side has been private and on-premises deployment — running the whole retrieval stack inside a customer's boundary — which is usually the reason a regulated buyer ends up here rather than with a larger vendor. Last scored 15 Sep 2026.
Alternatives on this board
- Gemini 3.6 Flash — 84 at $0.75/$3.75 with Context 93 and Speed 88. Better at the same job if you do not need on-premises deployment.
- Claude Haiku 4.5 — Tool use 85 and Speed 94 at $1/$5, for tool-calling at volume where the context stays under 200k.
- Mistral Large 3 — 83 at $0.50/$1.50, the other European vendor with a serious private-deployment story.
- Llama 5 — 84 with Tool use 84 and open weights, if you want a general model rather than a retrieval specialist.
Sources
- Cohere's model documentation — current Command checkpoints, context limits and capabilities.
- Cohere's retrieval-augmented generation guide — how grounding and citations actually work, which is the point of this model.
- The Cohere Labs weights on Hugging Face for licence terms.
- On this site: context engineering, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.