fullauto.online

← Model leaderboard

Model · stat sheet

Kimi k3

Moonshot AI · open weights·open
Overall
84
Rank
#14 / 34
Coding87
Terminal80
Reasoning86
Tool use80
Context92
Speed74
CodingTerminalReasoningTool useContextSpeed

Moonshot's long-context specialist; enormous windows and steady reasoning, open weights.

What it is

Kimi k3 is Moonshot AI's long-context specialist: open weights from a Beijing lab whose Kimi line has been built around large windows since the K2 generation. The note on this page pairs enormous windows with steady reasoning, and the Context score of 92 is the number that carries the model.

Moonshot publishes weights and a hosted API together — the same pattern as DeepSeek and Zhipu. Verify the licence text and the exact window before building on either route; those details are in the linked repositories, not on this board.

How the composite is built

Overall 84 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. The spread is 74 to 92 — Context leads by a distance, and Terminal and Tool use lag at 80.

Coding 87
One behind Qwen 4 Max's 88 and five behind DeepSeek V4's 92. Strong for an open model whose selling point is context rather than code.
Terminal 80
Three behind Llama 5's 83 and fifteen behind Opus 5's 95. Long autonomous shell sessions should stay supervised.
Reasoning 86
Level with GLM-5 and Llama 5, seven behind DeepSeek R2's 93. Planning is solid for the price band.
Tool use 80
Level with GLM-5 and Qwen 4 Coder, ten behind Sonnet 5's 90. The weaker half of the profile; budget for schema validation and retries.
Context 92
Its best axis and the whole case for the model. Behind only Gemini 3.6 Pro's 98 and the 95s of the 1M-context Claude tiers, and two ahead of Terra's 90.
Speed 74
Level with DeepSeek V4 and GLM-5 on hosted inference. Long-context prompts are slow by nature; self-hosting makes the number a property of your hardware.

Where it fits

Long-document work: contract and literature review, whole-repository reads, research synthesis, anything where the alternative is chunking and retrieval. Because the weights are open, it is also a candidate where the data cannot leave your infrastructure but the context requirement is large.

Limits

  • Value is rated Poor — 43.6 index points per $6.00/M blended is 7.3, the weakest ratio among the priced models here. Context 92 is expensive.
  • Terminal 80 and Tool use 80 cap agentic autonomy. The model reads more than it operates.
  • Hosted inference is operated from China. For regulated workloads, use the weights or a third-party host rather than the first-party endpoint.
  • Moonshot ships frequent point releases, and the board carries several unscored Kimi listings — K2.6 at $0.95/$4 and K2.7 Code at $0.66/$3.30 among them. Check what is current before buying.

Price and access

$3 per million input tokens and $15 per million output on the hosted API, blending to $6.00/M — the highest blended rate among the open-weights models scored here. Listed on OpenRouter, OpenCode Zen and OpenCode Go, with weights published for self-hosting. Last scored 15 Sep 2026.

Alternatives on this board

  • DeepSeek V4 — 86 at $0.78/$1.57 with Coding 92. Six times cheaper blended and two points up overall, with Context 86.
  • Qwen 4 Max — 85 with Context 87 and Reasoning 87, the multilingual alternative from Alibaba.
  • GLM-5 — 82 at $0.60/$1.92 with the same three-gateway reach and open weights, at a sixth of the blended rate.
  • Gemini 3.6 Pro — 89 with Context 98, if closed weights are acceptable and context is the priority.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.