Kimi K2 Thinking
Moonshot AI · China · 2025
The thinking branch of Kimi K2, published in November 2025: the same trillion-parameter body taught to reason step by step and to hold a tool chain across 200-300 consecutive calls, shipped natively in INT4 with a 256K window.
Kimi K2 Thinking is what Moonshot AI built on top of Kimi K2 four months after that model's release: the July version answers on reflex, this one reasons step by step while calling tools as it goes. The weights appeared on 4 November 2025 and resellers began serving the model on 6 November. The body is unchanged — one trillion total parameters, 32 billion activated, 384 experts, a 160K vocabulary — with two engineering differences that matter in practice. The context window grows to 256K, and the model is released natively quantised to INT4, with quantisation-aware training applied during post-training; the maker reports a twofold speed-up in low-latency mode with no loss of quality, which is the opposite of the usual trade-off, where quantisation is something the user does afterwards and pays for in accuracy. The capability Moonshot puts first is endurance rather than raw score: stable tool use across 200 to 300 sequential calls. That is the difference between a model that can use a tool and a model that can run a long task without losing the thread. In the maker's tables the model leads on agentic search — BrowseComp 60.2 against 54.9 for GPT-5 (High) and 24.1 for Claude Sonnet 4.5 Thinking — and on Humanity's Last Exam with tools it reports 44.9 against 41.7 for GPT-5. On coding it stays behind both: SWE-bench Verified 71.3 against 74.9 and 77.2. The comparison against its own predecessor is the sharpest: K2 0905 scores 7.4 on BrowseComp where this model scores 60.2. One footnote deserves repeating, because it is unusually candid. Moonshot states that on HLE the model reaches 51.3 if access to Hugging Face is left open during the test, and that it blocked that access itself to avoid benchmark leakage. The published 44.9 is therefore the lower, harder number, by the maker's own choice.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!