Kimi-VL A3B Thinking 2506
Moonshot AI · China · 2025
A 16-billion-parameter open vision model that thinks before it answers, runs on about three billion parameters per token and reads screens well enough to drive them.
Kimi-VL A3B Thinking 2506, published on 21 June 2025, is Moonshot AI's small reasoning model for images and video. Its point is the ratio: 16.4 billion parameters in total, but a mixture of experts with 64 experts routes only six per token plus two shared, so generation costs roughly three billion parameters — light enough for a single accelerator, while the model still produces an explicit chain of reasoning before its answer. Images enter through the company's MoonViT encoder at native resolution, up to 3.2 million pixels in a single picture, which is four times the previous version and the reason the model can work on dense screenshots rather than only on photographs. The context window is 131,072 tokens. The 2506 revision is a correction of an earlier trade-off rather than a new model. The first Thinking release was good at reasoning and had visibly worse general perception than Moonshot's non-reasoning Instruct model of the same size. This version closes that gap: it reports MathVision 56.9 (up 20.1 points), MathVista 80.1, MMMU 64.0 and MMMU-Pro 46.3, while matching or beating the Instruct model on ordinary perception — MMBench-EN-v1.1 84.4, MMStar 70.4, RealWorldQA 70.0 — and it does so with roughly twenty per cent shorter reasoning chains. Moonshot also claims the open-model lead on VideoMMMU at 65.2, with Video-MME at 71.9. The screen-control figures are the ones worth noting for agent work: 52.8 on ScreenSpot-Pro, 52.5 on OSWorld-G and 83.2 on V* without external tools. One small inconsistency belongs in the record: the summary at the top of the model card gives MMVet as 78.4, while the comparison table two paragraphs below gives 78.1 for the same model and benchmark. The rest of the card's figures are internally consistent, and the comparison is drawn against Qwen2.5-VL-7B and Gemma 3 12B, with GPT-4o shown only for reference and marked as such. Weights are on Hugging Face under an MIT licence; there is no paid Moonshot endpoint for this model, so it is a download-and-run release. The technical report is arXiv:2504.07491.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!