The reasoning model that started Z.ai's vision line: nine billion parameters that matched a seventy-two billion rival on eighteen of twenty-eight tests.
GLM-4.1V-9B-Thinking is the model that turned Z.ai's vision-language work into a reasoning line. It is built on the company's own GLM-4-9B-0414 text model, extended with a vision encoder and then trained with reinforcement learning to reason step by step before answering — the same thinking paradigm the company later carried into GLM-4.5V and GLM-4.6V.
The headline claim is about size. Z.ai reports that on 28 benchmark tasks the model was best in class among models around ten billion parameters on 23 of them, and beat Alibaba's Qwen2.5-VL-72B — a model eight times larger — on 18. That claim comes from the publisher's own evaluation, but the download counts suggest the market took it seriously: it remains one of the most downloaded models the company has ever released.
One number deserves a footnote. The name says nine billion, and the text half of the model is indeed the 9B GLM-4; the published weight file totals 10.3 billion parameters because the vision encoder is counted too. Anyone sizing hardware should plan against the larger figure.
The model reads images at arbitrary aspect ratios up to 4K, handles video, works in Chinese and English, and holds a 64k-token context. The weights are MIT-licensed, and the base model without the reasoning training was published alongside it.
#vision language#reasoning#open weights#small model