GLM-4.5V
Z.ai (Zhipu AI) · China · 2025
Z.ai's first widely used vision model: 106 billion parameters that read images, video, documents and computer screens.
GLM-4.5V, published in August 2025, is the model that gave Z.ai's line eyes. It is a mixture-of-experts network with 106 billion total parameters and 12 billion active per token — the same skeleton as GLM-4.5-Air, which is why the two weigh almost exactly the same on disk — retrained to accept images, video, documents and text in a single request. The four workloads Z.ai names are ordinary but demanding: describing and reasoning about photographs, understanding video, parsing documents, and operating graphical interfaces. That last one matters most for where the industry went next: a model that can look at a screenshot and say which button to press is the missing half of a computer-using agent. At release the company claimed state-of-the-art results among open-source vision-language models of comparable size. The practical limits are worth knowing before choosing it. The context window is 64,000 tokens, half of what the text models of the same generation offer, and the answer is capped at 16,000 tokens — this is a model for reading, not for writing long reports. Pricing is $0.60 per million input tokens and $1.80 per million output. The weights are on Hugging Face under an MIT licence and have been downloaded around 97,000 times. Its successor, GLM-4.6V, ships weight files of identical size — the same network, retrained — at half the price and with twice the context, which makes GLM-4.5V mostly a historical reference today rather than a purchase.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!