GLM-4V-9B
Z.ai (Zhipu AI) · China · 2024
Z.ai's first open-weight vision model: reads 1120x1120 images, and its strongest single result is text recognition, not general reasoning.
GLM-4V-9B is the vision variant of the GLM-4 generation, published on 4 June 2024 alongside the text-only GLM-4-9B-Chat. It is the model that opened the producer's vision line, which today runs through GLM-4.1V-9B-Thinking to GLM-4.6V. It holds a bilingual Chinese-English conversation about images at 1120x1120 pixels, across several turns. One number in the name is worth correcting before anyone plans hardware around it. The published weights come to 13.9 billion parameters, not nine: the text tower is the 9-billion GLM-4 model with 40 layers, hidden size 4096 and 32 attention heads over 2 key-value heads, and on top of it sits a separate image encoder of 63 layers with hidden size 1792 that the name does not count. The context window is 8192 tokens, the shortest in this profile's whole neighbourhood. The model card claims the model beats GPT-4-turbo of April 2024, Gemini 1.0 Pro, Qwen-VL-Max and Claude 3 Opus, and the table below the claim backs that up. The same table also shows GPT-4o of May 2024 ahead of it in almost every column — 69.2 against 47.2 on the academic MMMU set, 63.9 against 58.7 on MMStar. Read column by column, the honest summary is narrower and more useful: this is a model that reads text in pictures unusually well for its size. Its 786 points on OCRBench are the highest figure in the producer's own comparison, above GPT-4o's 736, while on subject knowledge it sits a generation behind. The licence is the GLM-4 document, not an open-source one: derivative work must carry a 'Built with glm-4' notice and its name must begin with 'glm-4'.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!