ChatGLM2-6B
Z.ai (Zhipu AI) · China · 2023
The 2023 six-billion-parameter chat model that still pulls more monthly downloads than any current flagship on its maker's account.
ChatGLM2-6B arrived in June 2023, three months after the first generation, and it is the model that made the family practical rather than merely impressive. Three changes did the work: pre-training on 1.4 trillion bilingual tokens, a context window stretched from 2,048 to 32,768 tokens using FlashAttention, and multi-query attention, which shares two key-value heads across 32 query heads. The last one is the reason the model is still in use: the vendor measured 42 per cent faster inference than the first generation, and at INT4 quantisation the supported conversation length on a 6 GB card went from about 1,000 tokens to 8,000. The gains over the first generation were reported as percentages rather than absolute scores: MMLU up 23 per cent, CEval up 33 per cent, BBH up 60 per cent and GSM8K up 571 per cent. The last figure says more about how weak the first model was at arithmetic than about how strong the second one is. One caveat the vendor states plainly and this catalogue repeats: the 32K context belongs to the base model, while the dialogue alignment was trained at 8K, and the release notes admit limited comprehension of single very long documents. The download counter is the most interesting number here. Three years after release, this repository is fetched more often in a rolling month than GLM-5, GLM-5.1 or GLM-4.7 - the company's own current models. Part of that is inertia in tutorials and course material, part is that a 12 GB file still fits hardware that a 350-billion-parameter flagship never will. The licence is a bespoke vendor document, not an open-source licence: free for academic research, free for commercial use only after registering through the vendor's form.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!