MDL-8649EST.2025 · IDX.395
Language modelIn production

GLM-4.6V-Flash

Z.ai (Zhipu AI) · China · 2025

The small twin of Z.ai's vision flagship: ten times fewer parameters, free in the API, and downloaded twenty times more often.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

GLM-4.6V-Flash is the small member of the vision pair Z.ai published on 7 December 2025. Both models went out the same day, both carry an MIT licence, and both can be downloaded by anyone. The difference is size: the flagship GLM-4.6V holds 107.7 billion parameters, this one 10.3 billion. That difference decides who can actually use them, and the download counters show it. In a thirty-day window the small model was pulled about 82,600 times against roughly 4,000 for the flagship — twenty times more, even though neither costs anything to download. The reason is hardware, not price: 10 billion parameters fit on a single accelerator, and quantised community builds of this model run on a laptop, while 107 billion parameters need a server. Whichever model is technically better, only one of them is within reach of the people doing the downloading. In the API the same model is billed at zero. Z.ai keeps exactly three free text tiers — this one, GLM-4.5-Flash and GLM-4.7-Flash — and prices the paid vision line above them, with GLM-4.6V at $0.30 per million input tokens. The technical envelope is not a cut-down one. The context window is 131,072 tokens, the same order as the paid models, and the vision encoder is a full 24-layer stack at 336-pixel input with temporal patching, meaning the model reads video frames and not only stills. The text side is a dense 40-layer stack, 4,096 wide, on Z.ai's usual 151,552-token vocabulary.

#open weights#MIT license#vision language model#free tier#single GPU
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review