MDL-4083EST.2025 · IDX.509
AudioIn production

GLM-ASR-Nano-2512

Z.ai (Zhipu AI) · China · 2025

A 1.5-billion-parameter open speech-recognition model from Z.ai, MIT-licensed, tuned for Cantonese and for speech too quiet for other models to catch.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

GLM-ASR-Nano-2512 is Z.ai's open speech-recognition model, announced on 10 December 2025 and published under the MIT licence with 1.5 billion parameters. It is the downloadable half of the company's speech work: the hosted sibling, GLM-ASR-2512, is sold through the Z.ai API at $0.03 per million tokens, roughly $0.0024 per minute of audio. Its model card makes a direct claim: the 1.5B model outperforms Whisper large-v3 on several benchmarks while staying small enough to run on a laptop. Z.ai reports the lowest average error rate among comparable open models (4.10) and particularly strong results on Chinese meeting and read-speech sets. The hosted variant is quoted separately at a character error rate of 0.0717. Two capabilities are unusual enough to be the reason to pick it. The first is dialect coverage: beyond Mandarin and English it is explicitly tuned for Cantonese, Sichuanese, Min Nan and Wu — dialects that general-purpose models transcribe badly or not at all. The second is quiet speech. The model was trained on deliberately low-volume audio, the kind that other systems drop as silence: dictation in a shared office, a whispered aside in a recording, an interview where the second speaker sits away from the microphone. The timing matters for anyone building on speech. OpenAI deprecated its own hosted whisper-1 endpoint on 26 August 2026 and removes it on 26 February 2027; the open Whisper weights survive, but the convenient hosted version does not. GLM-ASR-Nano arrived a year earlier as a smaller, MIT-licensed model claiming to beat those weights — and with about 61,000 downloads in a thirty-day window, it is being taken up at a rate that has nothing to do with anyone's price list.

#speech recognition#open source#multilingual#transcription
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review