GLM-ASR-2512
Z.ai (Zhipu AI) · China · 2025
Z.ai's hosted speech-recognition model, priced at about a quarter of a cent per minute — but capped at 30 seconds of audio per call, which rules out most of what the vendor advertises it for.
GLM-ASR-2512 is the hosted half of Z.ai's speech work, launched on 10 December 2025 alongside the downloadable GLM-ASR-Nano-2512. It is sold through the company's API at $0.03 per million tokens — roughly $0.0024 per minute of audio, about half what OpenAI charges for the model it currently recommends. Z.ai quotes a character error rate of 0.0717 and claims parity with the best speech systems in the world. The language coverage is the strongest argument for it. Beyond Mandarin and English in American and British accents, the model is tuned for Sichuanese, Cantonese, Min Nan and Wu, and handles dozens of other languages including French, German, Japanese, Korean, Spanish and Arabic. It also accepts a custom dictionary: project code names, drug names, uncommon surnames and place names can be loaded once in the request rather than corrected afterwards in every transcript. For mixed Chinese-English speech — the ordinary register of a Shenzhen engineering meeting — the vendor claims a clear lead over competitors. The limitation is where this profile has to be blunt, because the vendor's own two pages contradict each other. The model card advertises six uses, and the first three are real-time meeting minutes, live video captioning and dictated medical records. The API reference for the same model states the hard limits: an audio file of 25 MB or less, and audio duration of 30 seconds or less, per call. That ceiling applies to the streaming mode too — streaming here means the text comes back progressively, not that audio can be fed in continuously. A meeting, a broadcast or a dictated patient history cannot be sent to this endpoint in one piece; the caller has to slice the recording into half-minute fragments and stitch the results, losing sentence context and speaker continuity at every seam. Nothing in the documentation offers a long-form path. That leaves a well-priced, genuinely multilingual model with a shape that fits short utterances: voice input in an app, a command, a single question, a customer-service turn. Anyone transcribing an hour-long recording will find the open sibling — MIT-licensed, running on their own hardware without any per-call duration cap — the more practical choice, which is an unusual thing to say about a free model competing with a paid one from the same company.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!