MDL-4140EST.2026 · IDX.296
AudioIn production

IBM Granite Speech 4.1 2B

IBM · USA · 2026

The most downloaded model IBM has ever published is not a chatbot but a transcriber - and its version number grew while the model itself did not.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Granite Speech 4.1 2B, published on 29 April 2026, converts speech into text and translates it in both directions between English, French, German, Spanish, Portuguese and Japanese. It was trained on 174,000 hours of audio from public corpora, plus synthetic material aimed at Japanese, keyword-prompted recognition and translation. The whole thing is Apache 2.0, with no territorial exclusion and no registration gate. The naming deserves a warning, because it is the sort of thing that misleads a buyer. This model is the successor to granite-4.0-1b-speech and has exactly the same parameter count. IBM changed the convention: the number now describes the real size of the finished system, 2.3 billion parameters, instead of the language model it was built on. Nothing doubled. Only the label did. What actually improved is worth more than the number. The encoder now carries two classification heads at once - one over characters, one over the tokeniser's subwords - and a single prompt change makes the model add punctuation and capitalisation. That sounds cosmetic until you use the output: German noun capitalisation reaches a 99.5 capitalisation F1, and a transcript that already has commas and full stops does not need a second model to clean it up. Keyword biasing lets an operator hand the model a list of names, acronyms or product codes before transcription, which is where general recognisers usually fail. One fact says more about the market than any benchmark. This model is downloaded roughly 279,000 times a month - more than any language model IBM publishes except the 8B flagship, and more than double its own vision sibling. The specialised transcriber, not the chatbot, is what people run. For a Polish reader the limitation is hard: Polish is not among the six supported languages, and IBM makes no claim about it. Two sibling variants exist for narrower jobs - a plus edition that labels who is speaking and adds word-level timestamps, and a non-autoregressive edition built for throughput.

#open-weights#speech recognition#transcription#translation#apache-2.0#IBM
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review