MDL-4091EST.2026 · IDX.229
Language modelIn production

IBM Granite Embedding 311M Multilingual R2

IBM · USA · 2026

The one IBM model that was explicitly trained on Polish: 311 million parameters, 52 languages, a 32,000-token window and vectors you can shorten from 768 numbers to 128 when storage costs more than accuracy.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Granite-Embedding-311M-Multilingual-R2 turns text into 768-number vectors for search and retrieval. It does not write, answer or reason — but for readers outside the English-speaking world it is the most relevant model IBM currently publishes, because it is the only one in this catalogue whose training explicitly covers Polish. That point deserves emphasis. Every Granite chat model we describe lists twelve tested languages, and Polish is in none of them. This model lists 52 languages that received explicit retrieval-pair and cross-lingual training — Polish among them, alongside Czech, Slovak, Ukrainian, Lithuanian and most of the region — on top of an encoder pretrained on more than 200. A Polish company that wants to search its own documents with an IBM model has exactly one option, and this is it. The jump over the previous generation is unusually large: 65.2 on multilingual MTEB retrieval against 52.2 for granite-embedding-278m-multilingual, thirteen points in one generation, and an average of 56.3 across all retrieval benchmarks against 41.8. Long-document search nearly doubles, from 37.7 to 71.7. IBM attributes this to the switch to the ModernBERT architecture and a context window that grew from 512 tokens to 32,768. The practical feature is Matryoshka: the 768-number vector can be cut to 512, 384, 256 or even 128 numbers without re-running the model, and the catalogue of scores shows what that costs. Trimming from 768 to 256 loses one point on English retrieval and half a point on multilingual — a six-fold saving in vector storage for an error most search systems will not notice. At 128 numbers the loss becomes visible, about two points. A smaller sibling exists for latency-sensitive work: granite-embedding-97m-multilingual-r2, 97 million parameters, 384-number vectors, five points lower on multilingual retrieval and about 40 percent faster. This model encodes about 1,828 documents per second on one H100; the small one manages 2,534. Licence is plain Apache 2.0, with no territorial exclusion — worth noting in a year when several Chinese and American labs added Europe-specific carve-outs to their weights. ONNX and OpenVINO conversions ship alongside the PyTorch weights.

#open-weights#apache-2.0#IBM#multilingual#long-context
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review