MDL-8823EST.2026 · IDX.362
Language modelIn production

IBM Granite 4.2 3B

IBM · USA · 2026

The smallest Granite that still reasons - and calls tools as accurately as the model twice its size, though its multilingual scores collapse in a way the shared spec sheet does not warn about.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Granite-4.2-3B is the entry point to IBM's reasoning family, published on 25 August 2026 with the same design as its larger siblings: a dense transformer, chain of thought in <think> tags, three switchable thinking modes and a plain Apache 2.0 licence. At 3.66 billion parameters it is meant for laptops, edge boxes and anywhere a GPU is not guaranteed. One result stands out and it is not the one IBM highlights. On BFCL v4, the standard test of function calling, the 3B scores 52.41 against the 8B's 52.39 - a dead heat with a model more than twice its size. For the common enterprise job of turning a sentence into a correctly formed API call, the small model gives up essentially nothing. Its reasoning holds up better than the size suggests too: 78.33 on AIME25 and 69.71 on LiveCodeBench v6. The collapse is elsewhere, and it matters more to a European reader than any of the headline scores. On MMLU-ProX lite, IBM's own multilingual test, the 3B scores 27.78 while the 8B reaches 61.06. On Arena-Hard-V2 it falls to 34.96 against 65.19. Both models list the same twelve tested languages on the same shared spec sheet, but only one of them behaves that way in practice. The 3B should be treated as an English model that has seen other languages, not as a multilingual one - and Polish is not on IBM's tested list to begin with. The long-document limits are similarly firmer: RULER falls to 67.52 at 64K and 55.30 at 128K, so the 512K extension IBM advertises for the family is not a promise about this size. Where the 3B is genuinely strong is the job it was built for - a cheap, permissively licensed, locally runnable model that can plan a short chain of steps and drive tools without a network round trip. IBM ships GGUF, FP8, MXFP4 and NVFP4 conversions itself, and the quantised repository is downloaded roughly eight times more often than the full weights.

#open-weights#reasoning#on-device#apache-2.0#IBM
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review