MDL-3220EST.2026 · IDX.024
Language modelIn production

IBM Granite Vision 4.1 4B

IBM · USA · 2026

A four-billion-parameter model built for one unglamorous job - turning charts and tables printed in a document back into data - and it does it at frontier level.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Granite Vision 4.1 4B, published on 29 April 2026, is the only member of the Granite family that sees. It is not a general image chatbot, and IBM does not pretend otherwise: it was built to pull structure out of documents. Point it at a scanned bar chart and it returns the CSV table behind it, a written summary of it, or the Python code that would redraw it. Point it at a page of a PDF and it returns the tables as JSON, HTML or OTSL markup. Each task is switched by a short tag in the prompt, so no prompt engineering is needed to get machine-readable output. The size is the point. Under the hood sit a 3.4-billion-parameter Granite 4.1 language model and a 0.6-billion-parameter SigLIP2 vision encoder, tiling every input image into 384-pixel squares and encoding each tile separately. Visual features are compressed fourfold and then injected into the language model at eight separate points rather than pasted in at the front - a design that keeps the visual detail alive deep in the network without inflating the token count. On VAREX, a benchmark for pulling named values out of documents, it reaches 94.2 percent exact-match accuracy without any examples, a score IBM positions against models many times larger. For a reader deciding whether to use it, two limits matter more than the benchmarks. First, it accepts English instructions only - the images can contain anything, but the conversation is English, and there is no Polish claim. Second, the licence is plain Apache 2.0 with no territorial exclusion, so a European company processing its own invoices or reports faces no legal obstacle here, which is not true of every open vision model. It is downloaded roughly 111,000 times a month, second only to the speech models among IBM's specialised releases, and it plugs directly into Docling, IBM's open document-conversion toolkit. The dataset behind the chart work, ChartNet, was published alongside it.

#open-weights#vision#document understanding#apache-2.0#IBM
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review