IBM Granite Vision 4.1 4B
IBM · USA · 2026
A four-billion-parameter model built for one unglamorous job - turning charts and tables printed in a document back into data - and it does it at frontier level.
Granite Vision 4.1 4B, published on 29 April 2026, is the only member of the Granite family that sees. It is not a general image chatbot, and IBM does not pretend otherwise: it was built to pull structure out of documents. Point it at a scanned bar chart and it returns the CSV table behind it, a written summary of it, or the Python code that would redraw it. Point it at a page of a PDF and it returns the tables as JSON, HTML or OTSL markup. Each task is switched by a short tag in the prompt, so no prompt engineering is needed to get machine-readable output. The size is the point. Under the hood sit a 3.4-billion-parameter Granite 4.1 language model and a 0.6-billion-parameter SigLIP2 vision encoder, tiling every input image into 384-pixel squares and encoding each tile separately. Visual features are compressed fourfold and then injected into the language model at eight separate points rather than pasted in at the front - a design that keeps the visual detail alive deep in the network without inflating the token count. On VAREX, a benchmark for pulling named values out of documents, it reaches 94.2 percent exact-match accuracy without any examples, a score IBM positions against models many times larger. For a reader deciding whether to use it, two limits matter more than the benchmarks. First, it accepts English instructions only - the images can contain anything, but the conversation is English, and there is no Polish claim. Second, the licence is plain Apache 2.0 with no territorial exclusion, so a European company processing its own invoices or reports faces no legal obstacle here, which is not true of every open vision model. It is downloaded roughly 111,000 times a month, second only to the speech models among IBM's specialised releases, and it plugs directly into Docling, IBM's open document-conversion toolkit. The dataset behind the chart work, ChartNet, was published alongside it.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!