North Micro Vision Instruct
Cohere · Canada · 2026
Cohere's smallest vision-language model: 2.4 billion parameters under Apache 2.0, built to read a full A4 page at its native resolution rather than a downscaled copy of it.
North Micro Vision Instruct is the smallest vision-language model Cohere has published, released on 12 August 2026 under the Apache 2.0 licence with no conditions attached to commercial use. It has 2.4 billion parameters in total: a 2-billion-parameter in-house language model paired with a 400-million-parameter vision encoder trained from Google's SigLIP 2 SO400M. The design goal is documents. Most compact vision models shrink an image to a fixed square before looking at it, which is fatal for small print; this one accepts native resolution up to 1654 × 2339 pixels — an A4 page scanned at 200 dpi — and preserves the aspect ratio. Spatial structure is kept through a combination of two-dimensional rotary embeddings and learned one-dimensional positions, and visual patches are injected into several layers of the language model rather than only the first. The language backbone alternates three sliding-window attention layers with one global layer, which is what allows a 128,000-token context without the memory cost that a fully global model of this size would carry. Cohere is unusually candid about the limit of that number: the validated operating range for prompts that contain images is 8,000 tokens, and anything beyond that relies on extrapolation the company has not benchmarked. Eleven languages are declared, English through Arabic; Polish is not among them. Cohere's own evaluation places the model at the top of its size class on document work — 0.921 on DocVQA, 0.808 on ChartQA, 0.775 on AI2D and 0.732 average on RefCOCO grounding — while conceding losses elsewhere: 0.687 on MMBench against 0.731 for the larger Phi-3.5-vision, and 0.725 on CountBench against 0.764 for SmolVLM. These are vendor figures. Training ran in four stages, from 10 million fixed-resolution examples up to instruction tuning on 50 million and a final preference-optimisation pass on 500,000. It arrived in the same week as Liquid AI's LFM2.5-VL-3B, a comparably sized model aimed at the same on-device niche. The technical difference is one of emphasis — Cohere optimises for reading a page, Liquid AI for reading a screen at speed — but the licensing difference is sharper: Apache 2.0 here, against a licence with a revenue ceiling there.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!