LFM2.5-VL-3B
Liquid AI · USA · 2026
A 3.1-billion-parameter vision-language model built to run on a phone or laptop: it reads screens, documents and charts, points at objects by coordinates and calls tools, in about 3 GB of memory.
LFM2.5-VL-3B is an open-weights vision-language model published by Liquid AI on 12 August 2026. It is deliberately small — 3.12 billion parameters, bfloat16 weights, roughly 3 GB in memory — and the whole design points in one direction: running the model on the device in front of the user rather than in a data centre. The architecture is a hybrid rather than a conventional transformer. Of its 30 language layers, most are short-convolution blocks and only every third or fourth is a full-attention block, which is what keeps memory growth in check on long inputs. Vision is handled by a SigLIP2 encoder that splits an image into 512-pixel tiles, up to ten of them, and hands the language model at most 256 image tokens per tile. The text context window is 128,000 tokens and the vocabulary was doubled to 128,000 entries, which is what makes the model usable outside the Latin alphabet; sixteen languages are declared, Polish among them. What the model is for is narrower than the usual chatbot pitch. It reads user interfaces across mobile, web and desktop, grounds objects to bounding-box coordinates, extracts text and figures from documents and charts, accepts several images at once and can issue tool calls from either a text or an image prompt. Liquid AI reports an average of 69.4 across 28 vision benchmarks — level with InternVL-3.5-4B and 0.7 behind Qwen3.5-4B, both larger — and 80.7 on ScreenSpot-v2 for interface understanding. These are vendor figures and have not been independently reproduced. The speed claims are the point of the release. Liquid AI measures 228 tokens per second on an Apple M5 Max, 116 on an AMD Ryzen AI Max+ 395 and 20 on a Galaxy S26 Ultra — that last figure being a phone, unplugged from any cloud. On an H100 the model reaches first token in about 34 milliseconds on five-frame video. One caveat is worth stating plainly, because the company's own announcement does not. The weights are published under the LFM Open License v1.0, not a standard open-source licence. Downloading, fine-tuning, redistributing and research use are free; commercial use, however, is granted only to entities whose annual revenue stays below 10 million US dollars, with an exemption for registered non-profits using the model for research. Above that threshold the licence does not cover commercial deployment at all and a separate agreement with Liquid AI is required. For a hobbyist, a student or a small studio the model is effectively free; for a large company it is not.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!