MDL-5381EST.2024 · IDX.421
category.visionIn production

PaliGemma

Google · USA · 2024

Google's first open vision-language model: a 3-billion-parameter base built to be fine-tuned, shipped with fifty-four ready checkpoints for individual benchmarks — and still the most downloaded model of its line.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

PaliGemma is the first vision-language model in the Gemma family, announced at Google I/O on 14 May 2024. Architecturally it is a deliberately simple recipe: the SigLIP-So400m image encoder feeding the Gemma 2B language model, an approach inspired by Google's earlier PaLI-3. The weights file counts 2.92 billion parameters. Image input comes at one of three resolutions — 224, 448 or 896 pixels square — and the model answers in text: captions for images and short video, visual question answering, reading text inside a photograph, object detection and segmentation. The unusual thing about this release is not the model but what Google published around it. PaliGemma was never meant to be used straight out of the box; it is a base for transfer, and the accompanying technical report evaluates it on almost forty tasks. So the company shipped three kinds of checkpoint: pretrained ones (pt) for fine-tuning, one mixed checkpoint (mix) for trying the model out, and fifty-four task-specific checkpoints (ft) — separate weights for COCO captioning, DocVQA, TextVQA, ScienceQA, remote-sensing question answering, screen description, segmentation and more. Almost nobody else publishes weights for individual benchmarks, and the point is reproducibility: a researcher can check Google's published figure rather than take it on faith. The licence is where the caveat sits. Unlike Gemma 4 or SigLIP 2, which are Apache 2.0, PaliGemma is covered by the Gemma Terms of Use and its repositories are gated — a download requires accepting the terms and waiting for manual approval. Even so, this generation remains the traffic leader of the line: 619,679 downloads in the thirty days to 17 September 2026, about ten times the figure for the newer PaliGemma 2, which Google describes as a drop-in replacement. The bulk of it goes to two repositories — the plain 224-pixel pretrained base, the natural starting point for fine-tuning, and the COCO captioning checkpoint that sits inside a great many automated captioning pipelines.

#vision-language model#image captioning#object detection#OCR#open weights
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review