Aya Vision 8B
Cohere Labs · Kanada · 2025
Cohere Labs' multilingual vision-language model: an 8-billion-parameter Aya Expanse text model with a vision tower bolted on, reading images and answering in 23 languages including Polish.
Aya Vision 8B, published on 3 March 2025, extends the Aya line from text to images while keeping the language coverage intact: it takes pictures and text as input, returns text, and works across the same 23 languages as Aya Expanse, Polish included. Typical uses named by the publisher are image captioning and visual question answering - reading a photograph, a chart or a page of text and answering questions about it in the user's own language. The repository metadata shows plainly how the model was built. It holds 8,631,842,032 parameters against the 8,028,033,024 of Aya Expanse 8B, so the vision half of the model accounts for roughly 604 million parameters - about 7% of the total, bolted onto a text model that already existed. That is the standard construction for a vision-language model, and it explains why the multilingual behaviour transfers so cleanly. The licence is unchanged from the rest of the family: CC BY-NC 4.0 with Cohere Labs' acceptable use policy, no commercial grant, weights behind a gate. Cohere's production documentation lists only the 32B sibling as a hosted model, so this variant is weights-only. In the thirty days to 18 August 2026 it was downloaded 5,730 times and carries 324 likes.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!