MDL-7194EST.2025 · IDX.911
Language modelIn production

Aya Vision 8B

Cohere Labs · Kanada · 2025

Cohere Labs' multilingual vision-language model: an 8-billion-parameter Aya Expanse text model with a vision tower bolted on, reading images and answering in 23 languages including Polish.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Aya Vision 8B, published on 3 March 2025, extends the Aya line from text to images while keeping the language coverage intact: it takes pictures and text as input, returns text, and works across the same 23 languages as Aya Expanse, Polish included. Typical uses named by the publisher are image captioning and visual question answering - reading a photograph, a chart or a page of text and answering questions about it in the user's own language. The repository metadata shows plainly how the model was built. It holds 8,631,842,032 parameters against the 8,028,033,024 of Aya Expanse 8B, so the vision half of the model accounts for roughly 604 million parameters - about 7% of the total, bolted onto a text model that already existed. That is the standard construction for a vision-language model, and it explains why the multilingual behaviour transfers so cleanly. The licence is unchanged from the rest of the family: CC BY-NC 4.0 with Cohere Labs' acceptable use policy, no commercial grant, weights behind a gate. Cohere's production documentation lists only the 32B sibling as a hosted model, so this variant is weights-only. In the thirty days to 18 August 2026 it was downloaded 5,730 times and carries 324 likes.

#multilingual#vision-language#open weights#non-commercial licence#gated weights
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review