MDL-2687EST.2026 · IDX.229
VideoIn production

Marengo 3.5

TwelveLabs · United States · 2026

An embedding model built for video rather than adapted to it: one 512-dimensional representation covering picture, speech, music, ambient sound and PDF pages, with no limit on how long a file may be.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Marengo 3.5, released on 31 August 2026, is TwelveLabs' embedding model — the component that turns video into numbers a search engine can compare. It does not generate text; a companion model, Pegasus, does that. An embedding is a list of numbers describing a piece of content, so that similar content lands nearby. Most systems reach video by stitching together separate tools: one for frames, one for a speech transcript, one for the text. Marengo treats a recording as one object and returns a single 512-dimensional vector covering picture, speech, music and non-dialogue sound. Version 3.5 moved to one unified audio encoder, where 3.0 kept a separate transcription embedding — meaning a door slamming or a crowd roaring is now described in the same space as the words spoken over it. The practical changes over 3.0 are unusually concrete. Duration limits are gone: 3.0 accepted up to four hours of video or audio, 3.5 accepts any length. PDFs are now valid input, one embedding per page. Queries can combine text with images, video or audio at once, rather than text with images only. And each embedding comes with a per-dimension uncertainty vector, so a system can tell which parts of its own description are shaky — a rare thing to publish at all. There is a real catch, stated by the vendor rather than discovered: the platform's own /search endpoint does not accept Marengo 3.5. To search inside TwelveLabs you still need Marengo 3.0 on the index; 3.5 embeddings have to be queried in a system you run yourself. The two versions are also incompatible, so moving means regenerating every embedding you hold. Billing changed with the model. Marengo 3.0 was metered by duration; 3.5 is metered in input tokens, priced separately by content type, from $0.065 per million tokens for audio to $0.260 for video. Weights are not published; the model is available through the TwelveLabs API and on Amazon Bedrock.

#video understanding#embeddings#multimodal#search#closed model
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review