MDL-3974EST.2024 · IDX.436
Language modelIn production

text-embedding-3-large

OpenAI · USA · 2024

OpenAI's most capable embedding model: up to 3,072 dimensions, and short enough at 256 dimensions to still beat the model it replaced at 1,536.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Released on 25 January 2024, text-embedding-3-large is OpenAI's strongest embedding model. Embeddings are not answers — they are numeric representations of text used to measure how related two pieces of text are, which is what powers search, clustering, recommendation, anomaly detection and classification. This model produces vectors of up to 3,072 numbers. Against the ada-002 generation it replaced, OpenAI reports the MIRACL multilingual retrieval average rising from 31.4% to 54.9%, and the English-oriented MTEB average from 61.0% to 64.6%. The multilingual gain is by far the larger of the two — a reminder that the older model's weakness was concentrated outside English. The most useful property is one that only appears in the fine print. Both 2024 models were trained so that a vector can be truncated — the dimensions parameter simply drops numbers off the end — without losing its meaning. OpenAI's own example: shortened to 256 dimensions, this model still outscores a full-length 1,536-dimension ada-002 embedding on MTEB. A vector six times smaller, and better. For anyone paying for vector storage, that is the number that matters, and it is the reason the older model has no remaining technical argument in its favour.

#embeddings#retrieval#search#multilingual
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review