MDL-2600EST.2022 · IDX.826
Language modelIn production

text-embedding-ada-002

OpenAI · USA · 2022

The 2022 embedding model that became an industry default — and the only OpenAI model still sold at five times the price of its own, better replacement.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

text-embedding-ada-002 arrived on 15 December 2022, two weeks after ChatGPT, and quietly became the default embedding model of the early retrieval-augmented generation era. It replaced five separate GPT-3 era embedding models with one, cut the price of embeddings sharply, and produced 1,536-dimension vectors from up to 8,000-odd tokens of text. A large share of the vector databases built in 2023 are full of its output. Its numbers have since been overtaken by OpenAI's own 2024 pair. On the benchmarks the company publishes, ada-002 scores 31.4% on MIRACL multilingual retrieval and 61.0% on MTEB, against 44.0% / 62.3% for text-embedding-3-small and 54.9% / 64.6% for text-embedding-3-large. The model remains on the price list at $0.10 per million tokens — five times what text-embedding-3-small costs, for lower scores at the same vector length. There is a practical reason it has not been retired: embeddings are not interchangeable between models, so switching means re-embedding an entire corpus and rebuilding the index. OpenAI is, in effect, charging a premium for not having to do that migration. Anyone starting fresh has no reason to pick it.

#embeddings#retrieval#legacy#superseded
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review