MDL-8925EST.2024 · IDX.275
Language modelIn production

OpenAI o1

OpenAI · United States · 2024

The first model sold on thinking time rather than size — and the start of the reasoning era every lab now competes in.

wujec.ai score

8.9/10

Community score

no votes yet
Sign in to rate

OpenAI o1 broke the pattern that had governed language models until 2024: instead of making the model bigger, OpenAI made it think for longer before answering. The model generates a long internal chain of thought, trained with reinforcement learning, and its accuracy rises with the compute spent at answer time as well as during training. A preview appeared on 12 September 2024 and the full model reached ChatGPT on 5 December 2024. The gains were concentrated in mathematics, science and competitive programming rather than in writing. On the AIME 2024 mathematics olympiad qualifier o1 scored 74.4% on a single attempt, against a few per cent for GPT-4o; on the GPQA Diamond set of PhD-level science questions it reported 77.3% (OpenAI's materials also cite 78.3% for a later configuration), above the 69.7% scored by human experts holding doctorates in the relevant fields. On Codeforces it reached an Elo rating of 1673, around the 89th percentile of competitors. The design also introduced a new commercial unit: reasoning tokens. The internal deliberation is billed as output but not shown to the user — OpenAI publishes only a summary of it, citing safety and competitive reasons. That made o1 expensive and slow by the standards of the time ($15 per million input tokens and $60 per million output), and unsuitable for many everyday tasks, but it established the pattern that GPT-5, Claude's extended thinking, Gemini's thinking modes and DeepSeek-R1 all follow. A cheaper variant, o1-mini, was aimed at coding and STEM work; o1-pro, released through the API in March 2025 at $150 and $600 per million tokens, offered a higher-compute tier. The line was superseded by o3 and then folded into the GPT-5 family, where reasoning effort became a setting rather than a separate model.

#reasoning#historic#chain-of-thought#legacy flagship
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review