one year on
OpenAI unveils o1 models that 'think' before answering, achieving 83% of problems on IMO qualifying exam
The new Strawberry models use hidden chain-of-thought reasoning and test-time compute, outperforming GPT-4o on complex tasks but costing six times more and prompting OpenAI to withhold raw chains of thought in ChatGPT and show model-generated summaries instead.
OpenAI today released o1-preview and o1-mini, the company’s first models that use hidden chain-of-thought reasoning before answering. Dubbed internally as Strawberry, the models mark a new approach: they spend additional compute time at inference to fact-check themselves and plan responses holistically.
In benchmarks, o1 achieves 83% on the International Mathematical Olympiad qualifying exam, compared to GPT-4o’s 13%. On Codeforces programming challenges, it reaches the 89th percentile. However, Arredondo says o1 can take over 10 seconds to answer some questions, and it is significantly more expensive: $15 per million input tokens and $60 per million output tokens, roughly six times GPT-4o’s cost.
OpenAI says it decided against showing o1’s raw chains of thought in ChatGPT partly due to competitive advantage, and instead shows model-generated summaries. o1 can’t browse the web or analyze files yet, and its image-analyzing features are disabled pending additional testing, with weekly rate limits of 30 messages for o1-preview and 50 for o1-mini.
Reactions are mixed. Early testers praise its reasoning depth in law, science, and code, but note it still hallucinates. In a technical paper, OpenAI says it has heard anecdotal feedback from testers that o1 tends to hallucinate more than GPT-4o and less often admits when it doesn’t have the answer to a question. The company says it aims to experiment with o1 models that reason for hours, days, or even weeks.
The record
VP at Thomson Reuters said o1 is better than previous models at analyzing legal briefs and identifying solutions to LSAT logic games, describing it as more substantive, multi-faceted analysis.
Wharton professor who tested o1 for a month wrote that it solved a challenging crossword puzzle correctly but still hallucinated a new clue, and noted that errors and hallucinations still happen.
OpenAI research scientist said on X that o1 is trained with reinforcement learning and that the longer it thinks, the better it does.
One year later — open only if you can handle spoilers
o1-preview set a new paradigm for reasoning models, with competitors like Google and Anthropic quickly launching similar 'thinking' variants. The hidden chain-of-thought debate continued for months, eventually leading OpenAI to offer opt-in visibility in later versions.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy