The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

12SEPT2024replayed
one year on
model launchOpenAI

OpenAI unveils o1 models that 'think' before answering, achieving 83% of problems on IMO qualifying exam

The new Strawberry models use hidden chain-of-thought reasoning and test-time compute, outperforming GPT-4o on complex tasks but costing six times more and prompting OpenAI to withhold raw chains of thought in ChatGPT and show model-generated summaries instead.

OpenAI today released o1-preview and o1-mini, the company’s first models that use hidden chain-of-thought reasoning before answering. Dubbed internally as Strawberry, the models mark a new approach: they spend additional compute time at inference to fact-check themselves and plan responses holistically.

In benchmarks, o1 achieves 83% on the International Mathematical Olympiad qualifying exam, compared to GPT-4o’s 13%. On Codeforces programming challenges, it reaches the 89th percentile. However, Arredondo says o1 can take over 10 seconds to answer some questions, and it is significantly more expensive: $15 per million input tokens and $60 per million output tokens, roughly six times GPT-4o’s cost.

OpenAI says it decided against showing o1’s raw chains of thought in ChatGPT partly due to competitive advantage, and instead shows model-generated summaries. o1 can’t browse the web or analyze files yet, and its image-analyzing features are disabled pending additional testing, with weekly rate limits of 30 messages for o1-preview and 50 for o1-mini.

Reactions are mixed. Early testers praise its reasoning depth in law, science, and code, but note it still hallucinates. In a technical paper, OpenAI says it has heard anecdotal feedback from testers that o1 tends to hallucinate more than GPT-4o and less often admits when it doesn’t have the answer to a question. The company says it aims to experiment with o1 models that reason for hours, days, or even weeks.

P
Pablo Arredondo

VP at Thomson Reuters said o1 is better than previous models at analyzing legal briefs and identifying solutions to LSAT logic games, describing it as more substantive, multi-faceted analysis.

E
Ethan Mollick

Wharton professor who tested o1 for a month wrote that it solved a challenging crossword puzzle correctly but still hallucinated a new clue, and noted that errors and hallucinations still happen.

N
Noam Brown

OpenAI research scientist said on X that o1 is trained with reinforcement learning and that the longer it thinks, the better it does.

One year later — open only if you can handle spoilers

o1-preview set a new paradigm for reasoning models, with competitors like Google and Anthropic quickly launching similar 'thinking' variants. The hidden chain-of-thought debate continued for months, eventually leading OpenAI to offer opt-in visibility in later versions.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy