The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

19JUL2025replayed
one year on
researchOpenAI

OpenAI: experimental model scores IMO gold

An unreleased reasoning model solves 5 of 6 International Mathematical Olympiad problems under contest conditions, writing natural-language proofs graded by former medalists — a milestone many forecasters had penciled in for years later.

OpenAI researcher Alexander Wei announces that an experimental reasoning model has achieved a gold-medal score on the 2025 International Mathematical Olympiad: 35 out of 42 points, solving five of six problems.

The conditions matter: two 4.5-hour sessions, no tools, no internet, proofs written in natural language and graded by former IMO medalists. This is not a specialized theorem-prover with a formal-verification crutch — the claim is that general-purpose reasoning methods got here.

Sam Altman calls it a significant marker of how fast the field is moving, while noting the model is a research artifact: GPT-5 arrives soon, but “we don’t plan to release a model with this level of math capability for many months.”

The milestone instantly generates two arguments. One about timing and manners — announcing during the IMO’s own weekend, with self-graded results, while other labs waited for official certification. And one about calibration: as recently as last year, expert forecasts put this achievement years away. The proofs are public. The forecasts were wrong.

A
Alexander Wei@alexwei_

Stresses what's new: no tools, no formal verifiers, no internet — just 4.5 hours, natural-language proofs, and general-purpose reinforcement learning.

S
Sam Altman@sama

Says the model won't ship for many months — GPT-5 is coming soon, but this level of math capability stays in the lab.

T
The etiquette dispute

Mathematicians and rival labs object to the announcement's timing around the IMO's own closing ceremony, and to self-graded results — a controversy that says as much about the field's trust levels as its math.

From the record — 1 original post, embedded officially
One year later — open only if you can handle spoilers

DeepMind's officially-certified gold two days later validated the result wasn't a one-off — and by the following spring, olympiad-level reasoning had shipped in consumer models. The forecasting community's 'AI gets IMO gold' median, which sat at 2027+ as late as 2024, became a case study in systematically underestimating the pace.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy