The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

14FEB2023replayed
one year on
researchStephen Wolfram · OpenAI

Stephen Wolfram publishes lengthy explainer on what ChatGPT is doing and why it works

The Mathematica creator offers a deep dive into the mechanics of large language models, arguing that ChatGPT is fundamentally a next-word predictor and explaining the internal representations he calls "attractors" in a "meaning space".

Stephen Wolfram, creator of Mathematica and Wolfram Alpha, today published a sprawling essay titled “What Is ChatGPT Doing … and Why Does It Work?” that aims to demystify the workings of large language models for a general technical audience.

The piece walks readers through the core mechanism of ChatGPT: it is ‘what ChatGPT is always fundamentally trying to do is to produce a reasonable continuation of whatever text it’s got so far,’ selecting words one token at a time based on learned probabilities. Wolfram compares this to n-gram models but notes that LLMs use a 175-billion-parameter neural net trained on billions of webpages to estimate probabilities for sequences never seen verbatim in training data.

Wolfram emphasizes that the model’s internal representations—what he calls ‘attractors’ in a ‘meaning space’—allow it to produce coherent text. He writes that there is no mathematical theory yet for why it works, but that the results are striking. The essay includes interactive Wolfram Language code examples to illustrate token probabilities, temperature scaling, and the differences between GPT-2 and GPT-3.

On Hacker News, the essay received over 1,090 points and 496 comments, with many commenters debating whether LLMs truly learn world models or merely surface statistics. Some pointed to recent papers showing internal representations of board state in OthelloGPT as evidence of genuine understanding, while others maintained that the models remain probabilistic parrots, however sophisticated.

H
Hacker News commenter spion

Argues that the answer to Wolfram's question is that 'we don't really know' and points to recent papers on world models and internal gradient descent as fun finds.

H
Hacker News commenter codeulike

Highlights the OthelloGPT paper as evidence that GPTs build an internal world model, describing how researchers found 64 nodes representing the board and could flip bits to affect moves.

H
Hacker News commenter kilgnad

Expresses frustration that people dismiss ChatGPT as 'just a probabilistic word generator', arguing that the Othello paper shows deeper understanding, and that such dismissal stems from fear of the unknown.

One year later — open only if you can handle spoilers

Wolfram's essay became one of the most widely cited lay introductions to transformer mechanics. The debate over whether LLMs have world models only intensified later in 2023 with papers on causal abstraction and mechanistic interpretability. The OthelloGPT probes are now a standard example in interpretability tutorials.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy