one year on
Don Knuth publishes transcript of ChatGPT experiment, highlighting AI's confident errors
The legendary computer scientist reveals a 20-question test he gave ChatGPT in April, with answers that are often plausible but wrong, sparking debate about the limits of large language models.
Donald Knuth, the legendary computer scientist and author of “The Art of Computer Programming,” has released a transcript of a 20-question experiment he conducted with ChatGPT on April 7, 2023. The results, published on his Stanford University page today, show the AI delivering confident but often erroneous answers to a range of questions designed to probe its limits.
Knuth’s questions included prompts like “Tell me what Donald Knuth says to Stephen Wolfram about ChatGPT” and “Why does Mathematica give the wrong value for Binomial[-1,-1]?” In response, ChatGPT produced lengthy, plausible-sounding answers that were factually incorrect or evasive. For example, it claimed the sun would be directly overhead in Kagoshima, Japan on July 4, 2023, and it described a non-existent ballet version of “Flower Drum Song.” The AI also failed to write a sonnet that is also a haiku, ultimately producing a non-sonnet, and it generated an essay avoiding the word “the” that actually included the word.
The transcript has ignited discussion on Hacker News, where the thread has amassed 927 points and 622 comments. Commenters are divided: some see the experiment as a clear demonstration of the AI’s limitations, comparing its erratic performance to self-driving car failures, while others argue the technology is still improving rapidly. One commenter noted that ChatGPT is “a very capable and convincing liar,” while another pointed out that the AI’s inability to recognize its own contradictions reveals a fundamental lack of understanding. The conversation reflects a growing unease about trusting large language models for tasks where correctness matters.
The record
thread debaters whether ChatGPT's flaws are inherent to neural nets, drawing parallels to self-driving cars
argues that ChatGPT's training on all available data still yields deficiencies, suggesting a need for a fundamental leap forward
calls ChatGPT 'a very capable and convincing liar'
One year later — open only if you can handle spoilers
In the years that followed, Knuth's experiment became a frequently cited example of the 'hallucination' problem in large language models. While later models like GPT-4 showed improvements on some of these specific questions, the underlying issue of confident falsehoods persisted, shaping ongoing research into AI alignment and factual reliability.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy