The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

29JAN2025replayed
one year on
businessOpenAI · Microsoft · DeepSeek · David Sacks

OpenAI Furious DeepSeek Might Have Stolen All the Data OpenAI Stole from Us

Microsoft and OpenAI investigate whether DeepSeek trained its R1 model on OpenAI outputs, a practice that closely mirrors how OpenAI itself built its business by scraping the open web.

OpenAI and Microsoft are investigating whether DeepSeek improperly trained its R1 model on outputs from OpenAI’s technology, according to reports from Bloomberg and the Financial Times. The probe centers on whether a group linked to DeepSeek obtained data in an unauthorized manner that could violate OpenAI’s terms of service or circumvent rate limits.

The accusation has drawn sharp commentary because OpenAI itself has faced years of criticism for obtaining large amounts of data from the internet largely in an unauthorized manner, and in some cases in violation of terms of service. “It is, as many have already pointed out, incredibly ironic that OpenAI, a company that has been obtaining large amounts of data from all of humankind largely in an ‘unauthorized manner,’ and, in some cases, in violation of the terms of service of those from whom they have been taking from, is now complaining about the very practices by which it has built its company,” wrote Jason Koebler at 404 Media.

On Hacker News, the article’s headline and framing drew pushback. Several commenters argued there is zero evidence OpenAI is actually “furious” and called the piece clickbait. “The company is not angry, or furious, or enraged. They simply suspect that deepseek broke their usage agreement and are trying to verify that,” said one user. Others, however, found the irony amusing, with one saying it would have been even better as the headline of a satirical takedown from Private Eye or The Onion.

Venture capitalist and Trump administration AI czar David Sacks told Fox News there is “substantial evidence” DeepSeek used a technique called distillation to “suck the knowledge” out of OpenAI’s models. Distillation involves a student model learning from a parent model by asking millions of questions.

D
dang@dang

Moderator who moved comments to a different thread and reminded submitters to link to original sources, characterizing the 404 Media piece as a knock-off.

C
csallen@csallen

Argued the article is clickbait with no evidence OpenAI is actually furious, and that it dishonestly frames them as hypocrites.

R
roshin@roshin

Said OpenAI is not angry; they simply suspect DeepSeek violated usage terms and are trying to verify.

Z
ZeroTalent@ZeroTalent

Speculated based on Twitter chatter that the real issue is suspected corporate espionage and model theft, not just distillation.

One year later — open only if you can handle spoilers

The investigation never produced public evidence of wrongdoing, and DeepSeek continued to release competitive models. The episode became a recurring joke in AI circles about the pot calling the kettle black.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy