The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

20MAR2025replayed
one year on
communityOpenAI · Anthropic · Alibaba · KDE · GNOME · SourceHut · Fedora · Inkscape · Read the Docs · LWN · Diaspora · Curl · CPython · pip · urllib3 · Requests

AI scrapers are overwhelming FOSS infrastructure, maintainers say

Open-source projects from SourceHut to KDE to GNOME report that aggressive LLM crawlers ignoring robots.txt are causing outages and forcing desperate countermeasures.

Open-source infrastructure is buckling under an onslaught of AI crawlers that disregard robots.txt and mask their user agents, according to maintainers across multiple projects.

SourceHut founder Drew DeVault published a blogpost this week titled ‘Please stop externalizing your costs directly into my face,’ describing how LLM companies hit expensive endpoints like git blame and every commit log using random IPs and user agents. KDE GitLab was overwhelmed yesterday by IPs from an Alibaba range claiming to be MS Edge, forcing a ban on that browser version. GNOME GitLab switched to Anubis, a proof-of-work challenge, after 97% of 81k requests over 2.5 hours came from bots — a nuclear response that also delays legitimate users when links are shared in chatrooms.

Similar issues plague Fedora infrastructure (Fedora infrastructure has been regularly down for weeks, Neal Gompa says), Inkscape, LWN (Jonathan Corbet says the website might be ‘occasionally sluggish’ due to DDoS from AI scraper bots), and Diaspora (70% of traffic from AI crawlers, per Dennis Schubert). Read the Docs reported a 75% traffic drop after blocking AI crawlers, saving $1500/month. Meanwhile, AI-generated bug reports waste maintainer time on hallucinated vulnerabilities: Daniel Stenberg of Curl and Seth Larson of CPython both describe the burden.

Hacker News commenters are simmering. One calls it ‘burning their goodwill to the ground’; another argues AI companies operate from a position where ‘goodwill is irrelevant.’ The thread debates copyright, fair use, and the existential threat of AI to labor. For now, maintainers are patching with blocklists and CAPTCHAs, but the consensus is that robots.txt is dead, and the desperate mood in sysadmin conversations is palpable.

D
Drew DeVault

Complained that LLM crawlers ignore robots.txt and cause severe outages at SourceHut.

B
Ben (KDE sysadmin)

Reported that IPs claiming to be MS Edge from Chinese AI companies overwhelmed KDE GitLab.

B
Bart Piotrowski

Shared that 97% of 81k requests to GNOME GitLab in 2.5 hours were bots, per Anubis proof-of-work.

J
Jonathan Corbet

Warned that LWN traffic is mostly bots, causing sluggishness.

K
Kevin Fenzi

Fedora sysadmin who blocked entire country of Brazil due to AI scraper attacks.

M
Martin Owens

Inkscape developer built a 'prodigious block list' against companies spoofing browser info.

D
Daniel Stenberg

Reported AI-generated bug reports wasting developer time in Curl project.

S
Seth Larson

Noted uptick in low-quality, LLM-hallucinated security reports to CPython and other projects.

D
Dennis Schubert

Claimed 70% of Diaspora traffic is from AI crawlers, ignoring robots.txt and returning every 6 hours.

R
Read the Docs

Blocking AI crawlers cut traffic by 75% and saved $1500/month.

E
ericholscher

On HN: 'everyone I know has a similar story running large internet infrastructure'.

One year later — open only if you can handle spoilers

Over the following year, several major AI companies updated their crawler policies to respect robots.txt more consistently, and projects like KDE and GNOME maintained Anubis or similar rate-limiting solutions. However, the adversarial scraping dynamic persisted, with new crawlers emerging and the community pushing for legal or structural remedies.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy