one year on
AI scrapers are overwhelming FOSS infrastructure, maintainers say
Open-source projects from SourceHut to KDE to GNOME report that aggressive LLM crawlers ignoring robots.txt are causing outages and forcing desperate countermeasures.
Open-source infrastructure is buckling under an onslaught of AI crawlers that disregard robots.txt and mask their user agents, according to maintainers across multiple projects.
SourceHut founder Drew DeVault published a blogpost this week titled ‘Please stop externalizing your costs directly into my face,’ describing how LLM companies hit expensive endpoints like git blame and every commit log using random IPs and user agents. KDE GitLab was overwhelmed yesterday by IPs from an Alibaba range claiming to be MS Edge, forcing a ban on that browser version. GNOME GitLab switched to Anubis, a proof-of-work challenge, after 97% of 81k requests over 2.5 hours came from bots — a nuclear response that also delays legitimate users when links are shared in chatrooms.
Similar issues plague Fedora infrastructure (Fedora infrastructure has been regularly down for weeks, Neal Gompa says), Inkscape, LWN (Jonathan Corbet says the website might be ‘occasionally sluggish’ due to DDoS from AI scraper bots), and Diaspora (70% of traffic from AI crawlers, per Dennis Schubert). Read the Docs reported a 75% traffic drop after blocking AI crawlers, saving $1500/month. Meanwhile, AI-generated bug reports waste maintainer time on hallucinated vulnerabilities: Daniel Stenberg of Curl and Seth Larson of CPython both describe the burden.
Hacker News commenters are simmering. One calls it ‘burning their goodwill to the ground’; another argues AI companies operate from a position where ‘goodwill is irrelevant.’ The thread debates copyright, fair use, and the existential threat of AI to labor. For now, maintainers are patching with blocklists and CAPTCHAs, but the consensus is that robots.txt is dead, and the desperate mood in sysadmin conversations is palpable.
The record
Complained that LLM crawlers ignore robots.txt and cause severe outages at SourceHut.
Reported that IPs claiming to be MS Edge from Chinese AI companies overwhelmed KDE GitLab.
Shared that 97% of 81k requests to GNOME GitLab in 2.5 hours were bots, per Anubis proof-of-work.
Warned that LWN traffic is mostly bots, causing sluggishness.
Fedora sysadmin who blocked entire country of Brazil due to AI scraper attacks.
Inkscape developer built a 'prodigious block list' against companies spoofing browser info.
Reported AI-generated bug reports wasting developer time in Curl project.
Noted uptick in low-quality, LLM-hallucinated security reports to CPython and other projects.
Claimed 70% of Diaspora traffic is from AI crawlers, ignoring robots.txt and returning every 6 hours.
Blocking AI crawlers cut traffic by 75% and saved $1500/month.
On HN: 'everyone I know has a similar story running large internet infrastructure'.
One year later — open only if you can handle spoilers
Over the following year, several major AI companies updated their crawler policies to respect robots.txt more consistently, and projects like KDE and GNOME maintained Anubis or similar rate-limiting solutions. However, the adversarial scraping dynamic persisted, with new crawlers emerging and the community pushing for legal or structural remedies.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy