Creepy crawlies
- ID
- 22213
- Status
- summarized
- Published
- 08 Sep 2026, 7:08 AM
- Fetched
- 08 Sep 2026, 7:30 AM
- Provider
- Simon Willison
- Category
- developer-ai
- Original URL
- https://simonwillison.net/2026/Sep/7/creepy-crawlies/
- Source URL
- https://simonwillison.net/atom/everything/
Summary
- Score
- 7.0
- Created
- 08 Sep 2026, 7:30 AM
- Tags
- Audience
- developerssaas_startup_founders
What happened
Konstantin Ryabitsev reports that git.kernel.org now spends more CPU cycles rendering commits for abusive scrapers than on all legitimate access combined, with 14 cores across 5 geo-distributed nodes dedicated solely to serving crawler traffic. Simon Willison highlights this as a growing concern for any project serving large numbers of crawlable web pages, including his own Datasette.
Why it matters
If you run a public site or API with many crawlable pages, expect aggressive AI scrapers to become a dominant infrastructure cost. Consider rate-limiting, robots.txt rules, or bot-detection now rather than after your hosting bill or CPU usage spikes.
Discussion angle
What practical bot-mitigation strategies work when the crawlers ignore robots.txt and rotate IPs, and how should small teams budget for scraper-driven infrastructure costs?