AI Weekly Malaysia

Back to items Summaries

Creepy crawlies

ID
22213
Status
summarized
Published
08 Sep 2026, 7:08 AM
Fetched
08 Sep 2026, 7:30 AM
Provider
Simon Willison
Category
developer-ai
Original URL
https://simonwillison.net/2026/Sep/7/creepy-crawlies/
Source URL
https://simonwillison.net/atom/everything/

Summary

Score
7.0
Created
08 Sep 2026, 7:30 AM
Tags
Audience
developerssaas_startup_founders

What happened

Konstantin Ryabitsev reports that git.kernel.org now spends more CPU cycles rendering commits for abusive scrapers than on all legitimate access combined, with 14 cores across 5 geo-distributed nodes dedicated solely to serving crawler traffic. Simon Willison highlights this as a growing concern for any project serving large numbers of crawlable web pages, including his own Datasette.

Why it matters

If you run a public site or API with many crawlable pages, expect aggressive AI scrapers to become a dominant infrastructure cost. Consider rate-limiting, robots.txt rules, or bot-detection now rather than after your hosting bill or CPU usage spikes.

Discussion angle

What practical bot-mitigation strategies work when the crawlers ignore robots.txt and rotate IPs, and how should small teams budget for scraper-driven infrastructure costs?

Top