Kernel.org Spends More CPU Rendering Commits for Scrapers Than Serving Real Users

Kernel.org Spends More CPU Rendering Commits for Scrapers Than Serving Real Users

Konstantin Ryabitsev, who runs kernel.org’s infrastructure, has published figures on what AI scrapers are costing the project. They are worse than most people would guess.

git.kernel.org receives roughly 6 million requests a day for individual commit pages. Anubis, the proof-of-work challenge kernel.org deployed, blocks about two thirds of them. What still gets through consumes around 20 percent of kernel.org’s total CPU capacity.

Kernel.org runs 90 CPU cores across five locations. Of those, 14 to 16 are continuously occupied turning Git commits into HTML for automated clients.

“We spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones.”

The part that makes this absurd

None of the scraped data is restricted. All of it is available by cloning the repository, which is one request that transfers a compressed pack file and leaves you with the complete history to query locally, as fast as your disk allows.

Instead, crawlers request commit pages one at a time over HTTP, forcing the server to render each one individually. It is the most expensive possible way to obtain information that Git is specifically designed to hand over cheaply in bulk.

The structure of git.kernel.org makes this much worse. The mainline Linux repository holds about 1.48 million commits, the site hosts hundreds of forks, and each commit is reachable through several views: the commit page, the patch, the diff, the plain-text form. Multiply that out and the crawlable URL space is enormous, which to a naive crawler looks like an inexhaustible supply of pages rather than one repository rendered many ways.

Why the usual defences stopped working

Kernel.org started with the conventional approach: block the offending IPs and netblocks. That worked until crawlers moved onto residential and mobile proxy networks, at which point blocking by address means blocking real users on the same carrier ranges.

Anubis came next. It requires a small proof-of-work computation before serving content, on the theory that the cost is negligible for one human and prohibitive at scraping volume. Crawlers adapted and began solving progressively harder challenges, which tells you something about the economics: whoever is running these is willing to burn real compute to avoid cloning a repository.

What happens next

Kernel.org is now considering measures aimed at the structure of the problem rather than at identifying the offenders: reducing the number of crawlable URLs and restricting resource-intensive operations for anonymous users.

Ryabitsev has been explicit that kernel repositories and development data stay public. The change under discussion is about how many distinct ways the same data can be requested, not about who can have it.

This is the same pressure showing up across open infrastructure. It is a variant of what the kernel’s networking maintainers described during the 7.3 merge window, and of what pushed OpenSSH toward more frequent security releases. In each case a volunteer- or donation-funded project absorbs a cost created by companies that never asked and never contributed to the capacity they consume.

The uncomfortable arithmetic: kernel.org is a non-profit funded by the Linux Foundation, and roughly a fifth of its compute is currently a subsidy to AI training pipelines.

Ryabitsev’s full writeup is at people.kernel.org.

Background reading

Explainers for the concepts behind this story.