Three Dozen Lines Make Opening Files 39% Faster, After Two Years and Five Revisions

Three Dozen Lines Make Opening Files 39% Faster, After Two Years and Five Revisions

A patch queued for Linux 7.4 makes opening files for reading about 39% faster on a 20-core system. It is roughly three dozen lines, and it took two years and five revisions to get there.

What was wrong

Mateusz Guzik has been working on the VFS do_open() path, which runs whenever a pathname lookup resolves to a file you are opening.

His description of the problem is short enough to quote entirely:

“Opening a file grabs a reference on the terminal dentry in __legitimize_path(), then another one in do_dentry_open() and finally drops the initial reference in terminate_walk().

That’s 2 modifications which don’t need to be there — do_dentry_open() can consume the already held reference instead.”

A dentry is the kernel’s cached mapping from a path component to an inode. Every path lookup walks them, and every walk takes and releases references to keep entries alive while in use.

The sequence above takes a reference, takes a second one for the same thing, then releases the first. The net effect is identical to taking one reference and keeping it. The two extra operations accomplish nothing.

Why two redundant operations cost 39%

Because reference counts are atomic operations on memory shared between every CPU in the system.

An atomic increment is cheap in isolation and expensive under contention. When many cores hammer the same cache line, that line ping-pongs between them, and each core stalls waiting for exclusive access. The cost is not the instruction. It is the cache coherency traffic behind it.

Guzik’s benchmark is the case that exposes this precisely: will-it-scale opening the same file read-only across a 20-core VM. Every core touching one dentry is maximum contention on one cache line, and removing two of the four atomic modifications removes roughly half of it.

On a single-threaded workload opening different files, the gain would be far smaller. The headline number is a scalability result, not a universal speedup, and that distinction is worth keeping in mind.

Where it actually shows up

Workloads that open the same files repeatedly from many processes:

  • Web servers serving a shared document root
  • Build systems where every compiler invocation opens the same headers
  • Container hosts where many containers read the same base image layers
  • Anything doing heavy stat-then-open patterns across a large process count

The kernel build speedups landing in the same cycle attack the same general problem from the other end, and a compile job is squarely in the category that benefits from both.

Two years for thirty-six lines

The uncomfortable observation, and Phoronix makes it too, is that this is low-hanging fruit that sat in the tree for a long time.

There are reasons. The VFS is one of the most safety-critical parts of the kernel, reference counting bugs there produce use-after-free rather than wrong output, and “this reference is redundant” is a claim that has to be proven across every path that reaches the function, not just the common one. Five revisions is what that proof costs.

It is also a fair illustration of why the AI-assisted bottleneck hunting elsewhere in the kernel is interesting. Finding this class of redundancy is mechanical. Proving it safe to remove is not, and that part still took a human two years.

Status

Queued in the VFS tree’s vfs-7.4.lookup branch and labelled 7.4 material, so it should go in during the 7.4 merge window after 7.3 ships in October.

# rough local check once you are on 7.4
perf stat -e cache-misses,cycles \
  sh -c 'for i in $(seq 1 10000); do cat /etc/hostname >/dev/null; done'

Background reading

Explainers for the concepts behind this story.