A One-Line Kernel Fix Closes a Silent Data Loss Bug That Shipped in Every Release Since 6.6

A One-Line Kernel Fix Closes a Silent Data Loss Bug That Shipped in Every Release Since 6.6

A patch merged just after Linux 7.3-rc3 fixes a bug with the worst possible failure mode: writes that succeed, return no error, and are then discarded.

It has been in every kernel since 6.6, introduced in July 2023. The fix is one line.

What it takes to trigger

Three things have to line up:

  1. Transparent hugepages (THP) enabled
  2. The process running under cgroup memory limits
  3. A write that happens after a MADV_FREE call, under heavy memory reclaim pressure

Hit all three and the write can be lost completely. No error, no warning, nothing in the log.

Why MADV_FREE is the ingredient

MADV_FREE is an optimisation an allocator uses to tell the kernel “I am done with this memory for now, but I may write to it again shortly. Reclaim it if you actually need it.”

The bargain is specific. If the kernel does not need the memory, the pages stay and the contents remain valid. If the kernel does need it, the pages are dropped and the next read returns zeroes. Either outcome is fine, because the allocator knows it must not rely on the old contents.

What must never happen is the case in between: the kernel marks the pages reclaimable, the application writes fresh data into them, and the kernel then treats that new data as still-discardable and throws it away. The write is not stale. It happened after the hint. That is the bug.

glibc’s malloc and most modern allocators use MADV_FREE routinely, so the application code does not have to mention it, or even know it exists, to be exposed.

It reached production

This is not a theoretical race someone found by reading code.

The bug was reported at the beginning of September with a small C reproducer, and it came with real casualties: users of Polars, the DataFrame library widely used for data analytics, hit production data loss traced back to this kernel behaviour.

That path makes sense. A data analytics workload is close to a perfect trigger. Large allocations cycling in and out, an allocator leaning on MADV_FREE, container memory limits in place, and sustained reclaim pressure because the dataset is sized to the box. Our cgroups guide covers the memory limit mechanism that forms one leg of this, and it is worth noting that the limit is not the bug. The limit is just what keeps the system in the reclaim state where the bug becomes reachable.

Why nobody caught it for three years

Silent corruption is the hardest class of bug to attribute. There is no crash, no dmesg line, no failed syscall. The application writes a value, reads it back later, and finds something else. Every instinct sends you looking at your own code first, then your libraries, then your storage, and the kernel is somewhere near the end of that list.

The conditions also do not overlap with most testing. Kernel memory management tests rarely run under cgroup limits. Container workloads rarely run kernel memory management tests. The bug lived in the intersection.

The fix

One line, merged as part of the x86/urgent pull:

# the commit
git show 704340f1cd0dcef829eb62f5b48ae95a2ce17bdf

It is marked for backporting to every supported stable series, which is the right call for a data integrity fix present since 6.6.

What to do

If you run analytics, databases, or anything memory-heavy inside containers, take the stable update when your distribution ships it rather than on your usual cadence.

# are you using transparent hugepages
cat /sys/kernel/mm/transparent_hugepage/enabled

# are your services under memory limits
systemctl show --property=MemoryMax <service>

Disabling THP would also close the hole, and it is a considerably larger hammer than the situation calls for now that a fix exists. Waiting for the stable backport is the better move.

The broader lesson is one this site keeps arriving at from different directions: a bug that fails loudly is a bad afternoon, and a bug that fails silently is a slow leak you find out about weeks later. Our backup comparison makes the same point about storage, and a checksumming filesystem as covered in Btrfs versus ZFS is the equivalent defence one layer down. Neither would have caught this one, because the data never reached the filesystem at all.

Background reading

Explainers for the concepts behind this story.