Inodes, Dentries and the Page Cache: What Happens When You Open a File

Inodes, Dentries and the Page Cache: What Happens When You Open a File

open("/var/log/syslog") looks like one operation. It is a path walk, a cache lookup per component, an inode resolution and possibly a disk read, and knowing which is which explains several confusing errors.

Three things, not one

The dentry maps one path component to an inode. /var/log/syslog involves three of them.

The inode is the file’s metadata: permissions, owner, timestamps, size, link count, and where the data blocks are. The inode does not contain the filename. Names live in directory entries.

The page cache holds file contents in memory, so a second read does not touch the disk.

# the inode number behind a name
stat -c '%i %n' /etc/hostname
ls -i /etc/hostname

Path resolution, one component at a time

Opening /var/log/syslog means:

  1. Start at /, whose inode the kernel already has
  2. Look up var in it, getting an inode
  3. Look up log in that, getting an inode
  4. Look up syslog in that, getting the final inode

Each step consults the dentry cache. A hit costs a hash lookup. A miss means reading directory data from disk.

This is why deep paths are marginally slower than shallow ones, and why a directory with a very large number of entries can be slow to search on filesystems without good directory indexing.

It is also why the kernel does a lot of reference counting during a path walk. A patch queued for Linux 7.4 removes a redundant pair of dentry reference operations from the open path and gains roughly 39% more opens per second on a 20-core machine, purely by not doing atomic work it did not need to do.

Inode exhaustion

The error that makes no sense until you know the model:

$ touch newfile
touch: cannot touch 'newfile': No space left on device

$ df -h /
Filesystem  Size  Used Avail Use% Mounted on
/dev/sda1   100G   43G   52G  46% /

$ df -i /
Filesystem  Inodes  IUsed IFree IUse% Mounted on
/dev/sda1   6553600 6553600    0  100% /

Every file consumes one inode regardless of size. A million empty files use a million inodes and almost no blocks.

On ext4 the inode count is fixed when the filesystem is created and cannot be raised afterwards. Running out means backing up, recreating with a different ratio, and restoring.

# find the directories responsible
sudo find / -xdev -type f 2>/dev/null | cut -d/ -f2-3 | sort | uniq -c | sort -rn | head

Usual culprits: session files, mail spools, cache directories, and any application that writes one small file per event.

Btrfs, XFS and ZFS allocate inodes dynamically, so they do not have this failure mode. That is a genuine practical advantage, and our Btrfs versus ZFS comparison covers the wider tradeoffs.

The page cache is why your memory looks full

free -h
#               total   used   free  shared  buff/cache  available
# Mem:           16Gi  4.2Gi  312Mi   890Mi        11Gi       11Gi

312 MiB free looks alarming and is fine. 11 GiB is page cache, holding file data the kernel read earlier, and it is released instantly when anything needs it.

available is the number that matters. free is memory doing nothing, which on a healthy Linux system should be close to zero, because idle RAM is wasted RAM.

# demonstrate it
sync && echo 3 | sudo tee /proc/sys/vm/drop_caches
time cat largefile > /dev/null    # reads from disk
time cat largefile > /dev/null    # reads from RAM

The difference is usually an order of magnitude or more. It is also why any benchmark that does not drop caches between runs is measuring RAM rather than storage, which our fio guide covers avoiding.

Do not drop caches on a production system to “free memory”. You are discarding useful data and the kernel will simply read it again.

echo hello > original
ln original hardlink        # another name for the same inode
ln -s original symlink      # a small file containing a path

ls -li original hardlink symlink

The hard link shares the same inode number. Both names are equally real; neither is the original. The data lives until the link count reaches zero.

stat -c '%h %i %n' original hardlink

A symlink is its own inode containing a path string. It breaks if the target moves, and it can point across filesystems, which hard links cannot because inode numbers are only meaningful within one filesystem.

Deleted files that still occupy space

$ df -h /var
# 95% full

$ du -sh /var
# accounts for far less

An inode is freed when both conditions hold: link count zero and no process has it open.

Delete a 10 GB log a service is still writing to and the name disappears while the data stays. The service keeps writing into a file with no name.

sudo lsof +L1
sudo lsof -nP | grep '(deleted)'

The fix is restarting the holder, or truncating through its file descriptor rather than deleting:

: > /var/log/big.log      # truncate, keeps the inode, service keeps working

This is why logrotate uses copytruncate or signals the service to reopen. Deleting a log a daemon holds open accomplishes nothing.

Watching the caches

# dentry and inode cache sizes
grep -E 'dentry|inode' /proc/slabinfo | head

# summary
sudo slabtop -o | head -15

# per filesystem inode usage
df -i

A very large dentry cache is normal on a system that has walked many paths. The kernel reclaims it under pressure, and our memory pressure guide covers reading whether that reclaim is actually hurting anything.

Why this is worth knowing

Three specific payoffs:

“No space left on device” with free space is inode exhaustion, and df -i answers it in one command.

“The server has no free memory” is usually page cache, and available is the column to read.

“I deleted the file and space did not come back” is an open file descriptor, and lsof +L1 finds it.

Each of those is a problem people lose an afternoon to at least once, and each becomes obvious the moment the model is clear.

Frequently Asked Questions

Why do I get no space left on device when df shows free space?

You have probably run out of inodes rather than blocks. Each file consumes one inode regardless of its size, and most filesystems fix the inode count at creation time. Check with df -i, and the usual cause is millions of tiny files such as cache or session data.

What is a dentry and why does it matter?

A dentry is the kernel cached mapping from one path component to an inode. Resolving a path means walking one dentry per component, so the cache is what makes repeated path lookups fast. Without it every open would re-read directory data from disk.

Why does free -h show almost no free memory on a healthy system?

Because the kernel uses otherwise idle memory for the page cache, which holds recently read file data. That memory is reclaimed instantly when something needs it, so it is available rather than used. The available column is the number that matters, not free.

Why is reading a file the second time so much faster?

The first read pulls data from disk into the page cache. The second read is served from RAM without touching the device at all. This is also why benchmarks that do not drop caches between runs produce numbers that have nothing to do with your storage.

A hard link is a second directory entry pointing at the same inode, so both names are equally real and the data survives until all of them are removed. A symlink is a small file containing a path, so it breaks if the target moves and can cross filesystems, which hard links cannot.

Why does deleting a file not free space until I restart a service?

Because an inode is only freed when both its link count and its open file count reach zero. Removing the name drops the link count, but a process still holding the file open keeps the data alive. Find it with lsof and the deleted marker.