Hugepages and NUMA Explained
Two tuning areas that matter for databases, virtual machines, and anything with a large working set, and that are irrelevant on a desktop.
Why page size matters
The CPU translates virtual addresses to physical ones through page tables. Because walking those tables is slow, the result is cached in the translation lookaside buffer, a small hardware cache.
The TLB holds a limited number of entries, typically hundreds to a couple of thousand. With standard 4KB pages, 1,500 entries cover about 6MB of memory.
A process with a 64GB working set is therefore missing the TLB constantly, and each miss means a page table walk of several memory accesses before the actual access happens.
Hugepages fix the arithmetic. One 2MB page uses one TLB entry, so 1,500 entries cover 3GB. One 1GB page covers a gigabyte with a single entry.
grep -i huge /proc/meminfo
AnonHugePages: 786432 kB
HugePages_Total: 0
HugePages_Free: 0
HugePages_Rsvd: 0
Hugepagesize: 2048 kB
Transparent hugepages
The kernel assembles hugepages automatically where it can.
cat /sys/kernel/mm/transparent_hugepage/enabled
# [always] madvise never
always is the default on most distributions and is the setting databases tell you to change.
The problem is not hugepages. It is khugepaged, the kernel thread that compacts memory to assemble them. Compaction moves pages around, and it can stall an allocating process while it works.
For throughput-oriented work that is a good trade. For latency-sensitive work it produces unpredictable multi-millisecond pauses with no application-visible cause, which is exactly the failure mode that is hardest to diagnose.
So PostgreSQL, MongoDB, Redis, and Oracle all recommend disabling it:
echo never | sudo tee /sys/kernel/mm/transparent_hugepage/enabled
echo never | sudo tee /sys/kernel/mm/transparent_hugepage/defrag
Persist it on the kernel command line:
transparent_hugepage=never
Our kernel parameters guide covers editing that.
madvise is the middle setting: hugepages only for regions that explicitly request them with madvise(MADV_HUGEPAGE). This is arguably the best default, since applications that benefit ask and nothing else is affected.
Check what a process is getting:
grep AnonHugePages /proc/$(pgrep -f postgres | head -1)/smaps_rollup
Explicit hugepages
Reserved at boot or at runtime, and used only by applications that ask.
# runtime, may fail if memory is fragmented
echo 2048 | sudo tee /proc/sys/vm/nr_hugepages
# persistent
echo 'vm.nr_hugepages = 2048' | sudo tee -a /etc/sysctl.d/99-hugepages.conf
sudo sysctl --system
2048 pages at 2MB is 4GB reserved.
Reserve at boot for reliability. Assembling contiguous 2MB regions on a system that has been running for weeks frequently fails because memory is fragmented:
default_hugepagesz=2M hugepagesz=2M hugepages=2048
For 1GB pages, which must be reserved at boot:
default_hugepagesz=1G hugepagesz=1G hugepages=16
Reserved hugepages are unavailable to everything else. They do not fall back to normal use. Reserve 16GB and your 64GB machine has 48GB for everything that is not using hugepages, whether or not anything uses the reservation.
PostgreSQL
# postgresql.conf
huge_pages = try # or 'on' to refuse to start without them
shared_buffers = 8GB
Size the reservation from what the server actually wants:
sudo -u postgres postgres -D /var/lib/postgresql/data -C shared_memory_size_in_huge_pages
# 4200
Reserve a little above that number. Our PostgreSQL guide covers the surrounding settings.
QEMU and KVM
<memoryBacking>
<hugepages>
<page size="2048" unit="KiB"/>
</hugepages>
</memoryBacking>
Worthwhile for VMs, because the guest’s memory access goes through two levels of translation and the TLB pressure is correspondingly higher. Our KVM guide covers the rest.
Verifying
grep -E 'HugePages_(Total|Free|Rsvd)' /proc/meminfo
HugePages_Free equal to HugePages_Total means nothing is using them. You have taken memory away from the system for no benefit, which is the most common hugepage misconfiguration.
NUMA
On a multi-socket server, memory is attached to specific sockets. A core reaches its own socket’s memory quickly and another socket’s memory through an interconnect, more slowly.
numactl --hardware
available: 2 nodes (0-1)
node 0 cpus: 0 1 2 3 4 5 6 7
node 0 size: 65536 MB
node 1 cpus: 8 9 10 11 12 13 14 15
node 1 size: 65536 MB
node distances:
node 0 1
0: 10 21
1: 21 10
The distance matrix is relative: 10 is local, 21 means roughly twice the cost. A single-node machine is uniform and none of this applies.
Several single-socket AMD parts present multiple NUMA nodes depending on the BIOS NPS setting, so check rather than assume.
numastat # allocation statistics
numastat -p $(pgrep -f postgres | head -1)
numa_miss and numa_foreign counting up means allocations are landing on the wrong node.
Binding
# confine to node 0
numactl --cpunodebind=0 --membind=0 ./myapp
# prefer local, allow remote rather than failing
numactl --cpunodebind=0 --preferred=0 ./myapp
# interleave across all nodes
numactl --interleave=all ./myapp
--membind fails allocations when the node is full rather than falling back. --preferred is safer for anything you do not want to die.
--interleave=all spreads allocations evenly. Counterintuitive and correct for a large shared-memory workload accessed by threads on every node, because uniformly mediocre beats half-fast and half-slow with unpredictable variance.
For a systemd service:
[Service]
NUMAPolicy=bind
NUMAMask=0
The KVM case
A virtual machine should fit within one NUMA node where possible. A VM spanning two nodes has its vCPUs randomly accessing remote memory with no way for the guest to know.
<numatune>
<memory mode="strict" nodeset="0"/>
</numatune>
<cputune>
<vcpupin vcpu="0" cpuset="0"/>
<vcpupin vcpu="1" cpuset="1"/>
</cputune>
For a VM larger than one node, expose the topology to the guest so it can make its own decisions rather than being unaware.
When to bother
Yes: database servers, virtualisation hosts, HPC, in-memory caches, anything with a working set in the tens of gigabytes.
No: desktops, small VPS instances, web servers, anything where the working set fits comfortably in cache.
Measure first. These are tuning parameters, and applying them without a measured problem usually achieves nothing and occasionally makes things worse by reserving memory that goes unused.
perf stat -e dTLB-load-misses,dTLB-loads ./myapp
A high dTLB-load-misses ratio is the evidence that hugepages will help. Without it, you are guessing.
Frequently Asked Questions
What are hugepages and why do they help?
A hugepage maps 2MB or 1GB of memory with a single translation entry instead of the usual 4KB. Because the translation lookaside buffer holds a limited number of entries, larger pages cover far more memory with the same number of entries, which reduces the misses that force an expensive page table walk.
Why do databases tell you to disable transparent hugepages?
Transparent hugepages are assembled by a kernel thread that compacts memory, and that compaction can stall an application unpredictably. Databases care more about consistent latency than about peak throughput, so PostgreSQL, MongoDB, Redis, and Oracle all recommend disabling THP and using explicit hugepages instead.
What is NUMA?
Non-Uniform Memory Access describes a machine where memory is attached to specific CPU sockets, so a core reaches its local memory faster than memory attached to another socket. Nearly all multi-socket servers are NUMA, as are several single-socket AMD parts with multiple memory controllers.
How do I tell whether my machine is NUMA?
Run numactl —hardware, which lists the memory nodes and the distance matrix between them. A machine with one node is effectively uniform, and a distance of 21 or higher between nodes means cross-node access is measurably slower than local access.
Should I enable hugepages on a normal desktop?
No. The benefit appears in workloads with large working sets and heavy random access, mainly databases, virtual machines, and scientific computing. On a desktop the default transparent hugepages setting is fine and explicit hugepages just reserve memory nothing uses.
How do I confirm hugepages are actually being used?
Check the HugePages_Free and HugePages_Rsvd lines in /proc/meminfo, and look at AnonHugePages for transparent hugepage use by processes. If you reserved pages and HugePages_Free equals HugePages_Total, nothing is using them and the memory is simply unavailable to everything else.