grep -rl "performance" ./blog

Linux Performance Tuning

Finding out why a system is slow: load average, memory pressure, I/O statistics, and the difference between a machine that is busy and one that is struggling.

30 articles

Hardware Video Acceleration on Linux: VA-API in Firefox and Chrome Why video playback makes Linux laptops hot, how VA-API hardware decoding works on Intel, AMD, and NVIDIA, how to check it with vainfo, and how to confirm Firefox and Chrome are actually using the GPU. Speeding Up Linux Boot With systemd-analyze Find out where boot time actually goes with systemd-analyze time, blame, critical-chain, and plot, then fix the usual culprits: network-wait services, slow firmware, unneeded services, and timeouts on missing disks. zstd and 7-Zip on Linux: Fast Modern Compression Explained zstd has become the default compressor for packages, filesystems, and kernels because it is fast at every level. How to use zstd and tar --zstd, choose levels, use long-range mode, and when 7-Zip and its .7z format make more sense. sched_ext Explained: Swapping the Linux CPU Scheduler at Runtime With BPF Since Linux 6.12 you can replace the CPU scheduler on a running system with one written in BPF, and switch back without rebooting. How sched_ext works, what scx_lavd and the other schedulers do, and whether it helps gaming. zswap vs zram: Which Compressed Swap Should You Use? Both compress memory to stretch your RAM, but zram is a swap device in RAM while zswap is a compressed cache in front of disk swap. How each works, how to check and configure them, why you should not stack them, and which to pick. CPU Frequency Scaling and Governors: schedutil, the Intel Driver and EPP Your CPU does not run at its rated speed most of the time. Which driver you have decides which governors exist, why Intel systems only offer two, what energy performance preference actually controls, and when any of this is worth changing. Inodes, Dentries and the Page Cache: What Happens When You Open a File A path is not a file. Resolving one walks a cache of directory entries to find an inode, which points at data the kernel keeps in the page cache. Understanding those three explains disk full errors with free space, and why the second read is instant. io_uring Explained: Why Linux Needed a Third Async I/O Interface Two shared ring buffers replace a syscall per operation. What was wrong with epoll and AIO, how submission and completion queues work, why polled mode can do I/O with zero syscalls, and the security tradeoff that has distributions disabling it. iperf3 Explained: Measuring Network Throughput Properly A speed test to the internet measures your provider. iperf3 measures the link you actually care about. Server and client setup, why a single TCP stream understates fast links, and how to tell a bandwidth problem from a latency one. Memory Pressure, PSI and the OOM Killer: Deciding What Dies The kernel OOM killer picks by score and the score is mostly size, which is why it takes your desktop shell instead of the browser. How to read pressure stall information, set OOM scores that reflect importance, and why systemd-oomd and the kernel disagree. perf Explained: Finding Out Where the CPU Time Actually Goes strace shows syscalls and gdb shows a stopped program. perf samples a running process thousands of times a second and tells you which functions are burning CPU, with flame graphs, without recompiling anything. Valkey and Redis Basics: Caching, Persistence and Not Losing Data An in-memory store used as a cache, a queue and a session store. Why the fork happened, what persistence does and does not guarantee, the eviction policy that decides whether it is a cache or a database, and how it gets left open to the internet. systemd Socket Activation Explained How systemd holds the listening socket so a service starts on first connection, why that removes startup ordering problems, and where it does and does not help. eBPF and bpftrace Explained How eBPF runs sandboxed programs inside the kernel, what the verifier will and will not allow, and the bpftrace one-liners that answer production questions strace cannot. Hugepages and NUMA Explained How TLB misses cost performance, the difference between transparent and explicit hugepages, why databases disable THP, and reading NUMA topology properly. The OOM Killer Explained: Why Linux Killed Your Process How memory overcommit leads to the OOM killer, how it picks a victim, how to read the kernel log entry, and how to stop it choosing the wrong process. Traffic Shaping with tc: Fixing Bufferbloat and Limiting Bandwidth tc controls how the kernel queues outgoing packets. The most valuable thing it does is not limiting bandwidth but fixing the latency spikes that make a saturated link feel broken. fio: Benchmarking Disks Properly dd measures one thing badly. fio measures what you actually care about: random IOPS, latency percentiles, and behaviour under concurrency. Here is how to run it without fooling yourself. Linux I/O Schedulers: mq-deadline, BFQ, Kyber, and none The scheduler decides the order requests reach your disk. On spinning rust it matters enormously, on NVMe it often should be switched off, and the right default differs per device. zram and Swap Tuning: Compressed Memory Instead of Disk zram gives you swap that lives in RAM, compressed. On machines with limited memory it is dramatically faster than swapping to disk, and most distributions now enable it by default. ulimit Explained: Per-Process Resource Limits and Why Services Ignore Them Too many open files is one of the most common server errors, and raising the limit is more subtle than it looks. Here is how soft and hard limits work, and why the shell setting does not apply to your systemd service. cgroups Explained: Limiting CPU, Memory, and I/O With systemd Control groups are the kernel mechanism behind container resource limits, and you can use them directly on any service. Here is what cgroups v2 does and how to cap a process without writing a container. nice, renice, and ionice: Controlling Process Priority on Linux A backup job should not make your desktop stutter. Priority tools let you tell the kernel which work matters less, for both CPU time and disk access, and the disk one is usually the fix people actually need. sysctl Explained: Reading and Changing Kernel Parameters IP forwarding, swappiness, file descriptor limits, and connection backlogs are all kernel parameters you can read and change at runtime. Here is how sysctl works and which settings are actually worth touching. Swap Space Explained What swap actually does on Linux, how it differs from RAM, how to create a swap partition or swapfile, and how to tune swappiness for desktops, laptops, and servers. CPU Monitoring on Linux A practical guide to monitoring CPU usage on Linux, covering top, htop, and mpstat, how to read user/system/wait time breakdowns, per-core versus aggregate views, and identifying whether a bottleneck is CPU-bound or I/O-bound. iostat Explained iostat reports per-device disk I/O statistics, showing throughput, request sizes, queue depth, and utilization. This guide covers installing sysstat, reading iostat -x output, and identifying which disk is actually the bottleneck. Monitoring Memory Usage on Linux A practical guide to checking memory usage on Linux: reading /proc/meminfo, understanding what counts as "used" memory, finding which process is consuming RAM, and recognizing real memory pressure versus normal cache behavior. Understanding Load Average on Linux Load average is one of the most misread numbers in Linux administration. This guide explains what it actually measures, why Linux counts differently from other Unix systems, and how to judge whether a load average is actually a problem. vmstat Explained vmstat reports processes, memory, swap, I/O, and CPU activity in one compact table, sampled at an interval you choose. This guide breaks down every column and how to spot memory pressure, I/O bottlenecks, and CPU contention from its output.