grep -rl "monitoring" ./blog

Linux Monitoring and Performance

Watching what a system is doing: CPU, memory, disk, and network usage, load average, process inspection, and disk health checks before hardware fails.

21 articles

Malware Scanning on Linux: ClamAV, rkhunter, and Lynis Compared Three different tools that are often lumped together as Linux antivirus. What ClamAV, rkhunter, chkrootkit, and Lynis each actually detect, how to run them, how to read their false positives, and when scanning is worth it. lm-sensors Explained: Reading Temperatures, Fans and Voltages sensors-detect, what the readings mean, why a laptop reports 95C and is fine, how to tell throttling from a cooling problem, and how to control fan speed when the firmware will let you. Sending Mail From a Server: msmtp, Postfix Relay and Cron Alerts Running a full mail server to send yourself alerts is the wrong amount of work. Relaying through an existing provider takes ten minutes, and there is one setting that stops your credentials being world readable. Centralized Logging with Loki Loki indexes labels rather than log content, which makes it far cheaper than Elasticsearch and changes how you have to query. What that trade buys and where it hurts. Prometheus Alerting Rules That Do Not Wake You Up For Nothing Writing alert rules with the for clause, why alerting on symptoms beats alerting on causes, and configuring Alertmanager so one outage does not produce forty notifications. Cockpit: A Web Console for Linux Servers What Cockpit gives you, why it does not maintain its own state, the security considerations of exposing it, and when a web console beats SSH. inotify and File Watching Explained How the kernel notifies programs about file changes, why you keep hitting the watch limit, and what fanotify does that inotify cannot. lnav: Log Analysis Without grep Gymnastics A log viewer that parses timestamps, merges files into one timeline, and lets you query logs with SQL, which is what you actually wanted when you started piping grep into awk. Monitoring a Home Server: Netdata, Prometheus, and Grafana One tool gives you everything in five minutes, the other gives you history and alerting but takes an afternoon. Here is which to pick and how the pieces fit together. logrotate Explained: Keeping Log Files From Eating Your Disk logrotate renames, compresses, and eventually deletes aging log files on a schedule, and nearly every distro ships it preconfigured. Here is how the rotation cycle works and how to write configs for your own apps. Monitoring Disk Health with smartctl: Reading SMART Before the Drive Dies Every drive tracks its own health through SMART. smartctl reads those attributes, runs self-tests, and, read correctly, gives you warning before a disk fails. Here is what to check and which numbers actually predict death. The watch Command Explained: Rerunning Commands on Repeat watch reruns any command at a fixed interval and shows the latest output full-screen, turning one-shot commands like df, ss, or sensors into live monitors. Flags like -d highlight what changed between runs. CPU Monitoring on Linux A practical guide to monitoring CPU usage on Linux, covering top, htop, and mpstat, how to read user/system/wait time breakdowns, per-core versus aggregate views, and identifying whether a bottleneck is CPU-bound or I/O-bound. The free Command Explained free is the fastest way to check RAM and swap usage on Linux. This guide covers every column in its output, why "available" matters more than "free", and how to read swap usage correctly. iostat Explained iostat reports per-device disk I/O statistics, showing throughput, request sizes, queue depth, and utilization. This guide covers installing sysstat, reading iostat -x output, and identifying which disk is actually the bottleneck. Monitoring Disk Usage on Linux A practical guide to checking disk space on Linux using df and du, understanding the difference between filesystem-level and directory-level usage, tracking down what is actually filling a disk, and watching out for inode exhaustion. Monitoring Memory Usage on Linux A practical guide to checking memory usage on Linux: reading /proc/meminfo, understanding what counts as "used" memory, finding which process is consuming RAM, and recognizing real memory pressure versus normal cache behavior. Understanding Load Average on Linux Load average is one of the most misread numbers in Linux administration. This guide explains what it actually measures, why Linux counts differently from other Unix systems, and how to judge whether a load average is actually a problem. The uptime Command Explained uptime is the fastest way to check how long a Linux system has been running and how busy it is. This guide covers reading its output, the three load average numbers, and how uptime relates to system reboots and patching. vmstat Explained vmstat reports processes, memory, swap, I/O, and CPU activity in one compact table, sampled at an interval you choose. This guide breaks down every column and how to spot memory pressure, I/O bottlenecks, and CPU contention from its output. ps vs top vs htop: Monitoring Linux Processes ps, top, and htop are the three main tools for inspecting running processes on Linux. Each serves a different purpose. This guide covers what each tool is best for, the most useful commands and keyboard shortcuts, and how to get the information you actually need.