Monitoring Disk Health with smartctl: Reading SMART Before the Drive Dies
Hard drives and SSDs keep a diary. Every reallocated sector, every read error, every degree of temperature and hour of operation is logged inside the drive by SMART, and smartctl is how you read it on Linux. Interpreted correctly, it often turns “the disk died last night” into “the disk warned us for three weeks.” The skill is knowing which of the dozens of numbers mean anything.
Getting the data
smartctl ships in the smartmontools package:
sudo apt install smartmontools
sudo smartctl -a /dev/sda # everything: identity, health, attributes, logs
sudo smartctl -H /dev/sda # just the overall verdict
The -H verdict prints PASSED or FAILED, and here is the first lesson: PASSED is nearly meaningless. The overall assessment only trips when the drive considers itself terminal; drives log months of escalating errors while cheerfully reporting PASSED. The attributes are where the truth lives.
The attributes that predict failure
For SATA spinning disks, smartctl -A prints the attribute table, and four rows deserve your attention, keyed by the RAW_VALUE column:
ID# ATTRIBUTE_NAME ... RAW_VALUE
5 Reallocated_Sector_Ct ... 0
197 Current_Pending_Sector ... 0
198 Offline_Uncorrectable ... 0
188 Command_Timeout ... 0
Reallocated sectors (5): sectors the drive gave up on and remapped to spares. Pending sectors (197): sectors the drive could not read and is waiting to judge; these are your active data-loss candidates. Uncorrectable (198): reads that failed permanently. Command timeouts (188): operations that hung. Large-fleet studies, most famously Backblaze’s drive stats, show these attributes correlate strongly with imminent failure: zeros are what you want, and small nonzero values that grow between checks are the actionable signal. A drive that jumps from 0 to 40 reallocated sectors in a month is telling you its retirement date; note the values, recheck in a week, and let the trend decide.
Temperature (194) and power-on hours (9) round out the picture as context rather than alarms.
NVMe is simpler
NVMe drives replaced the attribute zoo with a clean health log, and smartctl reads it the same way:
sudo smartctl -a /dev/nvme0
# Percentage Used: 12%
# Available Spare: 100%
# Media and Data Integrity Errors: 0
# Temperature: 38 Celsius
Percentage Used is rated write endurance consumed (past 100% the drive is beyond its rating, not necessarily dead). Available Spare falling toward its threshold and any nonzero Media Errors are the NVMe equivalents of reallocated sectors: take them seriously.
Self-tests
Beyond passive attributes, drives run active self-tests on command, entirely internally, while the system stays usable:
sudo smartctl -t short /dev/sda # ~2 minutes: electronics and a sample of the surface
sudo smartctl -t long /dev/sda # hours: full surface read scan
sudo smartctl -l selftest /dev/sda # results log
A long test that completes without error is decent evidence the surface is currently readable end to end; Completed: read failure with an LBA is the opposite, a concrete bad spot, and consistent failures at the same point are disqualifying. A sensible cadence: short test weekly, long test monthly, which is exactly what the companion daemon automates.
Automating with smartd
smartmontools includes smartd, which polls drives, runs scheduled tests, and alerts on changes. A minimal /etc/smartd.conf:
DEVICESCAN -a -o on -S on -s (S/../.././02|L/../../6/03) -m root -M exec /usr/share/smartmontools/smartd-runner
Translated: monitor everything, short test nightly at 02:00, long test Saturdays at 03:00, and notify on failures and attribute changes. For a homelab with several drives, the self-hosted app Scrutiny wraps this exact data in a web dashboard with historical trends, which turns “check smartctl occasionally” into something you will actually notice.
The honest limits
SMART sees media wear and read errors; it does not see a controller board about to quit or a firmware bug, and studies consistently find a substantial minority of drives fail with clean SMART data to the end. Treat SMART as an early-warning system that often works, never as a guarantee, and let it change when you replace drives, not whether you keep backups. A pending-sector warning with a current backup is a maintenance task; the same warning without one is a crisis.
Frequently Asked Questions
What is SMART?
Self-Monitoring, Analysis and Reporting Technology, a standard where drives track their own error counts, wear, temperature, and operating history, exposing the data for tools like smartctl to read.
Does PASSED in smartctl mean my drive is healthy?
Not really. The overall assessment only fails when the drive is already near death. A drive can report PASSED while logging reallocated and pending sectors that signal serious trouble. Read the attributes, not just the verdict.
Which SMART attributes matter most?
For spinning disks: reallocated sectors, pending sectors, uncorrectable sectors, and command timeouts. Nonzero and growing values in these correlate strongly with failure. For SSDs and NVMe: percentage used, available spare, and media errors.
How do I run a disk self-test?
smartctl -t short /dev/sda for a quick test or -t long for a full surface scan, then view results later with smartctl -a. Tests run inside the drive and the system stays usable meanwhile.
Do SMART checks work on NVMe drives?
Yes. smartctl reads the NVMe health log with the same commands, reporting percentage used, spare capacity, temperature, and media errors in a simpler format than SATA attribute tables.
Can SMART predict every failure?
No. A meaningful share of drives die with no SMART warning at all, especially from electronics failure rather than media wear. SMART shifts odds, it does not eliminate them. Backups remain the only real protection.