Reading a Kernel Panic or Oops

Reading a Kernel Panic or Oops

A kernel panic is the kernel deciding it cannot safely continue. Reading one is a skill, and the message contains more than it appears to.

Oops or panic

An oops is recoverable. The kernel kills the offending task and carries on.

The system is now untrustworthy. Locks may be held by a dead task, memory may be inconsistent, and a subsystem may be half-initialised. An oops means reboot at the first opportunity rather than continue indefinitely.

A panic is unrecoverable and halts the machine. An oops in a critical context, while holding a spinlock or in an interrupt handler, escalates to one.

dmesg | grep -i -E 'oops|BUG:|panic|call trace'
journalctl -k -b -1 | tail -50

-b -1 reads the previous boot, which is where a panic that rebooted the machine left its evidence.

Reading one

BUG: kernel NULL pointer dereference, address: 0000000000000010
#PF: supervisor read access in kernel mode
#PF: error_code(0x0000) - not-present page
PGD 0 P4D 0
Oops: 0000 [#1] PREEMPT SMP NOPTI
CPU: 3 PID: 4821 Comm: nginx Tainted: G           OE     6.18.50-generic
Hardware name: Dell Inc. PowerEdge R640/0H21J3, BIOS 2.19.1
RIP: 0010:myfs_read_inode+0x2a/0x180 [myfs]
Code: 48 8b 47 10 48 85 c0 74 25 ...
RSP: 0018:ffffb3c1c0f4fc58 EFLAGS: 00010246
Call Trace:
 <TASK>
 myfs_lookup+0x8c/0x1d0 [myfs]
 __lookup_slow+0x84/0x150
 walk_component+0x14b/0x1b0
 path_lookupat+0x6e/0x1c0
 filename_lookup+0xe3/0x1d0
 do_sys_openat2+0x9c/0x160
 __x64_sys_openat+0x55/0xa0
 do_syscall_64+0x5c/0xc0
 </TASK>

The first line

NULL pointer dereference, address: 0000000000000010

Something read a struct member at offset 0x10 through a null pointer. The small non-zero address is the giveaway: not a wild pointer, a null one plus a field offset.

RIP

RIP: 0010:myfs_read_inode+0x2a/0x180 [myfs]

This is the single most useful line.

myfs_read_inode is the function. +0x2a is 42 bytes into it. /0x180 is the function’s total size.

A small offset like 0x2a means the fault was near the entry, which very often means a parameter was null and never checked. A large offset means it got well into the work first.

[myfs] is the module. A module name here points at your culprit immediately, and out-of-tree modules are disproportionately represented. Our DKMS guide covers what gets rebuilt against your kernel and is therefore a candidate.

Tainted

Tainted: G           OE

Read these. They tell you how much the trace is worth:

FlagMeaning
GClean, no proprietary module
PProprietary module loaded
OOut-of-tree module loaded
EUnsigned module loaded
DA previous oops occurred
MMachine check exception, meaning hardware

OE above means an out-of-tree unsigned module. Kernel maintainers will ask you to reproduce without it before looking, and they are right to, because the module can corrupt memory in ways that surface anywhere.

D matters for a different reason: it means this is not the first oops, so the state was already bad. The first one is the one to investigate.

M means hardware told the CPU something was wrong. That is not a software bug.

The call trace

Read it top to bottom as most recent first:

myfs_read_inode      <- crashed here
myfs_lookup          <- called by
__lookup_slow
walk_component
path_lookupat
filename_lookup
do_sys_openat2
__x64_sys_openat     <- the openat() syscall
do_syscall_64

Someone called openat(), path resolution reached the myfs filesystem, and its inode read faulted.

Comm: nginx and PID: 4821 name the process, which gives you the workload that triggers it.

Entries with a ? prefix are unreliable guesses from stale stack data. Ignore them; they send people chasing functions that were never called.

Capturing one

The hard part is that a panicked machine frequently shows nothing, or reboots before anyone reads it.

kdump

The proper answer.

sudo apt install kdump-tools crash linux-crashdump
# reserve memory for the crash kernel
crashkernel=512M

kdump keeps a second small kernel in reserved memory. On panic, the machine kexecs into it, and that kernel writes a dump of the failed one to disk.

kdump-config status
ls /var/crash/
crash /usr/lib/debug/boot/vmlinux-$(uname -r) /var/crash/*/dump.*

Inside crash: bt for the backtrace, log for the dmesg buffer, ps for the task list at the moment of death. Our gdb guide covers the same discipline for userspace, and the symbol requirement is identical: without matching debug symbols you get addresses and nothing useful.

netconsole

When there is no disk to write to.

sudo modprobe netconsole netconsole=6666@203.0.113.10/eth0,514@203.0.113.50/00:11:22:33:44:55

Kernel messages go over UDP to another machine. On the receiver:

nc -u -l 514 | tee panic.log

Crude and effective, and it captures output the console never displayed.

Serial console

console=ttyS0,115200 console=tty0

The traditional answer and still the most reliable on hardware that has a serial port or an IPMI serial-over-LAN.

pstore

ls /sys/fs/pstore/

Some firmware provides persistent storage that survives a reboot. If it is available, the panic is sitting there after you restart, with no configuration needed.

Stop the instant reboot

sysctl kernel.panic          # seconds before reboot, 0 means halt
sudo sysctl -w kernel.panic=0

Servers commonly set this to reboot quickly, which is right for availability and wrong for diagnosis. Set it to 0 while investigating. Our sysctl guide covers making it permanent, and remember to set it back.

Suspect hardware first when it is random

A software bug fails in a consistent place. If traces point at a different subsystem every time with no pattern, that is the signature of bad memory.

sudo memtester 4G 3
# or boot memtest86+ from the GRUB menu

sudo apt install rasdaemon
sudo ras-mc-ctl --errors
journalctl -k | grep -i -E 'mce|machine check|edac'

ECC memory reports correctable errors before they become uncorrectable, which is the main practical argument for it and one our Btrfs versus ZFS guide touches on.

A Tainted: M flag or machine check entries in the log mean stop debugging software.

Reporting it

If it is a mainline kernel with no proprietary modules, it is worth reporting.

uname -a
cat /proc/version
lsmod | head -30
journalctl -k -b -1 > panic.txt

Include the full trace rather than a summary, the kernel version, whether it is reproducible, and what you were doing. scripts/decode_stacktrace.sh in the kernel source turns raw addresses into function names if your trace lacks them.

Reproduce without out-of-tree modules first. Otherwise the first response will ask you to, and correctly.

Frequently Asked Questions

What is the difference between a kernel panic and an oops?

An oops is a recoverable kernel error where the offending task is killed and the system continues in a degraded and untrustworthy state. A panic is unrecoverable and halts the machine. An oops in a critical context, such as while holding a lock, escalates to a panic.

What does the RIP line in a panic mean?

RIP is the instruction pointer, showing which function was executing and the byte offset into it when the fault occurred. The format is function+offset/size, so a small offset usually means the fault happened near the function entry, often while validating arguments.

How do I capture a panic when the screen shows nothing?

Configure kdump, which boots a small crash kernel that writes a dump of the failed one to disk. Alternatives are a serial console, netconsole to send messages over the network, or pstore which writes to firmware-backed persistent storage that survives a reboot.

What does tainted mean in a kernel oops?

It records that something happened which makes the kernel less trustworthy for debugging, most commonly that a proprietary or out-of-tree module was loaded. A taint flag of G means clean, P means proprietary, and O means out-of-tree, and upstream maintainers will usually ask you to reproduce without them.

Can a hardware fault cause a kernel panic?

Yes, and bad memory is the most common cause of panics that appear random and unreproducible. If traces point at different subsystems each time with no pattern, test the memory before spending time on software, because a software bug usually fails in a consistent place.

Why does my machine reboot instantly instead of showing the panic?

Because kernel.panic is set to reboot after a number of seconds, which is common on servers. Setting kernel.panic to 0 makes it halt so you can read the message, and configuring kdump is the better answer because it captures the information regardless.