perf Explained: Finding Out Where the CPU Time Actually Goes
strace tells you what a program asks the kernel for. gdb tells you what a stopped program looks like. Neither answers “why is this using 100% of a core”, and that is what perf is for.
Installing it
# Debian and Ubuntu, version must match the kernel
sudo apt install linux-tools-common linux-tools-$(uname -r)
# Fedora and RHEL
sudo dnf install perf
# Arch
sudo pacman -S perf
perf --version
The version tie to the running kernel is real on Debian and Ubuntu. A mismatch produces a warning and degraded functionality rather than a clean failure, which is confusing if you do not know to look for it.
Permissions
cat /proc/sys/kernel/perf_event_paranoid
| Value | What is allowed |
|---|---|
-1 | Everything |
0 | Raw tracepoints, no kernel profiling for users |
1 | CPU events, no tracepoints |
2 | Userspace only, the common default |
3 | Nothing without privilege, some distributions |
# for a session
sudo sysctl kernel.perf_event_paranoid=1
# and to see kernel symbols
sudo sysctl kernel.kptr_restrict=0
Lowering these on a shared machine widens what unprivileged users can observe, so prefer sudo perf or the cap_perfmon capability on a specific binary where that matters.
Start with stat
Before profiling anything, find out what kind of slow it is.
perf stat ./myprogram
perf stat -p 1234 -- sleep 10
4,231.55 msec task-clock # 0.998 CPUs utilized
1,204 context-switches # 284.532 /sec
12,481,033,120 cycles # 2.949 GHz
8,102,441,983 instructions # 0.65 insn per cycle
1,904,221,001 branches # 449.99 M/sec
41,201,882 branch-misses # 2.16% of all branches
Instructions per cycle is the headline number. Modern cores can retire several instructions per cycle. Below roughly 1.0 means the CPU is waiting rather than computing, usually on memory.
# cache behaviour
perf stat -e cache-references,cache-misses,LLC-load-misses ./myprogram
# is it CPU bound at all
perf stat -e task-clock,context-switches,cpu-migrations,page-faults ./myprogram
A high context-switch count with low CPU utilisation means the program is blocking, and no amount of CPU profiling will help. That is a syscall or lock problem.
record and report
# a running process, 30 seconds, with call graphs
sudo perf record -F 99 -g -p 1234 -- sleep 30
# a command from start to finish
perf record -g ./myprogram
# the whole system
sudo perf record -F 99 -a -g -- sleep 30
# then
perf report
-F 99 samples 99 times a second. The odd number is deliberate: it avoids locking in step with timers running at round frequencies, which would bias the sample towards whatever happens on those ticks.
perf report is interactive. Enter expands a call chain, a annotates down to instructions, / searches.
# non-interactive, for a quick look or a script
perf report --stdio --sort=overhead,symbol | head -30
# by shared object rather than function
perf report --sort=dso
Call graphs and why fp often fails
# frame pointer, cheap, frequently broken
perf record -g --call-graph fp ./myprogram
# DWARF, works on optimised binaries, large captures
perf record --call-graph dwarf ./myprogram
# hardware branch records, fast and accurate, needs Intel
perf record --call-graph lbr ./myprogram
Distribution binaries have historically been compiled with -fomit-frame-pointer, which makes frame-pointer unwinding produce truncated or nonsensical stacks. Several distributions have since re-enabled frame pointers precisely because profiling mattered more than the register.
If your stacks are one frame deep or obviously wrong, that is the cause. dwarf works regardless, at the cost of much larger capture files.
Symbols, or the hex address problem
42.31% myprogram [.] 0x00000000004011a6
That is a profile telling you nothing. The binary has no symbols.
# Debian and Ubuntu, enable the debug symbol repository then
sudo apt install myprogram-dbgsym
sudo apt install libc6-dbg
# Fedora
sudo dnf debuginfo-install myprogram
# your own code
gcc -g -fno-omit-frame-pointer -O2 ...
Do not strip binaries you intend to profile, and keep -g even in optimised builds. Debug info does not slow anything down; it only takes space.
Kernel symbols come from /proc/kallsyms, which needs kptr_restrict=0 to be readable.
Flame graphs
perf report gives you a list. A flame graph gives you the shape.
git clone https://github.com/brendangregg/FlameGraph
sudo perf record -F 99 -g -p 1234 -- sleep 30
perf script | ./FlameGraph/stackcollapse-perf.pl | ./FlameGraph/flamegraph.pl > out.svg
Or the built-in version on recent perf:
perf script report flamegraph > out.svg
Reading one correctly:
- Width is share of samples. Wide means the time went there
- Height is call depth, nothing more
- Horizontal order is alphabetical. Left to right carries no meaning, and reading it as a timeline is the standard mistake
- Look for wide plateaus, not tall towers. A deep stack that is narrow is irrelevant
A differential flame graph compares two captures, which is how you check whether a change helped:
perf script -i before.data | ./FlameGraph/stackcollapse-perf.pl > before.folded
perf script -i after.data | ./FlameGraph/stackcollapse-perf.pl > after.folded
./FlameGraph/difffolded.pl before.folded after.folded | ./FlameGraph/flamegraph.pl > diff.svg
perf top
sudo perf top
sudo perf top -p 1234
sudo perf top -e cache-misses
Live, like htop for functions. Good for “something is eating the CPU right now and I do not know what”. Our process monitoring guide covers finding the process; perf top tells you which function inside it.
Events beyond CPU cycles
perf list | head -50
perf list hw cache
# page faults, where memory pressure shows
sudo perf record -e page-faults -g -p 1234 -- sleep 30
# cache misses, where memory layout shows
sudo perf record -e cache-misses -g ./myprogram
# scheduler activity, where blocking shows
sudo perf record -e sched:sched_switch -g -a -- sleep 5
# block I/O
sudo perf record -e block:block_rq_issue -a -- sleep 10
perf trace is a syscall tracer with far lower overhead than strace:
sudo perf trace -p 1234
sudo perf trace -s -p 1234 # summary counts
When a profile is mostly kernel time, this is the follow-up: it tells you which syscalls, and the answer is usually to make fewer of them rather than to make each faster.
Off-CPU time, the thing perf record misses
A process blocked on I/O or a lock generates no samples, because it is not running. A flame graph of a program that spends its life waiting looks nearly empty, and the conclusion “the CPU is fine” is technically true and useless.
# scheduler latency, who is waiting and how long
sudo perf sched record -- sleep 10
sudo perf sched latency --sort max
For proper off-CPU analysis, bpftrace and the BCC tools are better suited, as our eBPF guide covers. The rule of thumb: perf for on-CPU, eBPF for off-CPU.
A realistic session
A service is using more CPU than it should.
# 1. what kind of busy is it
sudo perf stat -p "$(pgrep -f myservice)" -- sleep 10
# 2. where, cheaply
sudo perf record -F 99 -g -p "$(pgrep -f myservice)" -- sleep 30
perf report --stdio | head -25
# 3. if the stacks are broken, pay for DWARF
sudo perf record -F 99 --call-graph dwarf -p "$(pgrep -f myservice)" -- sleep 30
# 4. the shape
perf script | stackcollapse-perf.pl | flamegraph.pl > profile.svg
# 5. if it is mostly kernel time
sudo perf trace -s -p "$(pgrep -f myservice)" -- sleep 10
Five steps, each narrowing the last, and none of them requiring the program to be rebuilt or restarted.
What not to do
Do not profile a debug build. Unoptimised code has a completely different profile, and you will spend time on functions the compiler would have inlined out of existence.
Do not profile for two seconds. Short captures catch startup rather than steady state.
Do not trust a single run. Frequency scaling, cache state and other load all move the numbers. Our CPU governor guide covers pinning frequency while benchmarking.
Do not optimise a narrow tower because it is at the top of the graph. Width is the only thing that matters.
Clean up after yourself. perf record drops perf.data in the current directory, and those files are large.
Frequently Asked Questions
What is the difference between perf and strace?
strace intercepts system calls and shows what a program asks the kernel to do, at a large slowdown. perf samples the instruction pointer at intervals and shows where CPU time is spent, in kernel and userspace both, at a few percent overhead. Use strace for what a program is doing and perf for where it is slow.
Why does perf show hex addresses instead of function names?
Because the binary has no symbols. Install the debug symbol package for the distribution binary, or build with -g and without stripping for your own code. For JIT languages you need a runtime specific map file, and for anything static a rebuild is the only route.
What is a flame graph and how do I read one?
It is a visualisation of sampled stacks where width means share of samples and vertical position means call depth. Width is the only thing that matters: a wide box is where the time went. The horizontal ordering is alphabetical, so left to right carries no meaning at all.
Do I need root to use perf?
Profiling your own processes in userspace can work unprivileged depending on the perf_event_paranoid setting. Kernel symbols, tracepoints and system-wide profiling need elevated privileges or the cap_perfmon capability. Most distributions ship a restrictive default, so you will usually either lower that value or use sudo.
How much does perf slow down the program being profiled?
Sampling at the default rate typically costs a few percent, which is low enough to run on production. Raising the frequency, capturing full call graphs, or tracing every occurrence of a frequent event costs considerably more, so widen the scope only after a cheap run tells you where to look.
What does it mean when most samples land in the kernel?
It means the program is spending its time on system calls rather than computation, which points at I/O, locking, memory management or network work. That is the moment to look at which syscalls with perf trace or strace, since the fix is usually doing fewer of them rather than making each one faster.