eBPF and bpftrace Explained

eBPF and bpftrace Explained

eBPF runs sandboxed programs inside the Linux kernel. You attach a small program to a kernel function, a tracepoint, or a network hook, and it executes in kernel space without a module and without a reboot.

The reason this is not insane is the verifier. Before any eBPF program loads, the kernel analyses it and proves several properties: it terminates, it only accesses memory it is permitted to, and it cannot dereference a pointer that might be null. Programs that cannot be proven safe are rejected. That guarantee is what makes running arbitrary code in kernel context acceptable on a production machine.

What it is actually used for

eBPF underpins a lot of modern infrastructure that people use without knowing it:

Observability. bpftrace, BCC, and Pixie trace kernel and application behaviour at low overhead.

Networking. Cilium implements Kubernetes networking and policy with eBPF instead of iptables. Katran does load balancing at Facebook scale. XDP processes packets before the kernel network stack touches them, which is how you filter at millions of packets per second.

Security. Falco and Tetragon detect suspicious syscall patterns at runtime. seccomp filters, which confine containers and browsers, are classic BPF.

Performance. Profiling with negligible overhead, which is a different proposition from sampling profilers that interfere with what they measure.

Why not just use strace

strace works via ptrace, which stops the traced process at every syscall, copies data out, and resumes it. That is a context switch pair per call, and it can slow a syscall-heavy process by ten times or more. On a loaded production server, running strace can be the outage.

bpftrace runs in the kernel. No process stops, no ptrace, and it traces the whole system rather than one PID. It can also see things strace structurally cannot, because strace only observes the syscall boundary while eBPF can attach to internal kernel functions.

Our strace and ltrace guide covers where those tools still make sense, which is mostly interactive debugging of one process on a machine you do not mind slowing down.

Installing bpftrace

sudo apt install bpftrace        # Debian, Ubuntu
sudo dnf install bpftrace        # Fedora, RHEL
sudo pacman -S bpftrace          # Arch

sudo bpftrace -l | head          # list available probes
sudo bpftrace -l 'tracepoint:syscalls:*openat*'

bpftrace -l is where to start on any unfamiliar system. It enumerates every probe point the running kernel exposes, which on a modern kernel is tens of thousands.

The one-liners worth knowing

Which processes are opening files

sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat {
  printf("%s -> %s\n", comm, str(args->filename));
}'

comm is the process name. This answers “what is touching this file” and “what config is this daemon actually reading”, which is otherwise surprisingly hard to determine.

Count syscalls by process

sudo bpftrace -e 'tracepoint:raw_syscalls:sys_enter {
  @[comm] = count();
}'

Ctrl+C prints the map. @ denotes a map, and bpftrace prints all maps on exit. This is the fastest way to find which process is generating syscall load.

Latency distribution as a histogram

sudo bpftrace -e '
kprobe:vfs_read { @start[tid] = nsecs; }
kretprobe:vfs_read /@start[tid]/ {
  @us = hist((nsecs - @start[tid]) / 1000);
  delete(@start[tid]);
}'

This is the pattern that makes bpftrace valuable. kprobe fires on function entry, kretprobe on return, you store a timestamp keyed by thread ID, and hist() produces a log-scale histogram.

@us:
[0, 1)     12043 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@|
[1, 2)      3021 |@@@@@@@@@@                              |
[2, 4)       891 |@@@                                     |
[4, 8)       112 |                                        |
[8, 16)        4 |                                        |

An average would have told you reads take 0.4ms. The histogram tells you almost all are instant and a handful take 8 to 16 microseconds, which is a completely different diagnosis.

Who is executing what

sudo bpftrace -e 'tracepoint:syscalls:sys_enter_execve {
  printf("%s %d -> %s\n", comm, pid, str(args->filename));
}'

Every process launch, system-wide. Useful for finding what a cron job actually runs, and for noticing things you did not expect to be running at all.

TCP connections as they happen

sudo bpftrace -e 'kprobe:tcp_connect {
  printf("%s connecting\n", comm);
}'

Signals being sent

sudo bpftrace -e 'tracepoint:signal:signal_generate {
  printf("%s -> pid %d signal %d\n", comm, args->pid, args->sig);
}'

This answers “what keeps killing my process”, which is otherwise a genuinely difficult question. Our process signals guide covers what the numbers mean.

The language

bpftrace is deliberately awk-shaped: probe, optional filter, action block.

probe /filter/ { action }

Probe types worth knowing:

TypeFires on
tracepoint:Stable kernel instrumentation points
kprobe: / kretprobe:Any kernel function, entry and return
uprobe: / uretprobe:User-space functions in a binary
profile:hz:99Sampled, 99 times per second
interval:s:5Every 5 seconds

Prefer tracepoints over kprobes. Tracepoints are a stable interface the kernel commits to; kprobes attach to arbitrary internal functions that get renamed or inlined between versions, so a kprobe script can silently stop matching after a kernel upgrade.

Built-in variables: comm (process name), pid, tid, nsecs, cpu, uid, retval, args.

Aggregations: count(), sum(), avg(), min(), max(), hist(), lhist().

profile:hz:99 uses 99 rather than 100 deliberately, to avoid sampling in lockstep with periodic kernel activity that runs at 100Hz.

A real script

Slow reads only, with stack traces:

#!/usr/bin/env bpftrace

kprobe:vfs_read
{
    @start[tid] = nsecs;
}

kretprobe:vfs_read
/@start[tid]/
{
    $duration_ms = (nsecs - @start[tid]) / 1000000;
    if ($duration_ms > 10) {
        printf("%s (%d) slow read: %d ms\n", comm, pid, $duration_ms);
        print(kstack);
    }
    delete(@start[tid]);
}

END
{
    clear(@start);
}

$duration_ms is a scratch variable, @start is a map. kstack prints the kernel stack, which tells you not just that a read was slow but where in the kernel it was waiting.

Limits worth knowing

The verifier will reject things. Unbounded loops, large stack usage, and unchecked pointer access. The error messages are improving and are still frequently opaque.

Overhead is low, not zero. Attaching a probe to a function called millions of times per second costs real performance. Filter aggressively, and prefer aggregation in-kernel over printing every event to user space.

Root is required. Loading eBPF needs CAP_BPF and usually CAP_PERFMON. Unprivileged eBPF is disabled by default on most distributions, for good reasons involving the attack surface of a JIT compiler in kernel space.

Kernel version matters. Most tracing works on 4.9 and later. BTF and CO-RE, which let compiled tools run across kernels without recompiling against local headers, need 5.2 or later with BTF enabled. Any current distribution kernel is fine.

Portability is imperfect. A script using kprobes on internal functions may need adjusting between kernel versions. Tracepoint-based scripts are considerably more durable.

Where to go next

The bpftrace package ships a large collection of ready-made tools in /usr/share/bpftrace/tools: biolatency, execsnoop, opensnoop, tcpconnect, runqlat. Reading those scripts is the best available tutorial, because they are written by people who do this professionally and they are short enough to understand.

ls /usr/share/bpftrace/tools/
sudo /usr/share/bpftrace/tools/execsnoop.bt

For anything beyond one-liners and short scripts, BCC lets you write eBPF programs in C with a Python front end. More work, considerably more control, and the right tool once bpftrace stops being expressive enough.

Frequently Asked Questions

What is eBPF in simple terms?

eBPF lets you run small sandboxed programs inside the Linux kernel without writing a kernel module or rebooting. A verifier checks each program before loading it to prove it cannot crash the kernel or loop forever, which is what makes running arbitrary code in kernel space safe enough for production.

Is eBPF safe to use on a production server?

Yes, and that is largely the point. The verifier rejects any program that could crash the kernel, loop unboundedly, or access memory it should not. Overhead for typical tracing is low, often a few percent, though attaching high-frequency probes to very hot code paths can still cost measurable performance.

How is bpftrace different from strace?

strace uses ptrace to intercept syscalls, which stops the traced process on every call and can slow it by an order of magnitude. bpftrace runs in the kernel with no process stops, can trace the whole system at once rather than one process, and can observe kernel internals that strace cannot see at all.

Do I need root to run bpftrace?

Effectively yes. Loading eBPF programs requires CAP_BPF and usually CAP_PERFMON, and in practice that means running as root or granting those capabilities explicitly. Unprivileged eBPF is disabled by default on most distributions for security reasons.

What kernel version do I need for eBPF?

Most useful tracing works on 4.9 and later, with the ecosystem assuming 5.x or newer. BTF and CO-RE, which let compiled tools run across different kernel versions without recompiling, need 5.2 or later and a kernel built with BTF enabled. Any current distribution kernel is fine.

What is the difference between bpftrace and BCC?

bpftrace is a high-level tracing language for one-liners and short scripts, comparable to awk for kernel events. BCC is a toolkit for writing eBPF programs in C with a Python or C++ front end, which is more work and gives far more control. Start with bpftrace and move to BCC only when you need something it cannot express.