strace and ltrace Explained
When a program fails with a vague or missing error message, strace and ltrace answer the question ordinary debugging often can’t: what is this program actually doing, at the level of system and library calls, right up to the moment it fails. For diagnosing permission errors, missing files, unexpected network behavior, or a program that simply hangs with no explanation, they are often the fastest path to the real cause.
strace: tracing system calls
strace intercepts and logs every system call a program makes, the requests it sends directly to the kernel for things like opening files, reading and writing data, and network operations.
strace ls /nonexistent
openat(AT_FDCWD, "/nonexistent", O_RDONLY|O_NONBLOCK|O_CLOEXEC|O_DIRECTORY) = -1 ENOENT (No such file or directory)
Each line shows the call name, its arguments, and its return value. This example shows exactly why ls /nonexistent failed: the openat call returned -1 with error ENOENT (no such file or directory), directly on the path being opened. For a real, more mysterious failure, this same pattern, scanning for a call that returned a negative value with an error code, is usually the fastest way to find the actual problem.
Attaching to a running process
sudo strace -p 1234
-p attaches to an already-running process by PID instead of starting a fresh one, useful for diagnosing a process that is already hung or misbehaving rather than one that fails immediately at startup. This typically requires root, or being the owner of the target process, since attaching a tracer is a privileged operation. Press Ctrl+C to detach without affecting the traced process.
Filtering output
Unfiltered strace output for anything beyond a trivial program can be enormous. Narrow it down with -e trace=:
strace -e trace=open,openat,read myprogram # only file access calls
strace -e trace=network myprogram # only network-related calls
strace -e trace=process myprogram # only process management calls (fork, exec, etc.)
Matching the filter category to the symptom you are chasing, file operations for a “file not found” mystery, network for a connectivity problem, dramatically cuts down the noise you have to scan through.
A realistic debugging pattern: permission denied
A program failing with “permission denied” despite ls showing you have access is one of the most useful strace patterns:
strace -e trace=open,openat myprogram 2>&1 | grep EACCES
openat(AT_FDCWD, "/var/lib/app/data.db", O_RDWR) = -1 EACCES (Permission denied)
This immediately shows the exact file the program was trying to access when it failed, which is often not the file you assumed. Common real causes surfaced this way: a parent directory missing execute permission (needed to traverse into it, distinct from read permission on the file itself), an SELinux or AppArmor denial that plain Unix permission checks never reveal, or the process actually running as a different user than expected, commonly the case for services running under a dedicated system account rather than your own user.
Finding what files a program actually reads
strace -e trace=open,openat -f myprogram 2>&1 | grep -v ENOENT
-f follows child processes as well, useful for programs that fork subprocesses, since without it, strace only tracks the initial process, missing any file access performed by children it spawns. This pattern is useful for discovering exactly which configuration files, libraries, or data files a program is actually reading, especially useful when tracking down which config file a program is really loading when multiple candidate locations exist.
ltrace: tracing library calls
ltrace traces library calls instead of system calls, functions the program calls from shared libraries like libc, sitting at a somewhat higher level than raw kernel interaction:
ltrace myprogram
strlen("hello world") = 11
malloc(256) = 0x55a1b2c3d4e0
ltrace is generally more useful for understanding a program’s internal logic, memory allocation patterns, or string handling, than for diagnosing the kind of environmental problems (missing files, permission issues, network failures) that strace is typically reached for first. In day-to-day troubleshooting, strace sees far more use.
Performance impact
Both tools add real overhead: every traced call is intercepted, which can make a program run several times slower than normal, more for programs making a very high volume of calls. This is acceptable for interactive debugging of a specific issue, but neither tool should be left running continuously in production. It is also worth keeping in mind that a timing-sensitive program may behave slightly differently under tracing than without it, which occasionally means a traced run does not reproduce the exact symptom you were chasing.
Frequently Asked Questions
What is the difference between strace and ltrace?
strace traces system calls, the requests a program makes directly to the kernel: opening files, reading and writing data, network operations, and process management. ltrace traces library calls instead, the functions a program calls from shared libraries like libc, which sit at a higher level than raw system calls. strace is generally more useful for diagnosing issues like permission errors, missing files, or network problems, since those are fundamentally kernel-level operations. ltrace is more useful for understanding a program’s internal logic and which library functions it relies on, though it is less commonly reached for in day-to-day troubleshooting than strace.
How do I trace a program that is already running?
Use strace -p PID to attach to an already-running process by its process ID, rather than starting a new process under strace from scratch. This requires appropriate permissions, typically root or being the owner of the target process, since attaching a tracer to a process is a privileged operation for security reasons; an arbitrary user attaching to and observing another user’s process would be a significant information disclosure risk otherwise. Detach with Ctrl+C, which stops tracing without affecting the traced process itself.
Why is my program failing with “permission denied” but ls shows I have access?
This is one of the most common and useful strace diagnostic patterns: run the failing program under strace and search the output for the specific syscall that returned a permission error, such as openat(…) = -1 EACCES. The exact file or resource that call was operating on when it failed is usually right there in the same line, which frequently reveals the real cause: a parent directory without execute permission, an unexpected SELinux or AppArmor denial that manual permission checks with ls do not reveal at all, or the process actually running as a different user than expected.
What does the output of strace actually look like?
Each line represents one system call, showing the call name, its arguments, and its return value, in roughly this format: openat(AT_FDCWD, “/etc/config.conf”, O_RDONLY) = 3. This line shows the program attempting to open /etc/config.conf for reading, which succeeded and returned file descriptor 3. A failed call instead shows a negative return value and an error code, such as = -1 ENOENT (No such file or directory), which directly identifies both what the program was trying to do and exactly why it failed.
How do I filter strace output to just the calls I care about?
Use -e trace=CATEGORY to limit output to a specific class of system calls, for example strace -e trace=open,openat,read program to see only file access calls, or strace -e trace=network program to see only network-related calls. Without filtering, strace output for anything beyond a trivial program can be enormous and hard to scan through manually, so filtering to the relevant category (file access for a “file not found” mystery, network for a connectivity issue) makes the output dramatically more useful for the specific problem being diagnosed.
Does using strace or ltrace slow down the program being traced?
Yes, noticeably. Tracing intercepts every traced call, which adds meaningful overhead, often making the traced program run several times slower than normal, sometimes much more for programs that make a very high volume of system or library calls. This is generally fine for interactive debugging of a specific problem, but strace and ltrace are not tools to leave running continuously in production, and a program that is timing-sensitive (such as one with tight real-time constraints) may behave differently or fail differently under tracing than it would without it, which is worth keeping in mind when a traced run does not reproduce the exact same symptom as an untraced one.