trap and Signal Handling in Shell Scripts

trap and Signal Handling in Shell Scripts

A script that creates a temporary file and exits without removing it is a small leak. A hundred runs later it is a full /tmp.

trap fixes that in one line, and most scripts that use it use it wrong.

The one trap that matters

#!/usr/bin/env bash
set -euo pipefail

tmpdir=$(mktemp -d)
trap 'rm -rf "$tmpdir"' EXIT

# do work in "$tmpdir"

EXIT fires however the script ends. Normal completion, an error under set -e, Ctrl+C, a kill. One handler covers every case where the shell terminates.

That is why most scripts need only this. People write elaborate trap ... INT TERM HUP chains when EXIT already covers all of them.

Set the trap on the line immediately after creating the thing it cleans up. Any gap is a window where the script can exit having created the directory with no handler registered.

Signals worth knowing

SignalNumberSent by
INT2Ctrl+C
TERM15kill, systemd stopping a service
HUP1Terminal closing, or “reload config” by convention
QUIT3Ctrl+\
KILL9kill -9, uncatchable
USR1/USR210/12Whatever you decide

Our process signals guide covers the wider mechanism.

KILL and STOP cannot be trapped. That is deliberate, so there is always a way to stop a process. It also means cleanup can never be guaranteed, which is why anything important needs to be recoverable on the next run rather than dependent on an exit handler.

A cleanup function

#!/usr/bin/env bash
set -euo pipefail

readonly LOCKFILE=/var/run/myjob.lock
tmpdir=""

cleanup() {
    local rc=$?
    trap - EXIT

    [[ -n "$tmpdir" && -d "$tmpdir" ]] && rm -rf "$tmpdir"
    [[ -e "$LOCKFILE" ]] && rm -f "$LOCKFILE"

    if (( rc != 0 )); then
        echo "failed with exit code $rc" >&2
    fi
    exit "$rc"
}
trap cleanup EXIT

echo $$ > "$LOCKFILE"
tmpdir=$(mktemp -d)

Three details that matter.

local rc=$? first. Capture the exit status before anything else, because every command in the handler overwrites $?.

trap - EXIT at the start. This removes the handler so it cannot recurse. Without it, a failure inside the handler can trigger it again, and you get an infinite loop.

exit "$rc" at the end. Preserve the original exit status rather than returning the handler’s.

Our exit codes guide covers why that status matters, and our bash error handling guide covers set -euo pipefail.

Handling Ctrl+C differently

Sometimes you want a distinct response to an interrupt:

interrupted=0

on_interrupt() {
    interrupted=1
    echo "" >&2
    echo "Interrupted, finishing current item..." >&2
}
trap on_interrupt INT

for item in "${items[@]}"; do
    (( interrupted )) && break
    process "$item"
done

The loop finishes its current item and stops cleanly rather than dying mid-operation. For a script writing files or updating a database, that is the difference between a clean stop and a half-written record.

A second Ctrl+C should still work. If the first is handled gracefully, make the second forceful:

on_interrupt() {
    if (( interrupted )); then
        echo "Forced." >&2
        exit 130
    fi
    interrupted=1
    echo "Interrupted, finishing current item. Ctrl+C again to force." >&2
}

130 is the conventional exit code for termination by SIGINT: 128 plus the signal number.

Reload on HUP

The daemon convention:

reload_config() {
    echo "reloading configuration" >&2
    source /etc/myapp/config
}
trap reload_config HUP

while true; do
    do_work
    sleep 60
done
kill -HUP $(cat /var/run/myjob.pid)

For a systemd service, that pairs with:

[Service]
ExecReload=/bin/kill -HUP $MAINPID

Waiting and signals

A subtlety that catches people.

# bad: the trap does not run until sleep finishes
trap 'echo caught' INT
sleep 300

Bash does not run a trap handler while waiting for a foreground command. The signal is noted and the handler runs after sleep returns, which for a five-minute sleep means five minutes.

# good: background it and wait
trap 'echo caught; kill "$pid" 2>/dev/null' INT
sleep 300 &
pid=$!
wait "$pid"

wait is interruptible, so the handler runs immediately.

This matters for any long-running script that should respond promptly to a stop request, including anything systemd might send TERM to before escalating to KILL after TimeoutStopSec.

Locking

readonly LOCKFILE=/var/run/myjob.lock

if ! mkdir "$LOCKFILE" 2>/dev/null; then
    echo "already running" >&2
    exit 1
fi
trap 'rmdir "$LOCKFILE"' EXIT

mkdir is atomic, which makes it a correct lock where [ -e file ] && touch file is a race.

The better answer for a scheduled job:

flock -n /var/run/myjob.lock /usr/local/bin/myjob.sh

Or use a systemd timer, which will not start a second instance while the first is running and needs no lockfile at all.

Common mistakes

Single quotes versus double.

trap "rm -rf $tmpdir" EXIT    # expands NOW, likely empty
trap 'rm -rf "$tmpdir"' EXIT  # expands when the trap fires

The first captures whatever $tmpdir held at trap time, which if you set the trap before creating the directory is an empty string. rm -rf "" is harmless, and rm -rf "$tmpdir"/* with an empty variable is not.

Forgetting subshells do not inherit traps.

trap cleanup EXIT
( some_command )   # no trap in here

Assuming cleanup always runs. kill -9, a power failure, or the OOM killer leave no opportunity. Design so a leftover temporary file is tolerable and the next run recovers.

Frequently Asked Questions

What is the simplest useful trap in a shell script?

trap cleanup EXIT, which runs your cleanup function however the script ends, whether normally, through an error under set -e, or after an interrupt. One line removes the entire class of bugs where a script leaves temporary files behind.

Why does my trap not run when I press Ctrl+C?

An EXIT trap does run on Ctrl+C, because the shell exits. If you trapped only specific signals such as TERM and not INT, then Ctrl+C is not covered. Trapping EXIT alone covers every case where the shell terminates normally.

How do I make a script clean up a temporary directory reliably?

Create it with mktemp -d, then immediately set an EXIT trap that removes it. Setting the trap on the line after creation means there is no window where the script can exit having created the directory but not yet registered the cleanup.

What is the difference between trapping INT and trapping TERM?

INT is sent by Ctrl+C from the terminal. TERM is the default signal from kill and what systemd sends when stopping a service. A script that might run interactively and as a service should handle both, or trap EXIT which covers each.

Can I trap SIGKILL?

No. SIGKILL and SIGSTOP cannot be caught, blocked, or ignored, by design, so there is always a way to stop a process. This means cleanup cannot be guaranteed and anything critical needs to be recoverable on the next run rather than relying on an exit handler.

Why should a trap handler reset the trap before exiting?

To avoid recursion, where the handler triggers the same condition and calls itself indefinitely. Setting trap - EXIT at the start of the handler, or using a guard variable, prevents a cleanup failure from turning into an infinite loop.