Hardening systemd Services: The Directives That Actually Confine
Our guide to writing service files covers getting a service running. This covers confining it.
systemd has accumulated a large set of sandboxing directives, built on the same kernel features containers use: namespaces, seccomp, and bind mounts. They are off by default, and most packaged units set very few.
Find out where you stand
systemd-analyze security
UNIT EXPOSURE PREDICATE
nginx.service 9.2 UNSAFE
sshd.service 9.6 UNSAFE
systemd-resolved.service 2.1 OK
Higher is worse. The first run on any system is usually sobering.
systemd-analyze security nginx.service
That prints every directive, whether it is set, and what each contributes to the score. It is the most useful documentation of these options that exists, because it tells you specifically what this unit is missing.
Two caveats. The score is a heuristic, not a measurement: a high number does not mean exploitable, and a low one does not mean safe. And systemd’s own units score well partly because systemd’s authors wrote both the units and the scoring.
The directives worth knowing
Filesystem
ProtectSystem=strict # entire filesystem read-only
ProtectHome=yes # /home, /root, /run/user inaccessible
PrivateTmp=yes # private /tmp and /var/tmp
ReadWritePaths=/var/lib/myapp /var/log/myapp
ProtectSystem=strict is the strongest of the three settings. yes makes /usr and /boot read-only; full adds /etc; strict makes everything read-only except /dev, /proc, and /sys.
With strict you must list what the service may write via ReadWritePaths. That is the point: you enumerate the writable surface instead of assuming it.
PrivateTmp=yes gives the service its own /tmp. This blocks a real and old class of attack, where a process predicts or races the temporary filename another process will use. It also means temporary files vanish on restart, which is usually desirable and occasionally surprising.
ProtectHome=yes matters more than it looks. A web server has no business reading /home, and our Flatpak permissions guide makes the same argument for desktop apps: home directories hold SSH keys, cloud credentials, and browser profiles.
Privileges
NoNewPrivileges=yes
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
AmbientCapabilities=CAP_NET_BIND_SERVICE
PrivateUsers=yes
NoNewPrivileges=yes is the single highest-value line here. It prevents the process and every child from ever gaining privileges, which disables setuid binaries entirely. An attacker with code execution in your service cannot escalate through sudo, pkexec, or a vulnerable setuid helper.
Set it unless the service genuinely needs to run a setuid binary, which almost none do.
CapabilityBoundingSet drops all Linux capabilities except those listed. Most services need none:
CapabilityBoundingSet=
The empty value drops everything. The common exception is CAP_NET_BIND_SERVICE, needed to bind ports below 1024, which is why web servers historically started as root. With ambient capabilities they no longer need to.
Kernel and devices
ProtectKernelTunables=yes # /proc/sys and /sys read-only
ProtectKernelModules=yes # cannot load modules
ProtectKernelLogs=yes # no dmesg access
ProtectControlGroups=yes # cgroup hierarchy read-only
ProtectProc=invisible # cannot see other processes
ProcSubset=pid # minimal /proc
PrivateDevices=yes # minimal /dev
ProtectProc=invisible is worth calling out. Without it, any process can read /proc entries for every other process on the system, including command lines that sometimes contain credentials. This restricts the service to seeing only its own.
ProtectKernelModules=yes blocks module loading, which removes one of the most direct paths from code execution to kernel compromise. Our kernel modules guide covers what that mechanism is.
Network
PrivateNetwork=yes # no network at all
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
IPAddressDeny=any
IPAddressAllow=10.0.0.0/8 localhost
PrivateNetwork=yes is the strongest confinement available for anything that does not need the network: a backup job, a report generator, a periodic cleanup task. The service gets an isolated namespace with only loopback.
RestrictAddressFamilies removes exotic socket types. AF_PACKET allows raw packet crafting and almost nothing legitimately needs it.
IPAddressDeny/IPAddressAllow applies BPF-based filtering per service, which is finer-grained than a host firewall can be.
System calls
SystemCallFilter=@system-service
SystemCallFilter=~@privileged @resources @obsolete
SystemCallArchitectures=native
@system-service is a curated set covering what a normal daemon needs. The ~ prefix removes groups from it.
SystemCallArchitectures=native blocks the 32-bit compatibility syscall interface on a 64-bit system. That interface has been a recurring source of kernel vulnerabilities and essentially no modern service needs it. One line, meaningful reduction.
DynamicUser
DynamicUser=yes
StateDirectory=myapp
LogsDirectory=myapp
CacheDirectory=myapp
systemd allocates a transient unprivileged user for the service’s lifetime. No persistent account, no useradd in a postinstall script, no orphaned files owned by a user nobody remembers creating.
It implies several other protections automatically, and it requires the service to store state through the *Directory options rather than arbitrary paths. Software that writes wherever it likes will not cooperate.
A worked example
[Unit]
Description=My application
After=network-online.target
[Service]
Type=exec
ExecStart=/usr/local/bin/myapp
DynamicUser=yes
StateDirectory=myapp
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectKernelLogs=yes
ProtectControlGroups=yes
ProtectProc=invisible
ProcSubset=pid
RestrictSUIDSGID=yes
RestrictRealtime=yes
RestrictNamespaces=yes
LockPersonality=yes
MemoryDenyWriteExecute=yes
CapabilityBoundingSet=
AmbientCapabilities=
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
SystemCallFilter=@system-service
SystemCallFilter=~@privileged @resources
SystemCallArchitectures=native
MemoryMax=512M
TasksMax=64
[Install]
WantedBy=multi-user.target
MemoryDenyWriteExecute=yes blocks memory that is both writable and executable, which defeats a broad class of exploit technique. It breaks JIT compilers, so anything running JavaScript, Java, or .NET will need it off.
The MemoryMax and TasksMax lines are cgroup limits, and they turn a runaway service into a contained failure rather than an OOM kill that takes down something else.
Do it incrementally
Adding all of the above at once and finding the service broken tells you nothing about which line did it.
sudo systemctl edit myapp # drop-in, leaves the packaged unit alone
sudo systemctl daemon-reload
sudo systemctl restart myapp
journalctl -u myapp -f
Add a few directives, restart, watch the journal. Failures are usually immediate and specific: a permission denied on a path, or a syscall being blocked.
# what actually got applied
systemd-analyze security myapp.service
sudo systemctl show myapp | grep -E 'Protect|Private|Restrict|NoNew'
Using systemctl edit rather than modifying the shipped unit means package updates do not overwrite your work, which is the same reasoning our Ansible guide applies to configuration generally.
Against containers
These directives use the same kernel mechanisms container runtimes use. For a service installed from your distribution, systemd hardening is frequently the better fit: no image to build, no registry, no runtime, and the service integrates with the journal and the rest of the system normally.
Containers additionally solve dependency packaging and distribution, which systemd does not attempt. Podman Quadlet sits between the two, defining containers as systemd units.
If you run services from packages, the hardening directives are available right now and mostly unused, and systemd-analyze security will tell you exactly where to start.
Frequently Asked Questions
What does systemd-analyze security tell me?
It scores every service on a scale from 0 to 10 based on which hardening directives are set, with higher meaning more exposed. Running it against a fresh system is usually uncomfortable, because most packaged units set very few of them and score in the unsafe range.
What is the difference between ProtectSystem and ProtectHome?
ProtectSystem makes system directories read-only, with strict mode making the entire filesystem read-only except a few writable paths. ProtectHome restricts access to home directories, with yes making them inaccessible and read-only making them visible but unwritable.
Will these directives break my service?
Some will, which is why you add them incrementally and test. The common breakages are a service that needs to write somewhere ProtectSystem made read-only, or one that needs a capability CapabilityBoundingSet removed. Failures are usually obvious and appear immediately in the journal.
What does NoNewPrivileges actually prevent?
It stops the process and all its children from ever gaining privileges, which means setuid binaries and file capabilities no longer take effect. An attacker who executes code in that service cannot escalate through a setuid helper such as sudo or a vulnerable setuid program.
Should I use DynamicUser for my services?
It is excellent where it fits. systemd allocates a transient unprivileged user for the service lifetime, so there is no persistent account and no stale files owned by it. It requires the service to store state through StateDirectory rather than arbitrary paths, which not all software does.
Is systemd hardening a replacement for containers?
It covers much of the same ground using the same kernel features, namespaces and seccomp, without an image or a runtime. For a service installed from your distribution it is frequently the better fit, while containers additionally solve dependency packaging and distribution, which systemd does not attempt.