Container Security Hardening: What the Defaults Leave Open
A default docker run gives the container more than most people realise. Our Docker and Podman basics covers getting containers working; this covers making them a boundary.
What you get by default
docker run -d nginx
That container runs as root inside, keeps roughly fourteen Linux capabilities, has a writable filesystem, unrestricted outbound network, and no memory or process limits.
The capabilities retained by default include:
CAP_CHOWN change file ownership
CAP_DAC_OVERRIDE bypass file permission checks
CAP_SETUID/SETGID change process UID/GID
CAP_NET_RAW raw sockets, allows packet crafting
CAP_MKNOD create device nodes
Very few applications need any of them. CAP_DAC_OVERRIDE alone lets a process ignore file permissions entirely, and CAP_NET_RAW has been used for ARP spoofing from inside containers.
The flags that matter
docker run -d \
--user 1000:1000 \
--read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=64m \
--cap-drop=ALL \
--cap-add=NET_BIND_SERVICE \
--security-opt=no-new-privileges:true \
--memory=512m \
--pids-limit=100 \
-p 127.0.0.1:8080:80 \
myimage
—user
The highest-value single change.
Without user namespaces, root inside the container is UID 0 on the host. A container escape lands you as host root. Running as an unprivileged UID means an escape lands as an unprivileged user.
Better to bake it into the image so it is not something a deployment can forget:
RUN useradd --system --uid 1000 --no-create-home appuser
USER 1000:1000
—read-only
An immutable root filesystem. Malware that cannot write cannot persist.
Almost everything needs somewhere writable, so pair it with tmpfs mounts scoped tightly:
--tmpfs /tmp:rw,noexec,nosuid,size=64m
--tmpfs /run:rw,noexec,nosuid,size=16m
noexec on those is the part worth keeping: a writable directory that cannot execute is considerably less useful to an attacker.
—cap-drop=ALL
Drop everything, add back what breaks. Usually nothing. NET_BIND_SERVICE if it binds below port 1024, though mapping a high port is cleaner.
—security-opt=no-new-privileges
Same mechanism as systemd’s NoNewPrivileges, covered in our service hardening guide. Setuid binaries inside the image stop working as an escalation path.
Resource limits
--memory=512m --memory-swap=512m --pids-limit=100 --cpus=1.5
Without these one container can consume the host. --pids-limit specifically blocks fork bombs. Setting --memory-swap equal to --memory disables swap for the container, which prevents a memory-hungry container from thrashing the host’s disk.
These are cgroup limits, the same mechanism systemd uses.
The networking trap
This one catches people who believe they are protected:
docker run -p 8080:80 myimage
That is reachable from the internet, regardless of your firewall. Docker inserts its own iptables rules ahead of the chains ufw and firewalld manage, so a published port is open even when your firewall rules say it should not be.
sudo iptables -L DOCKER -n
sudo ss -tlnp | grep 8080
Two fixes. Bind explicitly to loopback:
-p 127.0.0.1:8080:80
Or do not publish at all, put the container on an internal network, and reach it through a reverse proxy that you control the exposure of.
services:
app:
networks: [internal]
# no ports section at all
proxy:
networks: [internal, external]
ports:
- "127.0.0.1:443:443"
networks:
internal:
internal: true
external:
internal: true gives that network no route out, so a compromised application container cannot reach the internet to download a payload or exfiltrate data.
Compose, with everything applied
services:
app:
image: myimage:1.2.3
user: "1000:1000"
read_only: true
tmpfs:
- /tmp:rw,noexec,nosuid,size=64m
cap_drop: [ALL]
security_opt:
- no-new-privileges:true
mem_limit: 512m
pids_limit: 100
restart: unless-stopped
volumes:
- ./data:/data
- /etc/localtime:/etc/localtime:ro
networks: [internal]
Note image: myimage:1.2.3 rather than :latest. A floating tag means you cannot say what is running or reproduce it, and a compromised registry account can replace what latest points at.
Note ./data:/data without :ro because it needs writing, and /etc/localtime:ro because it does not. Mount read-only by default and make writable the exception.
Our Docker Compose guide and the compose generator cover the basics.
Rootless Podman
The structural improvement rather than an incremental one.
podman run -d --userns=keep-id myimage
podman info | grep -i rootless
Docker runs a daemon as root; membership of the docker group is effectively root on the host, which surprises people. Podman runs rootless: no daemon, and the entire runtime executes as your unprivileged user with user namespaces mapping container root to your host UID.
A container escape from rootless Podman lands as you, not as root. That is a categorically different outcome.
Podman is otherwise Docker-compatible enough that alias docker=podman works for most use, and Quadlet integrates containers as systemd units.
Images
Pin versions and digests. image: nginx:1.27.3@sha256:... is reproducible; nginx:latest is not.
Use minimal bases. Alpine, distroless, or scratch for a static binary. Fewer packages means fewer CVEs and less for an attacker to use. A scratch image containing one static binary has no shell to drop into.
Multi-stage builds keep build tooling out of the runtime image:
FROM golang:1.25 AS build
WORKDIR /src
COPY . .
RUN CGO_ENABLED=0 go build -o /app
FROM scratch
COPY --from=build /app /app
USER 65534:65534
ENTRYPOINT ["/app"]
Scan, repeatedly. An image is not static in its risk profile: it accumulates known vulnerabilities as its dependencies age.
trivy image myimage:1.2.3
grype myimage:1.2.3
Scan in CI and rebuild on a schedule. An image untouched for six months is very likely vulnerable to something disclosed since.
Our building container images guide covers the wider practice.
Being honest about the boundary
Containers share the host kernel. A kernel vulnerability reachable from inside a container is a host vulnerability, and that is not something container configuration can fix.
For genuinely untrusted code, a virtual machine is the stronger isolation because the boundary is the hypervisor rather than the syscall interface. gVisor and Kata Containers exist to narrow the gap and carry their own tradeoffs.
For your own applications, hardened containers are a real and worthwhile improvement over the defaults. The gap between docker run -d myimage and the configuration above is large, and closing it costs a few lines.
Frequently Asked Questions
Is a container a security boundary?
A well-configured one is a meaningful boundary, and a default one is weaker than most people assume. Containers share the host kernel, so a kernel vulnerability reachable from inside the container affects the host. For untrusted code a virtual machine remains the stronger isolation.
Why should containers not run as root?
Without user namespaces, root inside the container is UID 0 on the host, so a container escape lands as host root. Running as an unprivileged user means an escape lands as an unprivileged user, which is a substantially smaller problem.
What is the difference between rootless Podman and running a container as a non-root user?
Running as a non-root user changes who the process inside is. Rootless Podman means the whole container runtime runs without privilege, using user namespaces so container root maps to your unprivileged host UID. Rootless is the stronger property because there is no privileged daemon at all.
What does —cap-drop=ALL actually remove?
It drops every Linux capability the container would otherwise keep, which by default includes roughly fourteen including CAP_CHOWN, CAP_SETUID, and CAP_NET_RAW. Most applications need none of them, and you add back only what fails, usually NET_BIND_SERVICE for ports below 1024.
Why does publishing a port bypass my firewall?
Docker inserts its own iptables rules ahead of your chains, so a published port is reachable even when ufw or firewalld says otherwise. Binding explicitly to 127.0.0.1 in the port mapping avoids this, as does putting a reverse proxy in front rather than publishing directly.
Do I need to scan container images?
Yes, and repeatedly rather than once. An image built today from a current base accumulates known vulnerabilities as its dependencies age, so an image untouched for six months is very likely vulnerable. Scanning in CI and rebuilding on a schedule addresses it.