io_uring Explained: Why Linux Needed a Third Async I/O Interface

io_uring Explained: Why Linux Needed a Third Async I/O Interface

Linux has had three attempts at asynchronous I/O. The first two were unsatisfying for reasons worth understanding, because they explain the shape of the third.

What came before

epoll is a readiness interface. It tells you a file descriptor is ready, and then you still make a syscall to do the actual read or write. For network sockets that is fine, because waiting is the expensive part. It never worked usefully for regular files, which are almost always “ready” in the sense epoll means and still block when you read them.

Linux AIO was a completion interface with conditions. It required O_DIRECT, worked on a subset of filesystems, covered a narrow set of operations, and silently fell back to blocking when its preconditions were not met.

That last property is what killed it. An asynchronous interface that becomes synchronous without telling you is worse than no interface, because your carefully non-blocking event loop stalls and nothing reports why. Almost nobody built on it.

The idea: shared memory instead of syscalls

io_uring’s insight is that the expensive part of a syscall-heavy workload is the syscalls, so remove them.

Two ring buffers are mapped into memory shared between your process and the kernel:

  • The submission queue (SQ), where you write requests
  • The completion queue (CQ), where the kernel writes results

Because the memory is shared, placing a request costs a memory write, not a syscall. You can queue many operations and then make a single io_uring_enter() call to tell the kernel they are there, or in polled mode, none at all.

your process                     kernel
     |                              |
     |--- write SQE to SQ ring ---->|  shared memory, no syscall
     |--- write SQE to SQ ring ---->|
     |--- io_uring_enter() -------->|  one syscall, many operations
     |                              |
     |<---- read CQE from CQ -------|  shared memory, no syscall

The syscall-per-operation cost becomes a syscall-per-batch cost, and in polled mode it disappears.

The three modes

Interrupt mode, the default. You submit, call io_uring_enter(), and the kernel completes work and notifies you. Already a large improvement, because one call carries a batch.

SQ polling. A kernel thread watches the submission queue. You write a request into shared memory and it gets picked up without any syscall at all. You are paying a CPU core to spin so that your I/O path has no kernel entry cost. That is a good trade at high request rates and a waste at low ones.

IOPOLL. For very fast storage, the kernel polls the device for completions rather than waiting for interrupts. An interrupt costs more than a poll when the device answers in microseconds.

What it can do

The operation list is broad and has kept growing: reads and writes, fsync, openat, statx, accept, connect, send, recv, splice, timeouts, and more.

Two features change how you structure code:

Linked operations. You can chain requests so one runs only if the previous succeeded, without a round trip to userspace between them. Open, read, close as a single submission.

Registered buffers and files. Pre-registering them lets the kernel skip per-operation setup work, which matters when you are doing millions of operations.

The part that is still being fixed

Some operations have no non-blocking path in the kernel at all: fsync, statx, openat and the *at family, xattr, fadvise, splice.

io_uring cannot risk blocking the submitting thread, so it hands those to a worker pool unconditionally, which costs a thread wakeup and a context switch per request. The frustrating part is that most of those calls do not actually block in practice. An fsync with nothing dirty, a statx that hits the cache, an openat creating a temporary file: all would have completed instantly inline.

Jens Axboe has proposed running them inline and only offloading if they genuinely block, using an unusual trick where a worker thread assumes the submitter’s identity and returns from the syscall in its place. It is an RFC, and it is the kind of change that illustrates how much overhead was being paid defensively.

Using it

Do not use the raw interface. liburing exists because the ring mechanics are fiddly and easy to get subtly wrong.

# Debian and Ubuntu
sudo apt install liburing-dev

# is it available and permitted
sysctl kernel.io_uring_disabled

You mostly do not need to write io_uring code at all. It is increasingly used underneath things you already run: modern database engines, high-performance proxies, and in-memory stores like the ones our self-hosted apps section covers.

The security tradeoff, stated plainly

io_uring has produced a disproportionate share of kernel vulnerabilities for its age. That is not surprising: it is a large new interface that reaches into storage, networking, filesystems and process management, and it lets userspace hand the kernel a lot of structured work.

The consequence is real. Some hardened distributions and container runtimes disable it by default, and Google restricted it on Android and ChromeOS for exactly this reason.

# 0 permitted, 1 restricted to CAP_SYS_ADMIN, 2 disabled
sysctl kernel.io_uring_disabled

# restrict it
sudo sysctl -w kernel.io_uring_disabled=2

If you run untrusted workloads, disabling it is a defensible choice, and our container hardening guide covers the seccomp filtering that can block it more selectively.

Should you care

Most applications should not change anything. Syscall overhead is rarely the bottleneck, and an application waiting on a database query or a slow disk gains nothing from a faster path to making the request.

It matters when you are doing an extremely large number of small operations, which means high-throughput storage, high-connection-count networking, and databases. Those are exactly the projects that have adopted it.

If you think you are in that population, measure first. Our fio benchmarking guide covers establishing where your I/O time is actually going, which is worth knowing before rewriting anything.

Frequently Asked Questions

What problem does io_uring solve that epoll did not?

epoll tells you when a file descriptor is ready, then you still make a syscall to do the work. io_uring lets you submit the work itself and collect results later, so one syscall can carry many operations. It also works for regular files, which epoll never usefully did.

Why did Linux AIO never get adopted?

It only worked with O_DIRECT on a subset of filesystems, silently fell back to blocking behaviour outside those conditions, and covered a narrow set of operations. An interface that quietly becomes synchronous when its preconditions are not met is worse than no interface, so almost nobody built on it.

What are the submission and completion queues?

Two ring buffers in memory shared between your process and the kernel. You write requests into the submission queue and read results from the completion queue. Because the memory is shared, adding a request or reading a result does not require a syscall at all.

Can io_uring really do I/O without any syscalls?

In polled mode, yes. A kernel thread watches the submission queue, so you place a request in shared memory and it is picked up without io_uring_enter being called. You spend a CPU core on polling in exchange for removing syscall overhead entirely.

Why do some distributions and container platforms disable io_uring?

It has produced a disproportionate number of kernel vulnerabilities for its age, because it is a large new interface reaching deep into many subsystems. Some hardened environments and container runtimes block it by default, trading its performance for a smaller attack surface.

Should I rewrite my application to use io_uring?

Only if syscall overhead is genuinely your bottleneck, which is rare outside high-throughput storage and networking. Most applications are limited by something else entirely. Use liburing rather than the raw interface if you do, and measure before and after rather than assuming.