Jens Axboe's 'Crazy' io_uring Patches Swap a Thread's Identity Mid-Syscall

Jens Axboe's 'Crazy' io_uring Patches Swap a Thread's Identity Mid-Syscall

Jens Axboe, block subsystem maintainer and io_uring lead, posted an RFC he described himself as “crazy patches.” The description is fair.

The waste being removed

io_uring tries to run requests inline with IO_URING_F_NONBLOCK and offloads to a worker pool when that is not possible.

For a set of operations it is never possible, because the kernel has no non-blocking path for them at all:

  • fsync
  • statx
  • openat and the rest of the *at family
  • xattr
  • fadvise
  • splice

Those get offloaded unconditionally. In Axboe’s words:

“Those are punted unconditionally, and the punt costs a thread wakeup, a context switch and a task_work completion round trip per request. io_uring HAS to be cautious to prevent accidental blocking in the kernel, even if the operations predominantly never block. Sad story.”

That last sentence is the whole problem. Most of these calls do not actually block. An fdatasync with nothing dirty to write. A statx that hits the dcache. An openat creating an O_TMPFILE. Every one of those would have completed instantly inline, and every one paid for a thread wakeup, a context switch and a completion round trip because the kernel could not promise in advance that it would not block.

io_uring is paying the worst-case cost on every request to protect against a case that rarely happens.

The identity swap

The new approach runs the request inline in blocking mode, and only pays the offload cost if it actually blocks.

The difficulty is what to do at the moment it does block. By then the submitting task is deep inside the kernel with the request on its own stack. You cannot hand that work to another thread, because the work is the current thread’s stack.

So Axboe moves the other thing instead:

“What we can move is the identity. If the submitting task blocks, an idle io-wq worker takes over its user visible identity (tid, signal state, credentials, scheduling attributes, cgroup, user register state), finishes the io_uring_enter() call and returns to userspace as the submitter. The original task finishes the request as an io-wq worker and joins the pool. Userspace is none the wiser, hopefully, the same tid came back from the syscall, it’s just on a different task_struct.”

The two threads trade places. The worker assumes the submitter’s tid, credentials, signal state, scheduling attributes, cgroup membership and user register state, then returns from the syscall as though it had been the caller all along. The original thread stays behind to finish the blocking work and then joins the worker pool.

From userspace, the same thread ID returns from io_uring_enter(). Underneath, it is a different task_struct.

Axboe notes that people who have been around a while may remember similar attempts roughly twenty years ago, which is a polite way of flagging that this idea has a history and the history is not uniformly happy.

Why it is audacious

The list of things that have to be transferred correctly is the list of everything userspace can observe about a thread. Miss one and you get a bug that appears only when a specific operation blocks at a specific moment, which is close to the worst debugging scenario available.

cgroup membership is a good example of why it is delicate. If the worker inherits the wrong cgroup, the accounting and limits described in our cgroups guide silently apply to the wrong container. Credentials are worse: getting those wrong is a privilege boundary problem, not a performance one.

The upside is proportionate. Removing a thread wakeup and context switch from every fsync and statx in an io_uring workload is a large structural saving for exactly the applications that chose io_uring for latency.

Who this reaches

Databases, object stores, proxies and anything that adopted io_uring for high-throughput I/O. Dragonfly and similar in-memory stores sit squarely in that population, as do modern web servers doing heavy file serving.

It joins an unusually busy month of kernel performance work: 39% faster file opens, kernel builds up to 70% faster incrementally, and faster hibernation.

# is io_uring available and unrestricted on your system
sysctl kernel.io_uring_disabled

# what your kernel supports
grep -i io_uring /proc/kallsyms | head -3

Worth noting that some distributions and container platforms restrict io_uring for security reasons, since it has had a meaningful share of kernel vulnerabilities. That tension is not resolved by this work and is a separate decision from whether it is fast.

Status

An RFC. Axboe himself calls it crazy, and whether it reaches mainline for 7.4 depends on review going better than experience suggests it might. The idea is not obviously safe, and it is obviously interesting.

Background reading

Explainers for the concepts behind this story.