sched_ext Explained: Swapping the Linux CPU Scheduler at Runtime With BPF

sched_ext Explained: Swapping the Linux CPU Scheduler at Runtime With BPF

Every Linux system has a CPU scheduler: the part of the kernel that decides which runnable task gets a CPU next, and for how long. For almost all of Linux’s history there has been one general-purpose choice, and changing it meant patching and rebuilding the kernel.

sched_ext changes that. Since Linux 6.12, you can load a completely different scheduler, written as a BPF program, into a running kernel, try it under a real workload, and unload it again, all without rebooting.

Why schedulers are hard

A scheduler is a set of trade-offs that cannot all be won at once:

  • Throughput: keep every CPU busy and minimize time wasted switching tasks
  • Latency: when something important wakes up, such as a game’s render thread or your mouse input, run it immediately
  • Fairness: do not let one task starve another
  • Topology: keep tasks near their cached data, and on modern CPUs, decide between performance and efficiency cores or between chiplets that do not share a cache

The kernel’s default, EEVDF (Earliest Eligible Virtual Deadline First, which replaced CFS in Linux 6.6), is a very good compromise for everything from phones to database servers. A compromise, by definition, is not optimal for any one workload. Our guide to understanding Linux processes covers scheduling basics, and nice, renice, and ionice covers the priority knobs the default scheduler offers.

How sched_ext works

sched_ext adds a new scheduling class to the kernel whose decisions are made by BPF programs. A scheduler is split into two halves:

  1. A BPF part, loaded into the kernel, that makes the hot-path decisions: which queue a waking task goes on, which CPU picks it up, and how long its time slice is
  2. A userspace part, a normal program, that loads the BPF code and can feed it higher-level information, such as CPU topology, or run slower policy calculations

BPF is the same in-kernel virtual machine used for tracing and networking. The key property is that the kernel verifies BPF programs before running them: they cannot crash the kernel, read arbitrary memory, or loop forever. Our eBPF and bpftrace guide covers BPF in general.

The safety net

A scheduler that is merely verified can still make terrible decisions, such as never running a particular task. sched_ext handles that with a watchdog: if any runnable task waits too long, the kernel ejects the BPF scheduler and moves every task back to EEVDF automatically. The same happens if the userspace loader exits or crashes.

There is also an emergency switch: Alt+SysRq+S forces sched_ext off.

This safety net is the whole reason sched_ext is practical. Experimenting with a scheduler used to mean risking a hung machine. Now the worst case is a stall of a few seconds followed by a fallback.

Checking support

Your kernel needs to be built with CONFIG_SCHED_CLASS_EXT. Most distributions enable it on kernels from 6.12 onward: Arch, Fedora, CachyOS, and openSUSE Tumbleweed all do.

cat /sys/kernel/sched_ext/state
# disabled   (supported, nothing loaded)

cat /sys/kernel/sched_ext/root/ops 2>/dev/null
# scx_lavd   (when one is running)

If /sys/kernel/sched_ext does not exist, your kernel does not support it.

The schedulers

The main collection lives in the scx project, packaged on most distributions as scx-scheds. Each scheduler is a separate scx_* binary. The most widely used:

SchedulerAimed atIdea
scx_lavdGaming, interactive desktops, handheldsEstimates how latency-critical each task is from its wake and block patterns, and gives critical tasks earlier deadlines. Built by Igalia for Valve and the Steam Deck
scx_bpflandInteractive desktopsPrioritises tasks that block often (interactive ones) over CPU-bound ones
scx_rustyGeneral purpose, multi-socket and chiplet CPUsSplits scheduling between fast BPF decisions and topology-aware load balancing in userspace (Rust)
scx_flashLow-latency workloadsDeadline-based, emphasises predictable latency

There are many more, including experimental ones. That experimentation is the point: new ideas can be tested on real hardware in an afternoon, and the good ones feed back into how the default scheduler evolves.

scx_lavd in more detail

LAVD stands for Latency-criticality Aware Virtual Deadline. The insight is that you can infer which tasks matter for responsiveness from their behaviour. A game’s render thread wakes up, does some work, hands a frame to the GPU, and blocks. It sits in the middle of a chain of tasks that wake each other. LAVD measures these wake-up and blocking relationships, scores each task’s latency criticality, and uses that score to set deadlines and time slices.

The target is not higher average FPS. It is fewer frame time spikes and better 1% lows when the CPU is contended, for example while Steam compiles shaders in the background or a browser is open on another monitor.

Trying one

# Arch / CachyOS
sudo pacman -S scx-scheds

# Fedora (via COPR or the distro package, depending on release)
sudo dnf install scx-scheds

# Run one in the foreground; Ctrl+C returns you to EEVDF
sudo scx_lavd

To make it persistent, the scx project ships a systemd service and a small loader (scx_loader, driven by scxctl) that lets you switch schedulers without restarting anything:

scxctl start --sched lavd
scxctl switch --sched bpfland
scxctl stop

Some gaming-focused distributions, including CachyOS, expose this in a GUI.

Measuring whether it helped

Do not judge by feel. Scheduler differences are small and placebo is large.

  • For games, use MangoHud and compare 1% and 0.1% low frame times over the same repeatable scene, with a realistic background load
  • For desktops, compare responsiveness under a deliberate load, such as a kernel build with make -j$(nproc) in the background
  • For servers, measure the actual service latency percentiles you care about

Run each test several times with each scheduler. If you cannot tell them apart in the numbers, stay on the default.

When to use it

Worth trying on a gaming PC or handheld, especially one with many cores doing background work, and on a desktop that stutters under load. Also on unusual hardware, such as CPUs with a mix of core types, where the default heuristics may not be ideal.

Not worth it on a server that is working fine, where EEVDF is heavily tuned and predictable, or on a mostly idle desktop where there is nothing to schedule around.

sched_ext does not replace understanding your workload. If a machine stutters because it is swapping or waiting on disk, no CPU scheduler will fix it. Check CPU monitoring and load average first to confirm the CPU is actually the bottleneck.

Frequently Asked Questions

What is sched_ext?

sched_ext is a Linux kernel feature, merged in version 6.12, that lets a CPU scheduler be written as a BPF program and loaded while the system is running. When the BPF scheduler is unloaded or fails, the kernel automatically returns to its built-in scheduler.

Is it safe to try a sched_ext scheduler?

It is designed to be. BPF programs are verified before loading, and the kernel runs a watchdog that ejects any sched_ext scheduler that leaves a runnable task waiting too long. The worst realistic outcome is a brief stall followed by an automatic fallback to the default scheduler.

What is scx_lavd?

scx_lavd is a sched_ext scheduler developed by Igalia for Valve, originally targeting the Steam Deck. It estimates how latency-critical each task is from its wake-up and blocking patterns and uses that to set deadlines and time slices, aiming for smoother frame pacing and better 1 percent lows in games.

How do I check whether a sched_ext scheduler is running?

Read /sys/kernel/sched_ext/state, which shows enabled or disabled, and /sys/kernel/sched_ext/root/ops, which shows the name of the loaded scheduler. If the directory does not exist, your kernel was built without sched_ext support.

Will sched_ext make my games faster?

It rarely raises average frame rates. What it can improve is consistency: fewer stutters and better minimum frame times when the CPU is busy with background work such as shader compilation or a browser. On an idle system with a fast CPU the difference is often too small to notice.

How do I stop a sched_ext scheduler?

Stop the userspace program that loaded it, for example with Ctrl+C or by stopping its systemd service. The kernel immediately switches all tasks back to the default EEVDF scheduler. The SysRq key combination Alt+SysRq+S also forces sched_ext off in an emergency.