sched_ext lets a runtime-loaded BPF program provide Linux task-scheduling policy through callbacks in struct sched_ext_ops. The kernel supplies the framework; your scheduler decides how tasks are placed and dispatched. That makes sched_ext useful for workload-specific experiments and policies, but it does not guarantee better performance: the result depends on the workload, CPU topology, fairness goals, and implementation.
What sched_ext lets you control
A sched_ext scheduler implements some or all of a set of optional callbacks. Only ops.name is mandatory. The callbacks and scx_bpf_* helpers let the BPF program choose CPUs, enqueue and dispatch tasks, and manage scheduling queues. The kernel’s interface is documented in the sched_ext kernel guide; relevant implementation material is in the kernel source, including include/linux/sched/ext.h and sched_ext core files.
First decide which tasks the scheduler should govern. With the default switching mode, sched_ext schedules tasks using SCHED_NORMAL, SCHED_BATCH, SCHED_IDLE, and SCHED_EXT while the BPF scheduler is active. Setting SCX_OPS_SWITCH_PARTIAL changes that scope: only tasks explicitly using SCHED_EXT are switched, while the fair class continues to handle normal, batch, and idle tasks. A task assigned SCHED_EXT before a scheduler is loaded is treated as SCHED_NORMAL.
Use partial switching when the intended policy should apply only to explicitly selected tasks; use the default mode only when the scheduler is meant to take over the broader set. The mode determines the scheduler’s reach, not whether its policy is fair or suitable for a particular machine. See the kernel documentation on sched_ext operation and switching.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Check kernel support and run an example
Support depends on the running kernel’s configuration, not just on the availability of sched_ext documentation or source examples. The kernel guide lists CONFIG_SCHED_CLASS_EXT and BPF-related requirements, including BPF syscall and JIT support, along with debug BTF configuration options. Check the configuration of the kernel you will actually boot and the requirements for the scheduler you intend to load.
The Linux 6.12 documentation covers the interface; that establishes documentation for 6.12, not that it was the feature’s first upstream version. The guide’s simple in-tree example can be built and run from a Linux source tree with:
make -j16 -C tools/sched_exttools/sched_ext/build/bin/scx_simple
These are the documented example commands, not a promise that every distribution kernel has sched_ext enabled or that the example is suitable for production. Check the current kernel guide and the Linux 6.12 versioned guide against your target kernel.
Rank #2
Follow a task through the scheduling cycle
The core design question is how a runnable task moves from a wakeup to a CPU. A typical path is:
- Choose a candidate CPU.
ops.select_cpu()is called when a task wakes. Its choice is a placement hint, not a binding: an invalid or disallowed CPU choice can be ignored. - Decide where the task waits. If
select_cpu()does not dispatch the task directly,ops.enqueue()can put it in a built-in dispatch queue, a custom dispatch queue (DSQ), or scheduler-managed BPF data structures. - Make work available to a CPU. A CPU checks its local DSQ, then the global DSQ. If neither has runnable work,
ops.dispatch()can populate local work. The built-in local and global DSQs are FIFO queues; custom DSQs can support FIFO or priority behavior. - Handle changes in task state. When a task held by the scheduler leaves its custody—for example, when it is dispatched to a terminal DSQ or sleeps—
ops.dequeue()is called once for that departure. Account for that lifecycle when storing task state or maintaining queue metadata.
Direct dispatch from select_cpu() can skip enqueue(). Likewise, a task placed in a custom DSQ or retained in BPF-owned structures remains in scheduler custody, so its lifecycle is not identical to a task sent straight to a terminal built-in queue. The kernel guide’s callback and DSQ descriptions are essential when deciding which callbacks your design must implement.
Choose a queue and placement policy around the goal
Start by stating the intended result in terms you can measure: for example, how you will assess latency, throughput, CPU locality, fairness, or cgroup control. Then design around the target hardware and workload rather than selecting a queue structure in isolation.
- Use built-in global and local DSQs when their FIFO behavior is sufficient and you want a straightforward dispatch path. A global queue can make work available across CPUs, while local queues connect dispatch to a particular CPU.
- Use custom DSQs when you need queue behavior such as priority ordering, or need to organize dispatch beyond the built-in FIFO queues.
- Keep tasks in BPF-managed structures when selection policy requires scheduler-owned organization before dispatch. In exchange, you must manage the relevant task lifecycle and ensure runnable work reaches CPUs.
- Make CPU placement topology-aware where needed. A policy that fits a single-socket, uniform-LLC machine may not fit a system with a more complex or NUMA-like topology. Account for locality and load distribution in the target environment.
- Specify fairness and control semantics explicitly. Decide how the policy treats competing tasks, idle CPUs, priorities, and cgroups before implementation; queue order alone does not define a complete policy.
A practical design sequence is to define the workload objective and constraints, choose a CPU-selection and locality policy, select terminal queues or scheduler-owned structures, implement lifecycle handling, and then evaluate the result under representative load and topology. This is a design method derived from the documented callback and queue model, not a kernel-mandated algorithm.
What the in-tree examples illustrate
The examples are useful starting points for understanding policy choices, but the kernel README cautions that examples primarily demonstrate features and testing rather than practical, production-ready schedulers. The project guide also explicitly says scx_qmap is a feature illustration, not production ready. Treat each as a pattern to study and test, not a drop-in recommendation.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
| Example | What it demonstrates | Fit and caveat |
|---|---|---|
scx_simple |
Minimal global FIFO or weighted virtual-time scheduling. | The project guide says it may suit single-socket systems with uniform L3 topology. It warns that global FIFO can starve inactive tasks when saturating threads are present; suitability on another topology or workload is not established. |
scx_qmap |
Weighted FIFO levels and BPF queue/storage techniques. | The project guide characterizes it as a feature illustration, not production ready. |
scx_central |
Centralized scheduling decisions, including dispatching work so other cores can run with long slices and avoid timer ticks. | The in-tree README discusses possible usefulness for VM workloads. That is an example use case, not a guarantee of benefit on a given host. |
scx_flatcg |
Hierarchical cgroup CPU control by flattening compounded weights into one scheduling layer. | Useful for examining that cgroup-control approach; the cited example material does not establish production readiness. |
scx_pair |
Sibling-core and cgroup coordination. | Study it when coordination across sibling cores is relevant; performance and suitability depend on the target system. |
scx_userland |
A minimal user-space scheduling example. | It demonstrates a user-space example path; the cited guide does not establish production readiness. |
Descriptions and cautions above come from the sched_ext example-scheduler guide, the in-tree sched_ext README, and the kernel guide. Use them to identify design patterns; validate any chosen policy against your own CPU topology, workload, fairness requirements, cgroup expectations, and scheduling overhead.
Rank #4
Implement cgroup and nice behavior deliberately
Do not assume the fair scheduler will automatically apply its usual controls to a custom sched_ext policy. The kernel communicates cgroup controls and nice changes through callbacks, but the BPF scheduler is responsible for implementing the corresponding semantics and may choose to ignore them.
If your scheduler claims to honor cpu.max, cpu.weight, cpu.idle, or nice-derived weights, implement and verify each behavior. If it does not honor one of them, make that limitation clear to operators; otherwise, users may expect controls to work when the policy does not apply them.
Plan for failure, inspection, and recovery
sched_ext can abort a BPF scheduler if the program terminates, an internal error occurs, or a runnable task stalls. On abort, tasks return to fair-class scheduling. This recovery path reduces the risk of a scheduler failure leaving the system permanently governed by a broken policy, but it does not replace testing or diagnosis.
Best Value
For operational inspection, the kernel documentation describes state files under /sys/kernel/sched_ext/, the monotonically increasing enable_seq, scheduler event counters, per-task state in /proc/self/sched, and debug-dump mechanisms including the sched_ext_dump tracepoint. Use these to investigate scheduler state, event history, task behavior, and dumps when a policy stalls or behaves unexpectedly. The kernel guide documents these interfaces.
Build for the target kernel and measure the actual policy
sched_ext has no API stability guarantee. The kernel’s ABI Instability documentation states: “The APIs provided by sched_ext to BPF schedulers programs have no stability guarantees.” It further warns that interfaces may change without warning between kernel versions. Build against and verify the documentation and source for the kernel you plan to run; do not assume a scheduler written for one release will remain compatible with another.
Evaluate a custom policy on the intended kernel, machine, and workload. Compare it with the existing scheduling behavior using measurements relevant to the stated goal, and include competing or saturating tasks if fairness and starvation matter. Check CPU locality, load distribution, overhead, and cgroup behavior where relevant. No generic performance gain follows from using sched_ext or from adopting one of the examples.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

