Achieve real-time dynamic load balancing by measuring each core over a bounded interval, then moving only eligible tasks when the expected scheduling benefit exceeds migration and synchronization costs. Choose SMP when a single kernel can safely manage shared state across the cores; choose AMP when isolation and explicit inter-core communication matter more. In either design, preserve affinity for work that depends on predictable interrupt handling, cache locality, or safety constraints, and validate the policy against deadlines on the target hardware.
Choose SMP or AMP before designing the balancing policy
SMP and AMP describe different ownership models for multicore execution; they are not interchangeable load-balancing algorithms. The choice affects what the scheduler can move, how cores coordinate, and where synchronization responsibility lies.
| Architecture | Kernel and scheduling model | Implication for balancing |
|---|---|---|
| SMP | One kernel instance schedules tasks across multiple cores that share memory. FreeRTOS documents this model for identical processor cores. | A shared scheduler can place work across cores, but tasks may execute simultaneously. Revisit assumptions about priority ordering, shared data, and mutual exclusion. |
| AMP | Each core runs an independent RTOS instance. FreeRTOS describes inter-core communication using shared memory and stream or message buffers. | There is no single scheduler balancing a common task pool across all cores. Partition work deliberately and coordinate transfers through explicit inter-core communication. |
| Zephyr SMP | By default, any processor can run any Zephyr thread. CPU masks can restrict a thread to allowed CPUs; pin-only mode uses an independent run queue per CPU. | Use CPU masks to express eligibility. Pin-only mode favors per-core queue independence, while less restrictive placement gives the scheduler more room to move work. |
As Zephyr’s SMP documentation notes, real-time applications may deliberately partition work across physical CPUs rather than rely only on the scheduler to choose a CPU. That is especially relevant when a workload includes hard-deadline, interrupt-coupled, or safety-critical tasks that should not be freely migrated.
When to favor SMP
Choose SMP when cores can safely share memory and a single kernel’s scheduling authority is useful. The convenience of a common task pool does not remove the need for synchronization: two cores may run tasks concurrently, and an interrupt service routine can execute while other work is active. Treat shared-state protection and memory-ordering behavior as part of the architecture.
#1 Best Overall
When to favor AMP
Choose AMP when cores need distinct responsibilities, separate RTOS instances, or stronger partitioning. Define which core owns each task or resource, and specify how data and events cross core boundaries. Shared memory alone is not a complete communication design; synchronization and message ownership still need to be explicit.
Build a bounded balancing loop
A real-time balancing policy should have bounded measurement and action costs. It should not react to a single utilization sample or migrate work simply because one core appears busier at an instant.
Rank #2
- Classify tasks. Mark each task as hard-deadline, soft-deadline, interrupt- or driver-coupled, cache-sensitive, or background work. Record its priority or deadline, execution-time estimate, and allowed CPU mask. Pin critical or interrupt-coupled work unless analysis shows migration is safe.
- Measure over a fixed window. Sample each core’s idle fraction or scheduler runtime counters over a defined interval. Zephyr’s CPU-load module supports per-CPU scheduler runtime statistics and idle-hook measurement;
cpu_load_get_cpu()reports a value from 0 to 1000 per mille in the current Zephyr documentation. Interpret it as a scaled load reading, not proof that a task will meet its deadline. - Detect meaningful imbalance. Consider utilization difference together with ready-queue depth, task deadline slack, execution-time estimates, and recent migration cost. An instantaneous percentage alone does not show whether queued work is urgent or whether moving it will help.
- Choose an eligible task and destination. Select a task whose CPU mask permits the destination and whose remaining slack can accommodate the expected migration, synchronization, and cache costs. Avoid moving critical or interrupt-coupled tasks without schedulability analysis.
- Bound migrations. Cap the number of migrations in each scheduling window. Include cache warm-up, lock hold time, interrupt masking, and inter-processor interrupt (IPI) latency in the cost estimate.
- Re-evaluate and stop. Stop once imbalance is below a hysteresis threshold, or when the predicted response-time benefit no longer exceeds the migration overhead. Hysteresis prevents small fluctuations from causing repeated moves between cores.
Use slack and overhead, not utilization alone
A useful decision is whether a candidate task’s remaining deadline slack is greater than the estimated costs introduced by moving it, with enough margin for uncertainty. Those costs can include migration execution, synchronization, cache refill, and time until the destination CPU responds. This is a decision framework, not a universal numeric threshold: derive bounds from the application’s timing analysis and measurements on the target.
Match scheduler policy to the timing problem
Different scheduling choices change deadline behavior, queue costs, and the amount of placement control. A balancing policy must work with the RTOS scheduler rather than assume that all systems schedule tasks the same way.
Recommended Free Tools
Rank #3
| RTOS or policy detail | Documented behavior | Design consideration |
|---|---|---|
| FreeRTOS default scheduling | Fixed-priority preemptive scheduling with round-robin time slicing among equal-priority tasks. | In SMP, do not assume a higher-priority task on one core prevents a lower-priority task from running on another. Shared resources still need correct synchronization. |
| RTEMS EDF-based SMP scheduler | RTEMS documents an earliest-deadline-first (EDF) SMP scheduler and affinity options. | Consider it when explicit deadlines drive dispatch, while still accounting for affinity, shared-resource blocking, and migration effects. |
| Zephyr ready-queue backends | Zephyr offers multiple ready-queue backends. CPU-mask filtering may require broader queue scans; documented complexity includes O(N) scans for simple/scalable backends and O(P·N) worst case for multi-queue filtering. | Account for queue-search cost when using CPU masks. Pin-only mode keeps independent per-CPU queues, trading flexible placement for queue independence. |
A globally shared queue can make placement straightforward but may increase contention. Per-CPU queues can scale more effectively, but need a balancing or work-stealing mechanism to address uneven queues. Compare candidate designs using deadline predictability, migration overhead, affinity flexibility, lock contention, cache locality, interrupt interference, and scheduler run-queue cost.
Measure execution, migration, and deadline effects
CPU utilization is only one part of the evidence needed to judge a balancing policy. Measure both the load it is intended to redistribute and the timing or synchronization costs it may add.
Rank #4
- Per-core busy and idle time: Use bounded-window idle measurements or scheduler runtime counters.
- Task execution and response time: Record execution time and release-to-completion latency for relevant tasks.
- Queue and urgency state: Track ready-queue depth and deadline slack so an apparent load imbalance can be interpreted in context.
- Balancing overhead: Count migrations and measure their duration, including the effects of cache warm-up where observable.
- Interference: Measure interrupt latency, time spent in critical sections, mutex contention, and priority-inversion events.
- Outcome under stress: Count deadline misses during controlled overload, not only during normal operation.
Zephyr runtime statistics can report execution cycles per thread and aggregate usage that includes the idle thread, supporting utilization calculations. FreeRTOS’s Kernel Book explains that the run-time statistics clock is supplied by application code rather than by the RTOS tick, and documents the configuration requirements for vTaskGetRunTimeStatistics(). Confirm the counter source, resolution, wraparound behavior, and instrumentation overhead for the specific port before relying on measurements.
Use trace timelines to inspect when tasks run, preempt, and migrate. The official FreeRTOS site identifies Percepio Tracealyzer as a tracing tool for FreeRTOS applications. Combine traces with response-time and interrupt-latency measurements; a visually balanced timeline alone does not establish that deadlines are safe.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsBest Value
Protect shared state and validate core wake-ups
Multicore execution changes concurrency assumptions even if task priorities and source code appear unchanged. FreeRTOS warns that an SMP system can run a lower-priority task on one core while a higher-priority task runs on another, and that ISRs can execute concurrently. Priority ordering is not mutual exclusion.
- Protect shared state with appropriate mutexes, atomics, or message passing; keep critical sections bounded.
- Verify that the target SoC’s cache-coherency and memory-ordering behavior matches the assumptions made by the RTOS port and application.
- Include IPI delivery, CPU wake-up, and deferred or dynamically onlined CPU behavior in testing. Zephyr documents an edge case where an idle CPU may not wake to handle newly runnable load in some configurations.
- Run overload tests on the actual multicore target and record deadline misses, interrupt latency, migration counts, and lock contention alongside per-core load.
Heterogeneous platforms can also combine different operating environments across cores. NXP’s Real-time Edge Software User Guide documents configurations using Linux alongside FreeRTOS and/or Zephyr cores. Before selecting a board for AMP or SMP work, check its core topology, cache coherency, interrupt routing, supported RTOS port, and current toolchain.
Set thresholds from the application, not a generic percentage
There is no universal load threshold in the cited RTOS documentation that makes migration safe. A threshold that works for one workload, core topology, or interrupt pattern may be unsuitable for another. Set measurement windows, hysteresis, and migration limits from target-specific schedulability analysis, then validate them with controlled tests that include overload. Average utilization can reveal imbalance, but it cannot by itself establish worst-case deadline safety.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

