What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build fault tolerance in Elixir by putting processes under supervisors, choosing a restart strategy that matches their dependencies, and defining what should happen when a child repeatedly fails. Supervisors can restore processes, but they cannot recover lost in-memory state or make interrupted external work safe to repeat; those guarantees need separate design and testing.

How OTP supervision supports fault tolerance

An OTP supervisor starts, monitors, and stops child processes, then applies configured restart policies when they terminate. A supervision tree organizes this recovery into levels: a top-level supervisor owns application processes or smaller supervisors, and each subtree can form its own recovery boundary. Erlang’s supervisor documentation describes the basic aim as keeping child processes alive by restarting them when necessary.

That is process recovery, not a blanket guarantee that the application remains available or correct. A restarted GenServer starts a new process; anything held only in its old process state is gone unless the application can reconstruct it. Work interrupted by a crash may also need replay, deduplication, or reconciliation.

Design the supervision tree around dependencies

Start by listing the long-lived processes your application needs and identifying which processes rely on which others. Put independent workers under the same supervisor when their failures can be handled separately. Group tightly coupled processes under a nested supervisor when they need a shared recovery boundary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Supervisors start children in the order listed and stop them in reverse order. If a later child depends on an earlier one, preserve that order and choose a strategy that reflects the dependency. A flat :one_for_one tree is a reasonable starting point for independent children, not a universal production default. See the official OTP supervisor design principles.

Choose a restart strategy that fits the failure

Strategy What restarts Use it when
:one_for_one Only the failed child. Siblings are independent and can continue running.
:one_for_all The entire child group. The children need to recover together to return to a consistent operating state.
:rest_for_one The failed child and every child started after it. Later children depend on the failed child or on earlier children in the start order.

These strategies define recovery scope; they do not define the dependency graph for you. With :rest_for_one, for example, the position of each child in the list determines which later children restart. Keep the reason for that ordering clear in the application design.

Define child specs and restart behavior

A child specification tells a supervisor how to identify and start a child, and can also define its restart and shutdown behavior. The Elixir Supervisor API documentation describes the required :id and :start information and options including :restart, :shutdown, and :type. Modules that implement the relevant behavior commonly provide child_spec/1; when starting multiple instances of one module, assign each a distinct ID.

Choose the child’s restart type according to its intended lifecycle:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • :permanent: restart whenever the child terminates.
  • :transient: restart after abnormal termination, but not after a normal exit or shutdown.
  • :temporary: do not restart after termination.

A long-lived service that must remain available may need a permanent child policy. A task that finishes normally after returning a result usually should not be restarted simply because it completed. Confirm valid options and exact behavior against the Elixir and OTP versions used in your deployment.

Set restart intensity and plan for escalation

Restart intensity bounds how many restarts a supervisor tolerates within a time period. In the documented Elixir API, the relevant options are :max_restarts and :max_seconds; the Erlang manual describes the limit as a maximum number of restarts within a period. If the limit is exceeded, the supervisor terminates its children and itself. A parent supervisor can then handle that failed subtree according to its own policy.

This limit prevents an endless local crash loop; it does not guarantee uninterrupted service. Choose values based on how long children take to start, how the application should tolerate temporary dependency failures, and the cost of repeated initialization. There is no universal threshold. Check the version-specific Elixir API options and Erlang supervisor behavior for the release you deploy.

Supervise children created at runtime

Use DynamicSupervisor for a changing child set

Use DynamicSupervisor when the application creates and terminates workers at runtime instead of declaring every child in a fixed list. Provide an appropriate child specification and restart policy, and decide how the application will handle duplicate work and resource limits. The Elixir dynamic supervision guide covers starting processes inside supervisors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Task.Supervisor for supervised background work

A Task.Supervisor can own background tasks that should be supervised. Its start_child API starts a task as a child linked to the supervisor, rather than to the caller; this is useful for side-effecting work when the caller does not need a result. The documented default restart policy is temporary. Changing it to restart tasks can repeat side effects, so do so only if the work is designed to tolerate retries. See the Task.Supervisor API documentation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Test process recovery and application correctness separately

A focused recovery test can deliberately terminate a supervised worker and assert that a new process takes its place, as in the Elixir supervision guide. Then test the application guarantees that process replacement alone cannot establish:

  • Whether state held only in the crashed process can be reconstructed.
  • Whether messages or in-progress work can be lost during termination.
  • Whether retrying an external write can cause duplicate effects.
  • Whether a dependency outage can trigger repeated restarts.
  • Whether startup failures can exceed the configured restart intensity and take down a subtree.

Test observable outcomes, not only the existence of a replacement process. A supervisor can restart a child according to policy; application code and tests must establish whether the service’s data and external behavior remain correct.

Check the deployed Elixir and OTP versions

Elixir and OTP APIs evolve. The Elixir API reference cited here is for the main-branch v1.21.0-dev documentation, not a guarantee that every option or default matches an older release. Before using examples or relying on defaults, verify the documentation for the versions pinned by your application and deployed in production.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.