iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Erlang/OTP can give an agentic application well-defined process boundaries, concurrency, and supervision-based recovery. It does not provide the agent’s reasoning, connect a model, decide which tools are safe, or make external work durable. Treat OTP as runtime infrastructure for the application components that perform model calls and tool actions—not as an agent framework.
What OTP contributes to an agentic system
OTP is a set of design principles and components for structuring Erlang applications. Its building blocks include processes, modules, directories, and supervisors arranged in hierarchical supervision trees. The Erlang/OTP documentation describes the supervision tree as “a hierarchical arrangement of code into supervisors and workers, making it possible to design and program fault-tolerant software.” Erlang/OTP Design Principles
For an AI application, a process can own a session or perform a unit of work, while supervisors manage worker lifecycles and configured restart behavior. That helps structure concurrent work and contain some failures. The model request, prompt construction, tool policy, and business logic remain application responsibilities.
What OTP does not do for you
A process restart is not the same as recovering an agent task. OTP does not automatically supply planning, prompt management, tool authorization, durable workflow state, idempotency, or recovery from external side effects. Those concerns need explicit application design.
#1 Best Overall
- Planning and model integration: decide how the application calls a model, interprets results, and determines the next step.
- Tool authorization: enforce which tools a task may invoke and validate their inputs. A tool executor process is not, by itself, a security policy.
- Durable state: persist any task or session information that must survive a process or node restart.
- External effects: design timeouts, retry rules, idempotency, and compensation for actions such as payments or API writes. A supervisor cannot undo an already-completed external action.
A practical OTP-shaped architecture
The following is one possible design, not an OTP prescription. Choose process boundaries based on ownership and failure behavior, then define the messages and state each component is responsible for.
| Component | Possible responsibility | Design question |
|---|---|---|
| Session coordinator | Track orchestration for one session or task and decide which deterministic step runs next. | Which state must be persisted so a restart can resume safely? |
| Model-provider adapter | Make a model request and return a normalized result to the coordinator. | How are timeouts, provider errors, and repeated requests handled? |
| Tool executor | Validate and execute an allowed tool action. | Where are authorization checks, idempotency, and side-effect handling enforced? |
| Background worker | Perform a longer-running or deferred unit of work. | Should its failure restart only this worker or a related group? |
Keep deterministic orchestration distinct from model-driven decisions. For example, the coordinator may enforce that a tool result is validated and recorded before another model call, even when the model proposes the sequence of actions. Give messages clear contracts, and make state ownership explicit so a restarted worker does not depend on memory that no longer exists.
Rank #2
Choose supervision behavior around failure boundaries
Supervisors start, stop, and monitor child processes, applying the restart strategy configured for their children. The OTP supervisor manual describes three strategies; their operational consequences should be checked against the OTP version used by the application. Erlang supervisor manual
| Strategy | Effect when a child fails | When to consider it |
|---|---|---|
one_for_one |
Restart only the failed child. | Workers can recover independently and do not rely on a shared lifecycle with siblings. |
one_for_all |
Restart the group of children. | The children are interdependent enough that restarting only one could leave the group inconsistent. |
rest_for_one |
Restart the failed child and children started after it. | Later children depend on earlier-started children, so a failure should rebuild that downstream portion. |
These choices describe process recovery, not workflow recovery. If a model call succeeds and a tool action times out before its result is recorded, blindly restarting a worker could repeat the action. Persist workflow progress and give external operations safe retry semantics where needed.
Understand links, exits, and supervision
Erlang processes can be linked, and exit signals can propagate termination behavior between linked processes. This is a mechanism for coordinating process failure, not a complete recovery design. Erlang Reference Manual: Processes and links
Use links and exit handling deliberately within the lifecycle design. A supervision tree provides a hierarchical place to specify how child failures affect restarts; it does not make every linked process failure recoverable or ensure that application data and external effects are consistent.
Rank #4
Choose how components coordinate
Within one runtime, processes can communicate through messages. Distributed Erlang adds node connections and monitoring, remote process spawning, and message exchange. Its documentation describes it primarily as a means of Erlang-to-Erlang communication—not as a general-purpose public-network protocol. Erlang/OTP Distributed Applications
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallUse distribution only after deciding how nodes are connected and how the deployment handles authorization, encryption, network exposure, and operational failure. The existence of node communication features does not establish that a particular deployment is secure, nor does it provide cross-language interoperability. External queues or services may be a better coordination boundary when components or teams use different runtimes.
Questions to settle before implementation
- Failure boundary: which worker, or group of workers, should restart after each kind of failure?
- State durability: what must survive a worker restart, node restart, or deployment?
- External effects: how will the system handle retries, idempotency, timeouts, and compensation?
- Coordination: should components use local messages, distributed Erlang nodes, or an external queue or service?
- Operations: how will an operator inspect failures and trace a task across multiple steps?
- Decision boundary: which steps are deterministic application orchestration, and which are model-driven?
Answering these questions is more important than assigning every conceptual “agent” its own process. Use a process boundary where it clarifies ownership, concurrency, or recovery—not simply because the system includes an AI model.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

