Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An incident playbook helps you work out what is happening. An incident runbook tells you what to do once the cause is known. The two documents answer different questions, need different evidence and permissions, and stop at different points, so keep them separate. This guide covers both, then gives a triage sequence for failures in OpenAI’s Agents API, where the error surface is specific enough to act on.

Playbook or runbook: which one do you need?

AWS’s Well-Architected guidance states the investigation role plainly: “Playbooks are step-by-step guides used to investigate an incident.” (AWS Well-Architected Framework, OPS07-BP04.) Its security guidance adds that “Incident response playbooks provide a series of prescriptive guidance and steps to follow when a security event occurs.” (AWS Well-Architected Framework, SEC10-BP04.) A playbook therefore guides discovery and scoping toward a root cause. A runbook begins after that point and describes mitigation.

The distinction matters most at the moment an alert fires. AWS’s GuardDuty guidance frames the question a team faces after a finding as “Now what?” A playbook turns that question into a scoped investigation. A runbook should be opened only once the investigation has named a cause it can act on.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Question Investigation playbook Mitigation runbook
Purpose Discover symptoms, scope impact, and find the root cause Mitigate a cause that is already understood
Starting point An alert or symptom with no confirmed cause A confirmed cause and the matching procedure
Tools and permissions Name any special tools and elevated permissions before the first step Name the tools and permissions the mitigation needs, and confirm them before acting
Expected output A root cause, or a recorded statement that the cause is still unknown The affected resource returns to the state the runbook defines as expected
Escalation trigger Define one for the case where the cause is still unknown Not specified in the cited AWS guidance; define one in the runbook’s escalation section

What a reusable runbook needs

AWS’s guidance on security playbooks calls for each scenario to state its goal, prerequisites, owners and escalation path, technical response steps, and expected outcomes. Write each runbook scenario with those parts, in this order:

  1. Overview and goal. Name the scenario, the alert or symptom that triggers it, and what done looks like.
  2. Prerequisites. List the logs, detection mechanisms, tools, and the alert you expect to see. A responder missing any of these should know before starting.
  3. Contacts, responsibilities, and escalation. Name who owns each step and who to call when a step does not produce the expected result.
  4. Response steps. For each step, state what to inspect, the query or code to run, the result you expect, and the next decision that result triggers.
  5. Expected outcomes. Describe the state that confirms the scenario is closed.

Steps must be operational rather than descriptive. “Check the logs” is not a step. A usable step names the log source, the filter or query, and what a healthy result looks like. Where the cited guidance gives no query for your platform, write one yourself and test it during validation.

AWS’s security framework groups response actions into five phases: detect, analyze, contain, eradicate, and recover. Use them as the checklist of what a scenario must cover, not as a replacement for the scenario’s own commands and authorization limits.

Phase What the runbook must specify
Detect The signal that starts the runbook and the detection mechanism that produced it
Analyze How to confirm scope: affected resources, sessions, or environments, and the time window
Contain The action that stops further impact, and who authorizes it
Eradicate How the cause is removed, and what proves it is gone
Recover How affected resources return to service, and how that is verified

Investigating when the cause is still unknown

Operational troubleshooting should run outside in: start with what users and systems observe, then move toward the component that fails. Use this sequence when the cause is not yet known.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Discover the symptoms. Record what is observed, when it started, and which alert, if any, fired first.
  2. Scope the impact. List the affected sessions, environments, accounts, or workloads, and check whether the count is growing.
  3. Gather evidence. Collect logs, status values, and error objects for the affected items. Note the tool and permission each query needs before you start.
  4. Identify the root cause. Match the evidence to a failure layer (see the triage sections below), or record that the cause is still unknown.
  5. Hand off to mitigation. Once the cause is confirmed, open the matching mitigation runbook. If it is not confirmed, continue under the escalation route.

Treat an authorization error as its own branch. AWS-specific IAM troubleshooting material quotes the message “I am not authorized to perform an action.” When a responder sees that text, the question is about access, and the next step is to involve the owner of the relevant policy rather than repeat the same action.

Two communication rules belong in every investigation. First, send status updates on an interval agreed in advance, covering what is known, what is not, and who is working it. Second, set a time limit for diagnosis in the runbook. When that limit passes without a confirmed cause, the escalation route activates. The cited guidance does not set a duration, so choose one that fits your service’s impact tolerance.

What failed: the request, the turn, the session, or the environment?

In the OpenAI Agents API, failures are reported at four layers. The layer that reports the error tells you which object to retrieve next.

Layer Where the failure appears What to inspect first
Request HTTP status and the response error object The HTTP status and error object returned by the failing API call
Turn Turn status and error Retrieve the turn and read its status and error
Session Session status and error Retrieve the session and read its status and error
Environment Environment error event Read the environment error details, then follow the sandbox troubleshooting guidance

Classify the layer before you act. The repair for a request error is not the repair for a sandbox setup error, and applying the wrong one wastes time and muddies the record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I retry, repair, or recreate the session?

OpenAI’s error guidance draws the line that matters most here: “A failed turn doesn’t always mean the session has failed.” Check session status before deciding anything else.

  • Session still usable: determine whether it can continue from where it stopped.
  • Session failed: fix the underlying issue, then create a new session and supply the inputs it needs.

Known error classes call for specific responses. Handle each one as follows.

Connection failure or timeout

Inspect executor startup and network access. Confirm the executor is running and that the network path it needs is open before you change anything else.

sandbox_error

Check the setup commands, the packages being installed, the input files, and the environment error reported with the failure. Check each one against that reported error, since it points to the setup or environment layer without naming the single faulty item.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Incompatible executor version

The guidance calls for an upgrade before you create a new session, so the upgrade is the repair step. Creating a session first leaves the mismatch in place.

idle_timeout

The session has timed out for inactivity and must be replaced. Create a new session and supply the inputs again.

Blocked sandbox request

Inspect the network settings and the hosts the request reaches, including hosts reached through redirects. A redirect can carry a request to a host you did not list, so check the final destination, not only the first address.

Expired environment during file operations

Before live file operations, confirm the sandbox is connected. If the environment has expired, create a new session and resubmit the inputs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Hosted environments and self-hosted sandboxes

OpenAI’s hosted sandbox guide says OpenAI provisions and connects the environment. A self-hosted sandbox is for cases that need a custom image, compute, or a private network. Choose between them by deciding which of those you must control and who will operate it.

Factor Hosted environment Self-hosted sandbox
Who provisions and connects the environment OpenAI Your team
Fits cases that Do not need a custom image, custom compute, or a private network Need a custom image, custom compute, or a private network
Control over image and network Set by OpenAI’s provisioning; the cited guide does not describe customizing it Controlled by your team
Operational ownership OpenAI, for provisioning and connection Your team, for the image, compute, and network it uses
Setup and connectivity failure surface Environment error details and the sandbox troubleshooting guidance Not stated in the cited guidance as a separate failure surface

What to record and what to preserve before escalating

The cited vendor guidance says where to inspect and how to recover, but it does not prescribe a record format. The fields below are a practice this guide recommends, not a vendor requirement.

  • The observable symptom, in one sentence.
  • The event or error identifier: the HTTP status and error object, the turn or session error, or the environment error event.
  • The affected session or environment.
  • The change made, with the time it was made.
  • The expected outcome, and what actually happened.

Preserve request and session identifiers when you escalate. If a status or file-list request keeps returning server errors, OpenAI’s guide recommends keeping the request ID. Do not retry blindly. Repeated attempts without a changed cause add noise to the record the next responder will rely on.

Validating the runbook before a real incident

AWS recommends validating response arrangements before an actual incident. Its Incident Detection and Response guidance describes a scheduled GameDay as an end-to-end simulation in which participants observe how the runbook unfolds and refine its instructions. Scheduling requirements are service-specific, so check AWS’s current GameDay service page before planning one.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Review the runbook when any of these change:

  • The workload it covers
  • The alerts that trigger it
  • The permissions a responder needs
  • The tools it names
  • The escalation contacts

This review rule is a practical recommendation drawn from AWS’s emphasis on prerequisites, response contacts, and workload-specific runbooks. It is not a quoted requirement.

What these sources do not establish

  • A vendor-neutral error taxonomy for CLI agents. The error classes above come from OpenAI’s Agents API and do not transfer automatically to other vendors’ tools.
  • A universal diagnostic command. Commands for your platform must come from that platform’s own documentation.
  • Incident-rate, recovery-time, or error-reduction figures. The AWS GameDay material describes scheduling and process, not performance. Measure your own runbook results before claiming improvement.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.