iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
An incident playbook helps you work out what is happening. An incident runbook tells you what to do once the cause is known. The two documents answer different questions, need different evidence and permissions, and stop at different points, so keep them separate. This guide covers both, then gives a triage sequence for failures in OpenAI’s Agents API, where the error surface is specific enough to act on.
Playbook or runbook: which one do you need?
AWS’s Well-Architected guidance states the investigation role plainly: “Playbooks are step-by-step guides used to investigate an incident.” (AWS Well-Architected Framework, OPS07-BP04.) Its security guidance adds that “Incident response playbooks provide a series of prescriptive guidance and steps to follow when a security event occurs.” (AWS Well-Architected Framework, SEC10-BP04.) A playbook therefore guides discovery and scoping toward a root cause. A runbook begins after that point and describes mitigation.
The distinction matters most at the moment an alert fires. AWS’s GuardDuty guidance frames the question a team faces after a finding as “Now what?” A playbook turns that question into a scoped investigation. A runbook should be opened only once the investigation has named a cause it can act on.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall| Question | Investigation playbook | Mitigation runbook |
|---|---|---|
| Purpose | Discover symptoms, scope impact, and find the root cause | Mitigate a cause that is already understood |
| Starting point | An alert or symptom with no confirmed cause | A confirmed cause and the matching procedure |
| Tools and permissions | Name any special tools and elevated permissions before the first step | Name the tools and permissions the mitigation needs, and confirm them before acting |
| Expected output | A root cause, or a recorded statement that the cause is still unknown | The affected resource returns to the state the runbook defines as expected |
| Escalation trigger | Define one for the case where the cause is still unknown | Not specified in the cited AWS guidance; define one in the runbook’s escalation section |
What a reusable runbook needs
AWS’s guidance on security playbooks calls for each scenario to state its goal, prerequisites, owners and escalation path, technical response steps, and expected outcomes. Write each runbook scenario with those parts, in this order:
#1 Best Overall
- Overview and goal. Name the scenario, the alert or symptom that triggers it, and what done looks like.
- Prerequisites. List the logs, detection mechanisms, tools, and the alert you expect to see. A responder missing any of these should know before starting.
- Contacts, responsibilities, and escalation. Name who owns each step and who to call when a step does not produce the expected result.
- Response steps. For each step, state what to inspect, the query or code to run, the result you expect, and the next decision that result triggers.
- Expected outcomes. Describe the state that confirms the scenario is closed.
Steps must be operational rather than descriptive. “Check the logs” is not a step. A usable step names the log source, the filter or query, and what a healthy result looks like. Where the cited guidance gives no query for your platform, write one yourself and test it during validation.
AWS’s security framework groups response actions into five phases: detect, analyze, contain, eradicate, and recover. Use them as the checklist of what a scenario must cover, not as a replacement for the scenario’s own commands and authorization limits.
| Phase | What the runbook must specify |
|---|---|
| Detect | The signal that starts the runbook and the detection mechanism that produced it |
| Analyze | How to confirm scope: affected resources, sessions, or environments, and the time window |
| Contain | The action that stops further impact, and who authorizes it |
| Eradicate | How the cause is removed, and what proves it is gone |
| Recover | How affected resources return to service, and how that is verified |
Investigating when the cause is still unknown
Operational troubleshooting should run outside in: start with what users and systems observe, then move toward the component that fails. Use this sequence when the cause is not yet known.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems- Discover the symptoms. Record what is observed, when it started, and which alert, if any, fired first.
- Scope the impact. List the affected sessions, environments, accounts, or workloads, and check whether the count is growing.
- Gather evidence. Collect logs, status values, and error objects for the affected items. Note the tool and permission each query needs before you start.
- Identify the root cause. Match the evidence to a failure layer (see the triage sections below), or record that the cause is still unknown.
- Hand off to mitigation. Once the cause is confirmed, open the matching mitigation runbook. If it is not confirmed, continue under the escalation route.
Treat an authorization error as its own branch. AWS-specific IAM troubleshooting material quotes the message “I am not authorized to perform an action.” When a responder sees that text, the question is about access, and the next step is to involve the owner of the relevant policy rather than repeat the same action.
Rank #2
Two communication rules belong in every investigation. First, send status updates on an interval agreed in advance, covering what is known, what is not, and who is working it. Second, set a time limit for diagnosis in the runbook. When that limit passes without a confirmed cause, the escalation route activates. The cited guidance does not set a duration, so choose one that fits your service’s impact tolerance.
What failed: the request, the turn, the session, or the environment?
In the OpenAI Agents API, failures are reported at four layers. The layer that reports the error tells you which object to retrieve next.
| Layer | Where the failure appears | What to inspect first |
|---|---|---|
| Request | HTTP status and the response error object | The HTTP status and error object returned by the failing API call |
| Turn | Turn status and error | Retrieve the turn and read its status and error |
| Session | Session status and error | Retrieve the session and read its status and error |
| Environment | Environment error event | Read the environment error details, then follow the sandbox troubleshooting guidance |
Classify the layer before you act. The repair for a request error is not the repair for a sandbox setup error, and applying the wrong one wastes time and muddies the record.
Should I retry, repair, or recreate the session?
OpenAI’s error guidance draws the line that matters most here: “A failed turn doesn’t always mean the session has failed.” Check session status before deciding anything else.
- Session still usable: determine whether it can continue from where it stopped.
- Session failed: fix the underlying issue, then create a new session and supply the inputs it needs.
Known error classes call for specific responses. Handle each one as follows.
Connection failure or timeout
Inspect executor startup and network access. Confirm the executor is running and that the network path it needs is open before you change anything else.
sandbox_error
Check the setup commands, the packages being installed, the input files, and the environment error reported with the failure. Check each one against that reported error, since it points to the setup or environment layer without naming the single faulty item.
Recommended Free Tools
Incompatible executor version
The guidance calls for an upgrade before you create a new session, so the upgrade is the repair step. Creating a session first leaves the mismatch in place.
Rank #4
idle_timeout
The session has timed out for inactivity and must be replaced. Create a new session and supply the inputs again.
Blocked sandbox request
Inspect the network settings and the hosts the request reaches, including hosts reached through redirects. A redirect can carry a request to a host you did not list, so check the final destination, not only the first address.
Expired environment during file operations
Before live file operations, confirm the sandbox is connected. If the environment has expired, create a new session and resubmit the inputs.
Hosted environments and self-hosted sandboxes
OpenAI’s hosted sandbox guide says OpenAI provisions and connects the environment. A self-hosted sandbox is for cases that need a custom image, compute, or a private network. Choose between them by deciding which of those you must control and who will operate it.
| Factor | Hosted environment | Self-hosted sandbox |
|---|---|---|
| Who provisions and connects the environment | OpenAI | Your team |
| Fits cases that | Do not need a custom image, custom compute, or a private network | Need a custom image, custom compute, or a private network |
| Control over image and network | Set by OpenAI’s provisioning; the cited guide does not describe customizing it | Controlled by your team |
| Operational ownership | OpenAI, for provisioning and connection | Your team, for the image, compute, and network it uses |
| Setup and connectivity failure surface | Environment error details and the sandbox troubleshooting guidance | Not stated in the cited guidance as a separate failure surface |
What to record and what to preserve before escalating
The cited vendor guidance says where to inspect and how to recover, but it does not prescribe a record format. The fields below are a practice this guide recommends, not a vendor requirement.
- The observable symptom, in one sentence.
- The event or error identifier: the HTTP status and error object, the turn or session error, or the environment error event.
- The affected session or environment.
- The change made, with the time it was made.
- The expected outcome, and what actually happened.
Preserve request and session identifiers when you escalate. If a status or file-list request keeps returning server errors, OpenAI’s guide recommends keeping the request ID. Do not retry blindly. Repeated attempts without a changed cause add noise to the record the next responder will rely on.
Validating the runbook before a real incident
AWS recommends validating response arrangements before an actual incident. Its Incident Detection and Response guidance describes a scheduled GameDay as an end-to-end simulation in which participants observe how the runbook unfolds and refine its instructions. Scheduling requirements are service-specific, so check AWS’s current GameDay service page before planning one.
Free tools Windows power users keep installed
One-click scans. No signup required.
Review the runbook when any of these change:
- The workload it covers
- The alerts that trigger it
- The permissions a responder needs
- The tools it names
- The escalation contacts
This review rule is a practical recommendation drawn from AWS’s emphasis on prerequisites, response contacts, and workload-specific runbooks. It is not a quoted requirement.
Quick Recap
What these sources do not establish
- A vendor-neutral error taxonomy for CLI agents. The error classes above come from OpenAI’s Agents API and do not transfer automatically to other vendors’ tools.
- A universal diagnostic command. Commands for your platform must come from that platform’s own documentation.
- Incident-rate, recovery-time, or error-reduction figures. The AWS GameDay material describes scheduling and process, not performance. Measure your own runbook results before claiming improvement.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

