What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Reliable production agents need more than a retry loop. Classify failures before recovery, repeat only safe operations within a deadline and attempt budget, switch to a compatible fallback when a dependency remains unavailable, and preserve validated progress between stages. Then trace and account for the entire run—including retries, tools, and delegated agents—so reliability work does not hide its cost.

Start by identifying what failed

An agent run can involve several model calls, tools, asynchronous turns, and external services. A failed request does not necessarily mean that nothing happened: a turn may have completed an action before its final response failed. Recovery should begin with the error and the state of the work, not with an automatic replay.

Classify the failure before choosing a response

Failure class Examples Appropriate response
Request or configuration error Invalid or oversized input, bad credentials or permissions, unavailable model or resource, incorrect configuration Correct the underlying request or setup. Retrying unchanged is unlikely to help.
Transient service failure Throttling, temporary capacity shortage, timeout, temporary service error Retry only if the operation is safe to repeat, honoring any server-provided Retry-After value and staying within the run’s limits.
Persistent dependency or capability failure A tool, model, or server remains unavailable, or cannot perform a required operation Use a compatible fallback, degrade the feature, return a valid cached result where appropriate, defer the work, or send it for human review.
Uncertain completion or side effect A timeout or failed turn after a payment, booking, data write, or other external action may have begun Read the turn state and action results, inspect the downstream system, and determine whether the action completed before attempting it again.

Capture the HTTP status and structured error code and message when available. Recovery code should also tolerate unknown error codes and missing optional fields rather than failing while trying to handle the original failure. For asynchronous or multi-turn systems, inspect the relevant session or turn events and any saved outputs. OpenAI’s Agents API guidance says to check completed actions before asking a failed turn to repeat work, and to stop automatic retries if the error changes or the retry limit is reached.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check whether repeating the operation is safe

Before retrying a tool call, ask whether running it twice could cause a duplicate or irreversible effect. A read-only lookup is usually easier to repeat safely than an instruction to send a message, create a record, charge an account, or change a reservation. Where a downstream service supports idempotency keys, use them to make duplicate submissions recognizable. For other operations, maintain an action ledger or query the downstream system for the result before replaying. A model-level retry limit alone cannot protect against duplicated external side effects.

Set retry limits around a deadline

Retries are a recovery mechanism, not a substitute for a timeout, a capacity plan, or a fallback. Set a maximum number of attempts and a total deadline for each operation and the complete run. The retry schedule must fit inside the remaining latency budget; if it does not, stop retrying and follow the fallback or escalation path.

Use server guidance, then bounded backoff

  1. Honor Retry-After. When a service supplies this header, use its instruction rather than immediately sending another request.
  2. Otherwise, back off exponentially with random jitter. Increase the wait after successive transient failures and add randomness so many clients do not retry at the same moment.
  3. Cap each delay and the total retry window. No wait should consume the remaining time needed to return a useful result or take an alternative action.
  4. Stop when the error changes or the budget is spent. A new error class may indicate a permanent request problem rather than a transient outage.

AWS’s retry guidance illustrates six total attempts—one initial request and up to five retries—as an example, not a universal setting. SDK configuration can count attempts differently: the cited AWS guide says botocore’s total_max_attempts includes the initial attempt, while OpenAI and Anthropic SDK max_retries settings count retries only. Check the selected SDK’s current semantics before translating a policy into configuration.

Budget timeouts and concurrency together

Configure connection and read timeouts separately where the client allows it. A timeout that is shorter than a valid long-running inference can cause the caller to abandon a request that is still working, creating duplicate work if it is retried. On the other hand, a timeout longer than the user-facing deadline can leave the workflow stalled. Set these values from the operation’s real latency requirements and the end-to-end deadline.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For continuing 503 or 529 capacity errors, increasing retries can intensify pressure on an already constrained service. Stabilize the request rate, bound concurrent work, and use rate limits or queues to smooth bursts. Defer or shed low-priority work when necessary. Provisioned or cross-region capacity may help only where it is applicable and supported by the service and deployment.

Control agent-specific loops

Agent workflows need limits beyond the retry policy for an individual request. Set a maximum number of turns or recursion depth, cap completion tokens, and constrain concurrent runs. A call may grow as later prompts include earlier outputs and context; the number of calls and tokens used can therefore vary from run to run. AWS’s capacity-planning guidance relates thread rate, invocation rate, input length, completion limits, and recursion limits to requests-per-minute and tokens-per-minute estimates. Those relationships help plan capacity, but they are not universal performance benchmarks; the cited example assumes no prompt caching.

Use fallbacks when retrying is no longer useful

A fallback is the next safe way to deliver value when the primary dependency continues to fail or cannot meet the need. It should not silently change the meaning of the result. Keep response schemas and tool contracts compatible across routes, and make any degraded behavior visible to the calling application or user.

Choose a fallback that fits the failure

  • Compatible alternate model or tool: Route to another capability only if it can honor the same input, output, permission, and safety requirements.
  • Cached result: Return a prior result only when its age and context make it suitable; do not present stale information as current.
  • Degraded feature: Complete the parts of the workflow that remain reliable and clearly mark unavailable functionality.
  • Queue or defer: Preserve the task for later processing when the user does not need an immediate answer.
  • Human review: Escalate cases that require judgment, have uncertain side effects, or cannot meet the system’s quality or safety requirements through an automated fallback.

Use a circuit breaker or equivalent control when a dependency’s failure rate or latency indicates that continued calls are unlikely to succeed. The breaker should prevent repeated calls during the failure window and allow recovery checks rather than turning a temporary outage into an indefinite outage for the whole workflow. AWS’s Agentic AI Lens recommends testing fallback paths alongside primary implementations, including by injecting inference failures or inconsistent knowledge-base results.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Persist progress and validate each stage

Represent an agent workflow as stages with explicit inputs, outputs, and completion states. Save useful outputs at transitions and validate them before later stages consume them. If a late stage fails, the system can resume from the last valid checkpoint instead of replaying the entire task—and a malformed intermediate result is less likely to cascade into incorrect tool actions.

  1. Record stage state. Persist the stage identifier, its status, the relevant output or reference, and the correlation information needed to trace it.
  2. Validate before advancing. Check required fields, types, and business constraints before passing an output to another model call or tool.
  3. Resume from a known checkpoint. On recovery, determine which stages are complete and which side effects have occurred before restarting work.
  4. Keep recovery decisions observable. Record whether the run retried, switched routes, degraded, deferred, or escalated, along with the reason.

This staged approach follows the recovery pattern in the AWS Well-Architected Agentic AI Lens: persist and validate workflow stages, classify errors, retry transient failures within limits, and route persistent problems to degraded responses, cached answers, or human review.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Trace the run and measure its cost

Monitoring only the final answer hides where a workflow failed and what consumed resources. Trace the complete path across model calls, tools, queues, agent handoffs, and external services. Use a run or task identifier to connect those events, including asynchronous boundaries, and record the recovery decision at each failure.

Pair reliability signals with spend signals

Reliability view Cost and usage view Question it helps answer
Success and error rate by stage; error class Cost per completed task or outcome Which stage fails, and what does a successful outcome cost?
Retry and fallback rate Retries and tokens per task Are recovery attempts producing useful completions or mainly adding work?
Latency distributions, including percentiles Input, output, and reasoning tokens How does time to completion relate to model usage?
Throttling events and tool availability Subagent, tool, sandbox, and third-party charges Which dependency or delegated stage contributes to failure or spend?
Outcome quality and abandoned or no-op runs Spend on incomplete or abandoned work Is the system spending resources without delivering the intended outcome?

AWS CloudWatch documentation describes metrics including total and average invocations; total, average, input, and output token usage; average, P90, and P99 latency; errors and throttling; and cost attribution by application, user role, or user. Use the dimensions that fit the deployment, and make sure the trace covers communication between agent stages rather than only the top-level invocation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Attribute the whole task, not just its final model call

OpenAI’s Agents API documentation notes that one task may involve multiple model calls. Include root-agent and subagent work, retries, input and output tokens, cached input, reasoning tokens, and applicable tool, sandbox-compute, and third-party service charges in task-level accounting. Usage data may be best-effort, null, or revised, and is not necessarily a final bill. Cached input remains billable; a high cached-input share by itself does not show that the total task cost fell.

Cost per completed task or business outcome is generally more informative than cost per model request: a cheap call that triggers several retries, delegations, or paid tool actions may make the overall run expensive. Also track tokens and retries per task so a growing context or recovery loop is visible before it is mistaken for normal demand.

Turn the policy into a recovery path

A practical production policy can be expressed as a sequence of decisions for every stage:

  1. Capture the failure: Record the status, structured error, stage, and relevant turn or session state.
  2. Check progress and side effects: Read persisted outputs and determine whether external actions completed.
  3. Classify it: Correct permanent request problems; consider bounded retries for transient failures; route persistent capability failures to a fallback; escalate uncertain or judgment-heavy cases.
  4. Apply the budgets: Enforce attempt, delay, deadline, token, turn, and concurrency limits.
  5. Validate and continue: Save successful stage outputs, validate them, and resume from the last known-good state.
  6. Measure the outcome: Emit the final status, recovery path, latency, usage, and attributable costs for the complete run.

Test the recovery behavior as deliberately as the primary path: transient throttling, timeouts, persistent dependency failure, malformed intermediate output, and failure after an external action are distinct cases. Review whether the workflow stops safely, avoids duplicate side effects, preserves useful progress, and records enough information to explain both its reliability and its cost.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AWS’s Well-Architected Agentic AI Lens describes a common failure pattern in which retries are applied uniformly, including to non-retryable errors, without exponential backoff or jitter. The operational lesson is to make retry eligibility explicit and to connect it to the run’s deadline, state, fallback, and cost budget.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.