Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A tool call that a model emits is a request for your application to act. It is not authorization, and it is not proof that anything happened. The question to ask first is the one that breaks most agent systems: what happens to a tool call after a side effect may already have happened? If a payment, email, ticket, or file write was dispatched and the response was lost, the transcript cannot tell the agent whether the change was committed. The production answer has four parts. Validate arguments and permissions inside the executor. Classify each outcome as a known failure, a confirmed success, or an unknown outcome. Bound every retry. Reconcile an unknown mutation against the system of record before any replay.

Treat each tool call as untrusted input

A tool schema tells the model what shape of arguments to produce. It does not decide whether the caller is allowed to perform the action, and it does not guarantee that the values make business sense. OpenAI’s Programmatic Tool Calling documentation places that responsibility in the application that executes the operation. Your executor should validate and authorize every call, and a high-impact action should pass through an approval step in the application workflow rather than rely on instructions in the prompt. Schemas constrain structure. Authorization and business rules belong in code that runs at execution time.

Questions to answer for every tool

Write the answers down before you expose a tool to an agent. They determine every later decision about validation and retries.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Which fields are required, which are bounded by length, range, or format, which are enumerated, and which depend on each other? A refund amount that is valid only for an order in a given state is a cross-field rule the schema cannot express.
  • Is the tool read-only, or can it change an external system?
  • Who may perform the action, and is that checked against the acting user or service identity at execution time rather than when the plan was made?
  • Is the operation naturally idempotent, or does it need a stable idempotency key or a deduplication record?
  • What does the executor return for a known failure, a confirmed success, and an unknown outcome? If the answer is only an exception, the agent has no basis for choosing a next step.

What the executor must re-check

Even when the model was given a schema, re-check the following inside the executing service:

  • Argument types, bounds, and cross-field rules.
  • The caller’s permission for this specific record, not only for the tool category.
  • Approval state for high-impact actions, recorded by the workflow and not inferred from conversation text.
  • Whether an identifier carried over from earlier in the conversation still points to a record in an acceptable state. Stale IDs are a common source of wrong-target mutations.

Classify the failure before deciding to retry

An HTTP status code is not a retry policy by itself. The decision depends on operation semantics: was the request rejected before it could have an effect, might it have reached the service and then failed, or was it throttled? The table below sorts the common cases. The handling column is the one to implement; the source column records what each cited guidance says and where it stops.

Outcome Typical handling Source and caveat
Invalid arguments or business-rule rejection Correct the input, or return the error to the agent or user. Do not resend the same request unchanged. OpenAI’s recovery guidance says to fix invalid input before retrying.
Authentication, authorization, or billing and configuration problem Resolve the credential, permission, or configuration first. The recovery guidance does not treat these as transient retry cases.
Rate limit or overload Honor Retry-After when the server sends it. Otherwise use a bounded, delayed retry. OpenAI’s recovery guidance says to honor Retry-After and to set an attempt limit or a deadline.
Network timeout or temporary service failure Establish whether the request may have reached the service. Retry only if replay is safe, or after reconciliation. OpenAI’s recovery guidance warns that a failed turn may already have called external tools.
Mutation with unknown completion Query status, deduplicate by operation identity, or reconcile against the system of record before any retry. AWS guidance on idempotent agent task execution and Google Cloud’s retry strategy documentation both warn that retries without idempotency can duplicate side effects.
Streamed model response or model-call failure Apply the model-layer replay policy, which is separate from the tool retry policy. The OpenAI Agents SDK documentation describes replay-safety checks that block some replays, including streamed runs after output has started, and local side-effect vetoes.

Reads and mutations need different rules

A repeated read usually carries less risk than a repeated payment, email, ticket creation, or record write. Google Cloud’s retry strategy documentation lists the following as always idempotent: “Always idempotent: List operations (they don’t modify resources), get requests, token count requests, and embeddings requests.” The same documentation makes the opposite point for mutations: “Unconditionally retrying non-idempotent operations can lead to side effects, such as duplicate resources.” Classify each tool as read or mutation before you write any retry wrapper around it, because the wrapper alone cannot tell the difference.

Bound every retry

Every retry loop needs a maximum attempt count, a deadline, or both. Without one, a transient fault becomes an unbounded queue of duplicated work. Retry only errors that plausibly recover, such as rate limits, overload, timeouts, and temporary service failures, and use the provider’s guidance to decide which errors qualify.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Attempt limits and backoff

Google Cloud’s guidance recommends exponential backoff with jitter, retrying only specific retryable errors, setting a maximum number of retries, and logging each attempt. Its documentation gives an illustrative sequence in which delays grow from 1 to 2, 4, and 8 seconds. That sequence shows the shape of exponential backoff. It is not a measured result and not a required production setting. The sources do not provide a universal retry count or delay for agent tools, so set both from the provider’s documented limits and from the latency your users will tolerate.

Provider hints are not a retry budget

When a provider sends a retry hint such as Retry-After, honor it. Keep your own limits anyway. A hint tells you when the provider expects another request to be accepted. It does not tell you how many attempts your agent should make, and it does not say whether the operation is safe to repeat.

Handle uncertain side effects

A known failure and an unknown outcome can look identical in a timeout log, yet they need different handling. A timeout usually means only that your process stopped waiting. The remote service may have committed the change and then lost the response. For that reason, check downstream state before repeating work, and do not ask the model to infer from the transcript whether the side effect happened. The transcript records what the agent observed, and in a timeout case the agent observed only the timeout.

Operation states the executor must report

Return one of three distinct states from every mutation. Merging them into a generic error is the most common cause of duplicate writes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
State Meaning Next step
Confirmed success The downstream system reports that the operation completed. For any duplicate dispatch, return the stored result. Do not execute the mutation again.
Confirmed failure The operation was rejected, or the system shows it had no effect. Surface the error. Retry only if the failure class is transient and the attempt limit has not been reached.
Unknown outcome A timeout, lost response, or dispatch without acknowledgement. Query status or reconcile the operation record before any re-dispatch. If neither source answers, escalate rather than guess.

A replay-safe pattern for mutations

The steps below are an implementation approach. They synthesize the official guidance to make calls idempotent and to check whether an action already completed. The sources do not prescribe this exact record format, and they do not claim that every downstream API accepts idempotency keys or that a single key format applies across vendors.

  1. Give each intended mutation a stable operation identity, generated before dispatch. Reuse that identity on every retry of the same intent.
  2. Persist the intent and its normalized arguments before dispatch, where your architecture allows it. Keep credentials and sensitive payload fields out of this record, and store a reference or hash instead if you need to match later.
  3. Pass a downstream idempotency key when the target API supports one. If it does not, keep a deduplication record in the tool service, keyed by the operation identity.
  4. Record the outcome as confirmed success, confirmed failure, or unknown, as separate states.
  5. On a timeout, query the downstream system or reconcile the record before dispatching again.
  6. When a duplicate arrives for an operation already marked as a confirmed success, return the existing result instead of running the mutation again.

Native idempotency keys move the state-keeping burden to the downstream vendor, where it is offered. A wrapper deduplication store puts that burden on you: you must operate the store, define how long records live, and clean them up.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Keep model-call replay separate from tool replay

Agent-level retries and tool-level retries act on different layers, and they need separate policies. A model request can be unsafe to replay when streaming has begun, when run state is involved, or when local side effects are possible. The OpenAI Agents SDK documentation describes replay-safety checks and fail-closed cases for these conditions. Replay rules are version-specific, so confirm the behavior in the SDK version you run before relying on it.

Replaying a model turn also has a tool-level consequence. A replayed turn can emit tool calls that were already dispatched, so the mutation controls above must hold even when the model layer is replay-safe.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Log the retry path

Instrument every tool attempt so that you can see retries as they happen, not only after a duplicate appears in a downstream system. Record:

  • The attempt number and the retry policy that produced it.
  • The error class from the taxonomy above, not the raw error message if that message can contain sensitive data.
  • Elapsed time, including time spent waiting in backoff.
  • The operation identity.
  • The final disposition: confirmed success, confirmed failure, unknown and escalated, or abandoned at the attempt limit.

Avoid logging credentials and sensitive arguments. Google Cloud’s guidance recommends logging and monitoring retry attempts, error types, and response times for this reason.

What the evidence establishes

  • The official guidance cited here, from OpenAI, the OpenAI Agents SDK, Google Cloud, and AWS, reflects documentation as of October 2026. Provider behavior changes, so check the current pages before you codify retry rules.
  • No published incident rates or prevalence figures for duplicate side effects in agent systems were found in these sources. The absence of figures is not evidence that duplicates are rare. It means the risk has to be controlled by design rather than estimated from production statistics.
  • Further reading on the systems fundamentals behind these controls, including idempotency, replication, and failure handling, is covered in Martin Kleppmann and Chris Riccomini’s Designing Data-Intensive Applications, 2nd Edition. The book is not an agent-tool operations manual, but it covers the distributed-systems reasoning that the operation-state pattern depends on.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.