Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An AI agent usually runs the same action twice because it cannot tell whether an earlier attempt succeeded. A tool call reaches an external service and creates a side effect, the response is lost or delayed, a timeout fires, and the agent retries. Workflow replay after a crash and parallel workers handling the same task can also run one logical step more than once. Repeated execution does not have to mean a repeated effect. If the operation is idempotent, repeated requests for the same logical action produce the same final outcome as a single request.

The core problem: a timeout does not tell you what happened

A timeout means the caller did not receive an answer in time. It does not prove that the external action failed. A retry is often necessary to make progress, but if the first request already completed, a blind retry repeats the side effect. AWS’s Well-Architected Agentic AI Lens, in the guidance titled “AGENTREL06-BP04 Implement idempotent task execution patterns,” puts it directly: “Retry is the most common recovery mechanism, and without idempotency it can produce duplicate side effects.”

Three failure patterns produce most duplicate actions in agent systems.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pattern 1: the lost response

  1. The agent sends a request, such as creating an order or sending an email.
  2. The receiving service processes the request and commits the change.
  3. The response is lost on the network or arrives after the caller’s timeout.
  4. The agent records a failure and retries the same call.
  5. Unless the service recognizes the repeat, the change happens a second time.

This sequence is why the interruption point matters more than the retry logic itself. The same retry is harmless when the operation is read-only and harmful when it writes.

Pattern 2: replay after checkpoint recovery

Durable execution systems and agent orchestrators often recover by replaying a workflow from saved state. AWS’s Durable Execution SDK developer guide, in the section “Idempotency and retries,” describes both at-least-once and at-most-once behavior per retry and notes that replay can re-reach steps that have external side effects. Checkpoints tell the system where to resume, but they do not tell the external service that a request was already applied. Checkpointing therefore helps recovery without preventing duplicates on its own.

Pattern 3: concurrent workers on the same work

Distributed processing can assign the same logical work to more than one worker, especially after a failure. Google Cloud’s documentation on exactly-once processing in Dataflow states that a transform may run more than once or simultaneously on multiple workers after failures, and that output deduplication is what keeps results correct. That behavior is specific to Dataflow’s model, but the underlying problem is common to any system that retries or runs work in parallel.

Execution count is not effect count

Idempotency is a property of the effect, not of the code path. The same function can run twice while the external world changes only once. An idempotent payment API, for example, deduplicates requests at the service boundary and may return the first result when it sees the same request again. The agent still executed its step twice, but the customer was charged once.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This distinction matters because it defines what you can and cannot promise. You cannot prevent an agent from invoking a tool twice in every architecture. You can, in many cases, make the second invocation harmless. The phrase “exactly once” should be used only when you have evidence for the whole path, from the agent step through the orchestrator to the external service.

How to make retries safe

The following sequence is the practical pattern that AWS’s agentic guidance describes, adapted to a typical tool-calling agent.

  1. Assign one key per logical action, not per attempt. Derive the key deterministically from stable inputs such as the workflow ID, the task type, and the request content. A random ID or a timestamp generated at retry time creates a new identity, so the service cannot recognize the repeat.
  2. Check for an existing result before the side effect. If a prior success is stored under that key, return the stored result instead of executing again.
  3. Record the key and result at the side-effect boundary. Use a conditional write, a unique constraint, or an equivalent atomic deduplication record so that two concurrent attempts cannot both succeed.
  4. Reuse the same key on every retry. Pass it through delegated agent steps, checkpoint resumes, and external API calls that accept an idempotency key.

A correct setup shows one stored result per key, no matter how many attempts the logs contain. If your logs show two successful external records with one key, the deduplication boundary is missing or sits in the wrong layer.

Choosing a recovery policy

Recovery policies trade off between making progress and avoiding repeats. The table below compares the common options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Policy What happens after an interruption Duplicate-effect risk What you must add
At-least-once execution The step is retried until it completes Repeats are possible; the operation runs again if the earlier attempt is unconfirmed Idempotent operations and a stable key on every side effect
At-most-once per retry The step is not re-executed after an interruption No automatic repeat, but the outcome can remain unknown A reconciliation path that checks whether the effect happened
Workflow-level retry The orchestrator restarts a larger unit of work A new attempt can re-run steps that already had effects Key propagation across the whole workflow and idempotent side effects
Checkpoint resume Execution continues from the last saved state Steps executed after the last checkpoint may run again Idempotent side effects; checkpointing alone is not enough

The AWS durable-execution documentation describes at-least-once and at-most-once per-retry behavior as separate semantics with different trade-offs. Choose the policy per step, not once for the whole agent. A read-only lookup can safely run at least once, while a payment or email send usually needs a key or a reconciliation step.

Where the protection has to live

Deduplication protects only the layers that enforce it. A key stored inside the agent process does nothing if the orchestrator restarts on another machine and the external service never sees the key.

Protection layer What it covers What it does not cover
Agent step Retries inside a single agent call Orchestrator replay, parallel workers, or a restarted process
Orchestrated workflow All steps that share the workflow-derived key External services that ignore the key or lack idempotency support
External service The actual side effect, such as a charge or record creation Duplicate requests that never carry the key

AWS’s guidance recommends propagating keys through multistep workflows and into external systems that support them. The strongest protection is at the external boundary, because that is where the effect occurs.

When the external API has no idempotency support

Some APIs offer no idempotency key and no way to look up a prior request. For those, the design has to choose one of three paths:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Stop for human reconciliation. The workflow pauses on an uncertain outcome and a person confirms whether the action happened.
  • Use an application-side deduplication ledger. Your system records each action before calling the API and checks the ledger before every retry. This only works if the external state can be queried or is controlled by your own code.
  • Accept a documented risk. Write down which actions can repeat, how often you expect it, and who owns the cleanup.

Do not assume a generic exactly-once guarantee in this case. Without an idempotent boundary, there is no evidence that the whole path executes each effect once.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Test the interruption points before you trust the design

Retry logic is usually tested with a clean failure, such as an error response. Duplicate effects appear at the boundaries where the outcome is unknown. Test these three points for every side-effecting tool:

  • Before the request leaves the caller. Kill the process after the key is generated but before the call is sent. The effect count should be zero, and the retry should produce one.
  • After the receiver commits but before the response returns. Simulate a dropped response. The retry should return the stored result and create no second effect.
  • After the caller receives success but before it records the result. Crash the agent at that moment. On resume, the key should find the prior success and skip re-execution.

Run each test and count the external effects, not the tool invocations. A passing test shows one effect with one key, even when the trace shows several calls.

What the evidence does and does not establish

The guidance above comes from AWS’s Well-Architected Framework and Agentic AI Lens, the AWS Durable Execution SDK developer guide, and Google Cloud’s Dataflow documentation. These describe their own services. Before stating the retry defaults, idempotency support, key retention period, or delivery semantics of a particular agent framework or orchestrator, check that framework’s current documentation. The same applies to any API you call.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

No measured failure rate or cost estimate for duplicate agent actions was established in these sources, so this article does not offer one. The problem is well documented as a design hazard; how often it occurs in a given system depends on that system’s failure rate and retry settings, which you need to measure yourself.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.