Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An idempotency key only prevents duplicate work if every retry uses the same key and the server can consult a durable, concurrency-safe record of what happened. If a mutation commits but its response times out, a retry that is treated as a new operation can duplicate, overwrite, or otherwise corrupt state. The fix is to make the operation identity, its result, and recovery behavior part of the complete side-effect workflow—not just to add a request header.

How a retry can lose or duplicate a record

A timeout tells the client that it did not receive a response in time. It does not tell the client whether the server committed the mutation. The server may not have started processing; it may have failed partway through; or it may have completed successfully and lost the response on its way back. Stripe describes these as distinct ambiguous outcomes. If the client retries without a stable operation identity, the server cannot reliably tell which outcome occurred.

For example, suppose a client sends a request to create a payment or database record. The server commits it, but the connection drops before the client receives the result. If the retry uses a new key—or the original key’s record has disappeared—the server may create the operation again. Conversely, a weak implementation might see a pending key, discard the retry, and never finish the operation. Either way, the caller can end up with an incorrect view of what happened.

Idempotency is therefore not a guarantee that code executes only once. At-least-once delivery and retries can cause code to run more than once. The goal is for repeated attempts with the same operation identity to produce one intended effect and a consistent, recoverable outcome.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What idempotency does—and does not—mean

RFC 9110 defines an HTTP method as idempotent when multiple identical requests have the same intended effect on the server as one request. Safe methods, along with PUT and DELETE, are defined as idempotent by the standard. That does not mean every request using those methods is harmless to repeat in every application, nor that every response will be identical. The application’s actual side effects still matter.

Do not automatically retry a non-idempotent request just because a network error occurred. RFC 9110 cautions against retrying such a request unless the client knows its application semantics make the retry safe. An application can make an operation such as a POST safe to repeat by assigning it a stable idempotency key and implementing durable deduplication correctly.

A request header by itself does not provide exactly-once processing. AWS’s Well-Architected guidance describes an idempotent service in terms of repeated identical requests having the same effect as a single request. Achieving that effect requires the entire mutation path—including database writes, external calls, queues, and consumers—to respect the same operation identity.

Ten failure modes that make idempotency lose data

1. The retry gets a new key

A key identifies one logical operation, not one network attempt. Generating a fresh key for each retry makes each attempt look new to the server. Generate the key once when the logical operation is created, then reuse it for every retry and replay of that operation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Two different operations share a key

If unrelated operations collide, the server may return one operation’s stored result for the other or suppress a legitimate mutation. Use a high-entropy random value, such as a UUIDv4, rather than a predictable or low-entropy identifier. The key’s scope should also be clear: for example, bind it to the relevant account or operation type if keys are not globally unique in your system.

3. Concurrent requests race through check-then-insert

A sequence like “look up key; if absent, run mutation; then save key” is not safe by itself. Two workers can both observe that the key is absent and both execute the mutation before either saves a record. AWS recommends closing this race with a uniqueness constraint, a conditional write, or an atomic transaction. Examples include a unique database constraint with INSERT ... ON CONFLICT DO NOTHING or a DynamoDB condition such as attribute_not_exists.

4. The key record lives only in volatile or isolated storage

A cache eviction can erase the evidence needed to recognize a retry. A record held only in one region can be invisible to a retry routed elsewhere. Store the operation record durably and make it visible wherever retries may be handled; a cache can accelerate lookups, but should not be the only source of truth when losing it could repeat a side effect.

5. The mutation commits but its result is not saved

If the database commit succeeds and the process crashes before recording the operation’s result, a retry may not know whether to run the mutation again. Where possible, claim the key, apply the mutation, and save the outcome in one database transaction. If the side effect is outside that transaction, use an outbox or durable workflow and make the external call idempotent too.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

6. Key expiry is shorter than the retry or replay window

Once a key record has been removed, a late retry can be mistaken for a new operation. Stripe documents that it automatically removes keys only after they are at least 24 hours old; this is a minimum age before pruning, not a promise that every key remains available for exactly 24 hours or indefinitely. Choose retention to cover your own realistic client retries, queue delays, manual replays, and recovery procedures.

While a Stripe key remains available, reuse with different request parameters is rejected rather than treated as the original request. Your service should define an equally clear rule: bind each key to a request fingerprint or canonical parameters, and reject changed input instead of returning another request’s result.

7. A worker stops between side effects and completion

A worker may complete one step—such as charging an external service or publishing a message—and crash before marking the operation completed. On replay, blindly executing the step again can duplicate the side effect; refusing to resume can leave the operation stuck. Persist enough workflow state to resume, reconcile the outside system, or safely repeat each step.

8. The operation identity is lost downstream

An API can deduplicate a request correctly and still produce duplicates later if it sends a queue message or calls another service without propagating a stable identity. Carry the operation ID or a deterministic event ID across every boundary. Downstream services and message consumers must deduplicate too; an upstream token cannot protect a side effect that never sees it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

9. A retry increments a counter twice

Increment operations are especially easy to repeat accidentally: applying “add one” twice is not the same effect as applying it once. Guard increments with a condition, transaction, or deduplicated operation record rather than assuming that the request will arrive only once.

10. A timestamp is used as the key

Timestamps are not reliable uniqueness tokens. Simultaneous clients can generate the same value, and clock skew can create collisions or unexpected ordering. Use a random high-entropy key generated once per logical operation instead.

Build a durable idempotency flow

  1. Create one operation identity. Generate a high-entropy key when the user or system initiates the logical operation. Persist it with the client’s operation state and reuse it across retries; do not regenerate it after a timeout.
  2. Claim the key atomically. Use a database uniqueness constraint, conditional write, or transaction so only one concurrent request can claim a key. A preliminary read followed by an unguarded insert is not sufficient.
  3. Bind the key to the request. Save a fingerprint or canonical representation of the relevant parameters. If the same key arrives with different parameters, reject it clearly rather than replaying a result that belongs to different input.
  4. Persist explicit operation state. Track states such as pending, completed, and failed in durable storage. Define what each state means, how a stale pending operation is recovered, and who is allowed to resume it.
  5. Make the state change and result durable together where possible. In a single database boundary, commit the mutation and its idempotency record in one transaction. When an external side effect prevents that, use an outbox or durable workflow, and ensure the external operation itself accepts or derives a stable deduplication identity.
  6. Save enough outcome to replay consistently. Keep the original status and response body, or a durable reference that can reproduce the defined result. Stripe says it stores the first request’s status code and body for a key, including a 500 response. Decide explicitly whether failures are replayed as recorded or whether a particular failure state is safe to resume; do not leave the behavior accidental.
  7. Propagate identity to every next step. Include the operation or deterministic event ID in queue messages and downstream calls. Deduplicate at each consumer or service boundary, not only at the API entry point.
  8. Retry selectively and with backoff. Retry only when the operation contract makes it safe. Use bounded exponential backoff with random jitter to avoid synchronized retry storms, as Stripe recommends.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose an implementation that matches your side effects

The right storage and replay contract depend on where the mutation happens. A single database transaction is simpler when all relevant effects fit inside one database boundary. External services and asynchronous workflows need durable coordination and explicit recovery.

Design choice Useful when Trade-off to resolve
Unique constraint or conditional write Concurrent workers may claim the same key; the datastore supports atomic uniqueness or conditions. Decide what the losing request reads or returns after the atomic claim fails.
Single-database transaction The operation record and mutation can be committed in the same database. It provides a tight atomicity boundary, but cannot by itself include a separate external service or queue.
Durable workflow or outbox The operation spans a database and an external service, or includes asynchronous steps. Persist step state and define resume or reconciliation behavior; every external step still needs retry-safe identity.
Stored response replay Callers need the original status and body even if the resource later changes. Persist the response contract and its data for the required retention period.
Status plus resource lookup A stable resource can be retrieved to establish the operation outcome. Ensure the lookup is durable and unambiguous, and specify what clients receive for errors or incomplete work.
Cache-only key record Only when losing or expiring the entry cannot cause an unsafe duplicate effect. Eviction and regional isolation can erase deduplication evidence; it is not a durable record by itself.

Also choose a retention window deliberately and verify whether all regions and workers can see the same operation record. A design that works for a fast client retry may still fail for a delayed queue redelivery or an operator replay days later.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Test the ambiguous outcomes, not just the happy path

Tests should force the exact boundaries where the caller and server can disagree. A successful response on the first attempt does not demonstrate safe retries.

  • Commit the mutation, then drop the response or simulate a client timeout. Retry with the same key and verify that the intended effect occurs once and the client receives the defined outcome.
  • Send two requests with the same key concurrently. Verify that a database constraint or conditional write allows only one claim and that both callers converge on a consistent result.
  • Reuse a key with changed parameters. Verify that the server rejects the mismatch instead of replaying a result for different input.
  • Crash a worker after an external side effect but before it records completion. Verify that recovery reconciles or safely resumes rather than duplicating or abandoning the work.
  • Deliver the same queue message more than once and route retries through another worker or region. Verify that downstream consumers see the same stable identity and deduplicate it.
  • Exercise a retry after the configured retention period. Verify that expiry behavior is understood and cannot silently transform an old replay into an unintended new mutation.
  • Leave a record in pending and verify that a recovery path detects and resolves stale work rather than leaving it stuck forever.

Review checklist for a production implementation

  • Is one high-entropy key generated per logical operation and reused for every retry?
  • Can a database uniqueness constraint or conditional write prove that concurrent requests cannot both claim the key?
  • What happens when the server commits but the client receives a TCP timeout?
  • What happens if a worker dies after an external side effect but before marking completion?
  • Does a retry with changed parameters fail clearly?
  • Is key retention longer than the maximum realistic retry, queue-redelivery, and replay window?
  • Can any region or worker that receives a retry read the same durable operation record?
  • Can pending records be recovered, and are stale ones detected?
  • Do queue messages, downstream calls, and consumers carry and honor a stable operation identity?
  • Are inserts, deletes, and increments guarded according to their actual semantics?
  • Are automatic retries limited to safe cases and controlled with bounded backoff and jitter?

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.