What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Do not use a free inference tier as an unbounded drain for a dead-letter queue (DLQ). Replaying failed messages sends them back through model calls, and those calls consume request and token quota that is often shared, capped, and hard to observe. A small, rate-limited replay can fit inside a provider’s free allowance, but only after you confirm the current limits and watch the replay as it runs. The rest of this article explains how to tell those cases apart and how to run a replay safely.

Separate the two kinds of retry before you replay

Most confusion about replay comes from treating two different retry layers as one.

Delivery retries belong to the messaging or event service. When a consumer cannot accept a message, the service retries delivery according to its own policy, and after the policy is exhausted it can park the message in a dead-letter queue. Amazon’s documentation for Simple Notification Service describes the DLQ as “an Amazon SQS queue that an Amazon SNS subscription can target for messages that can’t be delivered to subscribers successfully.” Amazon EventBridge uses a similar model, with its own retry policy and a DLQ for events that never arrive. Cloudflare Queues also routes messages that exhaust their retry limit to a dead-letter queue.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Application-level inference retries belong to your code. Your worker calls a model endpoint, the call fails or is throttled, and your code tries again. Each of those calls is a request to the inference provider, and it counts against that provider’s quotas.

A DLQ replay is a second layer on top of both. It takes messages the messaging service has given up on and feeds them back into your pipeline, which then makes fresh inference calls. If your pipeline also retries on its own, one stored message can produce several model requests before it succeeds or fails for good.

Why free inference is the wrong place for an unbounded replay

A DLQ often holds a backlog that accumulated during an outage, a bad deploy, or a provider incident. Replaying that backlog all at once turns a slow trickle of work into a burst. Three things make the burst worse on a free allowance:

  • Quotas are per account and per model. A replay can exhaust the request-rate or token budget that your live traffic also depends on, so customer-facing requests start failing while the replay runs.
  • Retries multiply demand. Delivery retries, application retries, and the replay itself can all hit the same endpoint for the same message.
  • Queue operations also cost something. Cloudflare’s pricing example states the point directly: “Each retry incurs a read operation.” Its DLQ writes are also counted as operations. Replay therefore adds queue activity on top of inference activity, even when the queue is on a free allowance.

DigitalOcean’s guidance on 429 responses explains why a quota problem should not be treated as an ordinary error: “A 429 response means your account reached one of its own limits (a request-rate limit or a model’s token limit), or a platform overload.” A replay that ignores that signal keeps spending the same exhausted budget.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The recommendation here is an engineering inference from these documented behaviours: queue retries and DLQ writes are metered, inference quotas are enforced per account, and a replay adds demand that the original failures did not. It is not a universal rule that every free tier prohibits replay, and it is not a claim that replay always produces charges.

What the published figures do and do not establish

The numbers below come from different services and measure different things. They are useful for planning, but they are not interchangeable and they are not recommended settings.

Service Figure Unit and scope Qualification
Amazon EventBridge Default retry policy: 5 attempts or 300 seconds Delivery retries for an event target Default value on AWS’s “Retry policies and dead-letter queues” page as accessed in 2026. It is a default, not a recommendation.
Amazon EventBridge Configurable range: 0 to 185 attempts; 60 to 86,400 seconds Retry attempts and maximum event age Documented range on the same page, accessed in 2026. Confirm against the current page before configuring.
Cloudflare Queues Default DLQ retention: 4 days How long a dead-lettered message is kept From Cloudflare’s “Dead Letter Queues” documentation, last updated 2026-04-21.
Cloudflare Queues 1,000,000 free queue operations; $0.40 per additional million operations Queue operations, as shown in a pricing estimate From Cloudflare’s pricing page as accessed in 2026. Prices and allowances change; check before budgeting.
DigitalOcean Serverless Inference 5,000 requests per hour; 250 requests per minute Request limits, per provider and plan Values reported in DigitalOcean’s Serverless Inference API reference as indexed in 2026. Limits vary by plan and change over time.
DigitalOcean Serverless Inference Not stated Token quota values Token limits are per model and per account. Read the values returned in quota headers for your account rather than relying on a published figure.

Do not set a queue operation allowance against an inference request limit. They measure different things, and a message that fits comfortably inside one can still exceed the other.

Classify each failure before you replay it

Replaying a message without knowing why it failed spends capacity to learn nothing. Sort the failed messages first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Failure class Typical signal Replay? First action
Transient Timeouts, platform overload, intermittent errors that clear on their own Yes, bounded and with backoff Confirm the dependency has recovered, then replay in small batches
Quota 429 responses that name a request-rate or token limit Only after the applicable limit resets Read the reset information and schedule the replay around it
Malformed input Validation or parsing errors that repeat for the same message No, not unchanged Fix the producer or move the message to a quarantine path
Authentication or configuration Rejected credentials, wrong endpoint, missing permissions No, until the configuration is corrected Fix the key, endpoint, or permission, then test with one message
Model-specific Failures that cluster on one model or model version Only after changing the model or the input Route to a working model or repair the input, then replay

Permanent causes must be fixed before replay. Replaying an unchanged poison message repeats the same failure and spends quota each time.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A bounded replay procedure

  1. Stop the source of new failures. If the cause is still active, a replay only adds to it. Pause the producer or confirm the fix is deployed.
  2. Select the records. Use the failed record identifiers and timestamps to choose a specific range. AWS’s documented EventBridge redrive procedure follows this pattern: inspect the failed records, fix the named cause, then replay a selected range rather than the whole queue.
  3. Mark replayed work. EventBridge identifies replayed deliveries through service metadata; check the current documentation for the exact field. In your own pipeline, add a replay flag or field so downstream code and dashboards can tell replayed work from first attempts.
  4. Cap batch size and concurrency. Set both below the lowest limit you know applies to the model, and keep them well under the free allowance so live traffic keeps headroom.
  5. Track attempts per message. Store a retry counter with each message. AWS’s older Compute Blog example on SQS dead-letter queues shows the pattern: count retries, add a delay before each one, and escalate messages to human review once the count passes a threshold. That article is an illustration, not a current AWS guarantee.
  6. Back off between attempts. Increase the delay after each failure and honour any reset information the provider returns.
  7. Define a terminal path. After a set number of attempts, send the message to a manual-review queue or log it as failed. Do not let it cycle back into the DLQ indefinitely. Choose your own limit; AWS’s default of 5 attempts is a vendor default, not a recommendation.
  8. Make processing idempotent. Key each operation on a stable message identifier and store results so a repeated delivery does not produce duplicate side effects or a second paid call.
  9. Monitor the run. Watch queue depth, age of the oldest message, retry count per message, the number of 429 responses, and the number of successful completions. Stop the replay if 429s rise or completions stall.

Handling a 429 from serverless inference

A 429 during replay is a signal to slow down, not to try harder. The usual mistake is to retry immediately in a tight loop, which uses up the remaining quota and extends the outage for live traffic.

  • Read the quota information in the response. DigitalOcean documents quota-specific response headers for Serverless Inference, including request and token quota information and Retry-After behaviour. Use those values rather than a fixed guess.
  • Wait for the applicable reset. If the response indicates a reset time, schedule the next batch after it.
  • Stop the batch if 429s keep arriving. Return the unfinished messages to the queue with their attempt counts intact.
  • Do not treat every 429 the same way. A request-rate limit clears in a short window, while a token quota may need a longer reset. The response tells you which one you hit.

When a free allowance can carry a replay

A free allowance is reasonable for a small, bounded recovery when every condition below holds:

  • The number of messages is known and small enough to finish well inside one quota window.
  • You checked the provider’s current limits on the day you run it, including the plan that applies to your account.
  • Batch size and concurrency are capped below those limits, leaving headroom for live traffic.
  • Replayed work is marked, attempts are counted, and the run can be stopped at any time.
  • You accept that the trickle is subject to the vendor’s current terms, which can change.

If the recovery work regularly exceeds a free allowance, the decision is a cost and capacity question. Metered hosted inference lets you pay for the throughput the recovery needs, and that is a budgeting choice to make against the measured volume of your own replays, not a reason to drop the bounded procedure above.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common mistakes

  • Assuming a free queue makes replay free. Queue operations and inference requests are counted separately, and both can be metered.
  • Counting only one retry layer. Delivery retries, application retries, and replay can stack on the same message.
  • Replaying the whole DLQ at once. A bulk replay turns a backlog into a burst that can exhaust quota for live traffic.
  • Skipping idempotency. Repeated delivery without deduplication can produce duplicate outputs and duplicate paid calls.
  • Quoting vendor defaults as settings. Retry counts, retention periods, and request limits are documented values that differ by service, plan, and date.

Confirm the current figures on each provider’s documentation page before you configure a replay, since limits and prices change without notice.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.