Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Retry an LLM API call only when the provider indicates the failure may be temporary. For those failures, use bounded exponential backoff with jitter, respect any documented retry-after hint, and stay within both a per-request deadline and a service-wide retry budget. A circuit breaker complements retries: it temporarily blocks calls when failures persist, then allows limited probes to test recovery.
How should you decide whether to retry an LLM API call?
Classify the failure before scheduling another attempt. A retry is useful when waiting could change the outcome; it will not repair an invalid request, bad credentials, missing permissions, or an exhausted billing allowance. Check the provider’s structured error type and response body, not just the HTTP status.
- Potentially transient: documented throttling, temporary overload, server errors, network failures, and timeouts may be retryable, depending on the provider and endpoint.
- Usually requires a change first: malformed input, authentication or authorization failures, and payment or spend-limit errors should not be retried unchanged.
- Ambiguous: a timeout can mean the provider did not complete the request, or that the client lost the response after processing. Consider duplicate cost or side effects before repeating the operation.
Status codes are clues, not a universal LLM API contract. Providers can attach different conditions to the same status family, and even a 429 may reflect a limit that will not clear after a short wait.
Examples from provider documentation
Google’s Gemini troubleshooting guide recommends backoff for retryable errors and gives 429 RESOURCE_EXHAUSTED and 503 UNAVAILABLE as examples. It also advises retrying only selected transient failures, such as 429, 408, or 5xx, rather than client errors such as 400, 402, or 403. Google’s Gemini API error reference says 429 and 503 errors should wait and retry with exponential backoff; it says 500 errors may be retried, while 504 deadline-exceeded errors call for examining or adjusting the client deadline. Its payment-required guidance calls for addressing credit or auto-reload rather than repeating the unchanged request.
#1 Best Overall
Anthropic documents 429 rate-limit errors, 500 internal errors, 504 timeouts, and 529 overload errors for Claude. Its API error documentation recommends exponential backoff for 500 errors and notes that some spend-limit 429s may lack a retry-after header and persist until access resumes. Treat the subtype and provider’s current instructions as decisive, rather than assuming all 429s are short-lived rate windows.
What is the right exponential backoff for an LLM API?
There is no single schedule that fits every provider, endpoint, or application. For eligible failures, increase the wait between attempts exponentially, add random jitter, and impose three bounds: a maximum attempt count, a maximum delay, and a total elapsed-time budget. Jitter prevents many clients that failed together from retrying in a synchronized burst. Google puts the reason plainly: “Add random ‘jitter’ to the delay to help prevent all clients from retrying at the exact same time.”
If the response includes Retry-After or a provider-specific equivalent, follow its documented meaning. For Azure OpenAI, Microsoft guidance recommends honoring retry-after-ms when present. Do not assume that every provider supplies the same header or that it applies to every error subtype.
Fit the policy to the work
- Interactive requests: set the retry budget inside the user’s response deadline. A retry that finishes after the caller has timed out can add load without improving the experience.
- Background jobs: if the job can wait, a longer schedule may be appropriate. Keep attempts bounded and expose a final failure or queue the job for later rather than retrying indefinitely.
- Costly or side-effecting calls: check whether the provider and endpoint support idempotency. If a timeout leaves processing uncertain, a repeated call could incur duplicate cost or effects.
These are policy choices, not universal numeric settings. Choose them against your latency tolerance, provider behavior, request cost, and workload, then observe whether they are preventing recovery or worsening congestion.
Rank #3
How do SDK retries affect your retry policy?
Count retries across every layer. An SDK may retry internally while your application also retries the call; multiplying those policies can produce far more provider attempts and elapsed time than either layer suggests on its own. Read the SDK’s current documentation and decide whether to configure its behavior, disable it, or include it in your application’s total limits.
Google’s troubleshooting page gives one version-specific example for its Gemini Python SDK: up to four retries, an initial delay of approximately one second, and a maximum delay of 60 seconds. This is documented SDK behavior, not a general recommendation or a promise for every language or future SDK version.
Rank #4
What does a circuit breaker do for API calls?
A retry assumes the call may work if attempted again after a delay. A circuit breaker addresses persistent failure by temporarily stopping calls so they fail quickly instead of continuing to consume resources or adding pressure to an unhealthy dependency. Microsoft notes that “The Circuit Breaker pattern serves a different purpose than the Retry pattern.” The patterns can be combined: retry eligible failures through the breaker, and stop retrying when the breaker is open or signals a non-transient failure.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →The three breaker states
- Closed: calls reach the provider and relevant failures are counted over a configured window.
- Open: after the configured failure condition is reached, calls are rejected locally for a period rather than sent to the provider.
- Half-open: after a cooldown, a limited number of probe calls are allowed. Success closes the breaker; failure opens it again and restarts the cooldown.
Failure thresholds, observation windows, cooldowns, and probe counts are tunable policy, not established LLM-wide defaults. Set them in light of traffic, user latency tolerance, provider behavior, and request cost. Make an open-circuit response distinguishable from a provider error so callers and operators can tell why the request did not proceed.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How do you keep retries from making an outage worse?
A limit on each individual request is not necessarily a limit on total traffic. If many requests fail at once, each may retry within its own allowance and collectively burden a recovering provider. Use a service-wide retry budget, rate limit, or concurrency control to cap that combined load. When the budget is spent, return a bounded failure or queue delay-tolerant work where the user experience allows.
Also avoid independent retry loops at the SDK, application, job queue, and gateway layers unless their combined attempts and timing are understood. A retry should have one clear owner or be coordinated across layers, with the same overall deadline and attempt budget.
What should you check in each provider’s current contract?
Provider behavior changes, and quota rules can depend on model, endpoint, account, project, or tier. Before choosing a policy, verify the current documentation for the exact API and SDK version you use.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitches- Which structured error types and status codes are documented as transient?
- Does the response provide a retry or reset hint, and what does it mean for that specific error?
- Does the SDK retry automatically, and can those retries be configured or disabled?
- Are limits based on request rate, tokens, spend, project, organization, model, or usage tier?
- Could a timeout happen after the provider processed the request, and is idempotency supported?
For Gemini, Google’s rate-limit documentation describes requests per minute, input tokens per minute, and requests per day. Limits vary by model and usage tier, apply per project rather than per API key, and can change with account tier and status. Check the current account limits in AI Studio instead of treating a published example as a universal quota.
What should you log and monitor?
Record enough information to distinguish a provider incident from a client or policy problem. Track attempts and final outcomes, elapsed latency, classified error types, breaker state changes, and provider request identifiers when available. Breaker state-change events can help monitor dependency health; logs also make it possible to spot retry multiplication, repeated permanent errors, and duplicate-risk cases.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

