Build Anthropic API traffic handling around three steps: pace requests against your organization’s actual limits, classify errors before retrying, and make model fallback an explicit application decision. Anthropic enforces separate request, input-token, and output-token limits; its SDK retries transient failures twice by default, but not every 429 is temporary.
How Anthropic API rate limits work
For the Messages API, Anthropic measures limits in requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Limits depend on your organization’s tier and the model class; workspace limits can impose a lower ceiling beneath the organization limit. Treat these values as maximum permitted usage, not guaranteed capacity. Check your account in the Anthropic Console or Rate Limits API rather than hard-coding a tier-based assumption.
Anthropic says, “The API uses the token bucket algorithm to do rate limiting.” In practice, capacity replenishes continuously, so a workload that averages below its per-minute limit can still exceed available capacity during a short burst. Anthropic also documents acceleration-related 429s when usage rises sharply. Ramp traffic gradually and keep the request pattern steady.
Know which traffic shares a limit
Limits are organization-level and applied separately by model; requests using different inference_geo values share a pool. Most Claude models count uncached input tokens toward ITPM. Input usage is estimated when a request starts and adjusted as actual usage becomes known; OTPM is evaluated as the model generates tokens. The max_tokens setting does not itself consume OTPM.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Responses include headers for limit, remaining capacity, and reset time. Use them to observe headroom and tune pacing. A client-side queue or concurrency limiter is a practical way to smooth bursts, particularly when many workers or users can submit requests at once; it is an implementation recommendation, not an Anthropic requirement.
How to handle an Anthropic API 429
Do not treat every 429 as a signal to retry. Read the error type and response headers, especially retry-after, and distinguish a temporary rate limit from a spend cap.
Rank #2
| Response | Meaning | Recommended action |
|---|---|---|
429 with rate_limit_error and retry-after |
The organization exceeded a rate limit. Anthropic supplies a wait interval; requests sent sooner are expected to fail. | Wait at least the indicated interval before retrying, then resume at a controlled pace. |
| 429 caused by a tier spend cap | The account reached its monthly usage-tier spend cap. This case does not include retry-after and continues failing until access resumes. |
Stop automatic retries and address the account’s spend limit or wait for access to resume. |
| 429 caused by a Claude Code workspace spend limit | A workspace spend limit was reached; this is distinct from ordinary rate exhaustion. | Check the workspace limit and take the needed account or budget action rather than retrying indefinitely. |
Anthropic documents these 429 conditions in its API error guidance. If the response includes retry-after, honor it; if the header is absent, investigate whether a spend cap is responsible instead of assuming a short-lived throttle.
How many times does the Anthropic SDK retry?
Anthropic’s official SDKs retry transient failures—including connection errors, rate limits, and 5xx responses—with exponential backoff twice by default. They honor retry-after when supplied. You can adjust or disable automatic retries with max_retries; see the current Anthropic error documentation for SDK behavior.
Before adding application-level retries, account for the SDK’s attempts. A second unbounded retry loop can multiply traffic and latency during an outage. A sound implementation sets a finite attempt budget and an overall request deadline, records request IDs and error categories, and returns a clear failure once that budget expires. Those are application design recommendations, not platform guarantees.
Which errors should you retry?
Classify by error type as well as HTTP status. The same general response code can require different action, and a retry cannot fix every failure.
| Error | Anthropic’s description | Handling |
|---|---|---|
500 api_error |
Unexpected internal API error. | Retry with exponential backoff. If it persists, contact support and provide the request ID. |
504 timeout_error |
Request processing timed out. | For long-running Messages requests, consider streaming. Keep retries within the request’s deadline. |
529 overloaded_error |
Temporary API overload. | Retry with bounded backoff; if the operation still cannot proceed, use your application’s queue, error, or fallback policy. |
429 rate_limit_error |
May be rate exhaustion or a spend-limit condition. | Use retry-after where present; stop and investigate spend caps where it is absent. |
Streaming has a separate failure path: an SSE error can arrive after the server has already returned HTTP 200. Handle stream events as they arrive; ordinary request-level handling that only checks the initial HTTP status will not catch every mid-stream error.
When should you fall back to another model?
A model fallback is application policy, not an automatic consequence of retrying the same request. Anthropic’s direct API documentation does not define one universal fallback algorithm. Decide whether to queue or defer the operation, return a controlled error, or route it elsewhere only after classifying the failure and exhausting an appropriate bounded retry budget.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Best Value
- Classify the failure. Honor a temporary rate limit’s
retry-after; stop on a spend-cap 429 and surface the account action needed. - Retry only transient failures. Use finite exponential backoff and a request deadline, remembering that the official SDK already performs two retries by default.
- Check fallback eligibility. Confirm that the alternate model is active, authorized for the request, and acceptable for the task.
- Compare the operational fit. Assess task quality, output and tool/schema compatibility, latency, token costs, current lifecycle status, and geographic or data-routing requirements.
Model availability changes. Anthropic’s model deprecation guidance advises moving from deprecated models to suitable active replacements before retirement; requests to retired models fail. Verify the model’s status when configuring or revising a fallback rather than assuming an older model identifier will remain usable.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Should you use a gateway or managed fallback?
A single service can often pace its own traffic with an in-process limiter and queue. If multiple services or worker instances must coordinate, a shared gateway can centralize routing and observability, but it adds operational and security responsibilities. Choose based on the actual coordination problem rather than assuming a gateway is required.
| Approach | What it helps with | Trade-offs to consider |
|---|---|---|
| In-process limiter and queue | Simple pacing and burst smoothing inside one service. | Each process manages its own view; coordination across instances and centralized routing require additional design. |
| Shared gateway | Can coordinate traffic across clients and offer load balancing, fallback routing, usage tracking, and cost controls. | Adds a service to operate and a separate security and functionality review. Anthropic describes LiteLLM as a third-party proxy and explicitly says it does not endorse, maintain, or audit its security or functionality. |
Anthropic’s LLM gateway documentation describes LiteLLM as one third-party gateway option, not an Anthropic product. A separate note in Anthropic’s legacy Amazon Bedrock integration documentation points away from its server-side fallbacks parameter toward client-side fallback handling. That guidance is specific to that integration, not a universal direct-Claude-API setting. The same page distinguishes global endpoints, which dynamically route for availability, from regional endpoints intended for data-routing requirements.
Quick Recap
A practical operating checklist
- Read the current organization and workspace limits for the models you actually call; do not treat a published tier example as your account’s entitlement.
- Track RPM, ITPM, and OTPM separately, including remaining and reset headers.
- Smooth traffic with pacing or queuing, and ramp usage gradually to reduce burst and acceleration-related 429s.
- Branch on error type and relevant headers; distinguish spend-cap 429s from retryable rate limits.
- Understand SDK retry defaults, set a finite total latency budget, and log request IDs for persistent failures.
- Handle streaming SSE errors independently from initial HTTP status codes.
- Make fallback conditional on model status, task compatibility, authorization, and routing constraints.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

