Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

API rate limiting controls how often a defined client or traffic source can make requests. Set the limit against a clear identity and resource risk, reject excess requests with HTTP 429, and give clients a safe way to retry. It helps protect backend capacity, but it cannot cap the cost of an individual request or replace authentication and broader denial-of-service defenses.

What an API rate limit controls

A rate limit sets an allowance for requests over time, such as a maximum number of requests within a period or a steady request rate. When a client exceeds that allowance, the service can reject further requests, delay them, or apply another defined policy.

Every policy needs a scope—the routes, methods, or resources it covers—and a key that identifies whose traffic is counted. A limit by IP address, authenticated user, account, or credential protects different interests. A request-count limit also does not bound the work inside each request: expensive searches, large exports, high result counts, and large payloads need separate controls. OWASP recommends assessing resource consumption as part of API security, not relying on request frequency alone (OWASP API4: Unrestricted Resource Consumption).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I rate limit an API?

  1. Identify the risk. Decide whether you need to prevent brute-force attempts, protect a costly operation, enforce a customer quota, or preserve overall service capacity. Different risks may need different policies.
  2. Choose the policy scope and key. Specify which route or operation is covered, how requests are grouped, the allowance and time behavior, and what happens when capacity is exhausted.
  3. Select an algorithm that fits the traffic. Decide how much burst traffic is acceptable, whether the policy should apply over a rolling interval, and whether excess work should be rejected or smoothed.
  4. Enforce it where the relevant traffic can be identified. A gateway can apply limits before requests reach application services; application-level controls can handle identities or costs known only inside the service. Verify the actual enforcement scope and guarantees of the chosen system.
  5. Define rejection and recovery behavior. Return 429 when refusing a request for exceeding a rate limit, and provide a retry delay when the service can give clients a useful one.
  6. Test boundary and overload cases. Check bursts, shared client identities, distributed sources, retries, and costly requests. Confirm that the limiter protects the intended resource without unfairly blocking legitimate traffic.

Which rate-limiting algorithm fits?

Algorithms differ in burst tolerance, recovery, interval-boundary behavior, storage and distribution needs, and operational simplicity. The labels below describe common conceptual behavior; an implementation’s exact guarantees depend on its documentation and configuration.

Approach Behavior and trade-off Useful when
Token bucket Tokens refill at a configured rate and requests consume them. A bucket can hold a bounded burst, so requests may arrive faster than the steady refill rate until that capacity is used. You want a steady allowance with some controlled burst tolerance. AWS API Gateway documents this model for throttling.
Sliding window Counts or estimates requests over a moving interval, avoiding the sharp reset boundary of a fixed-window counter. Exact behavior depends on the implementation. You want a rolling-window policy rather than an allowance that resets at a fixed clock boundary.
Fixed-window counter Simple to implement, but requests can cluster on both sides of a window boundary, producing a burst larger than the nominal per-window figure over a short span. OWASP cautions against fixed windows in its bot-management guidance. Only where simplicity is acceptable and boundary bursts are understood and tolerable.
Leaky-bucket-style shaping Can smooth work by releasing it at a controlled pace. Shaping delays or queues excess work; a hard rejection policy instead refuses requests once its allowance is exhausted. You want to smooth outbound or queued work, provided added latency and queue capacity are managed.

A configured rate is not necessarily a perfect hard wall. AWS describes API Gateway throttles as best-effort targets rather than guaranteed request ceilings. Its HTTP API documentation explains token-bucket throttling and burst capacity (AWS: Throttle settings for HTTP APIs).

Should I rate limit by IP or API key?

Choose a key based on the abuse pattern and the fairness you need. No single key is reliable for every route or attacker.

  • IP address: Can constrain traffic from one source, including before a user signs in. But it may group many legitimate people behind a shared network address, and distributed sources can evade a per-IP limit.
  • User, account, or credential: Supports consumer-specific quotas after the identity is established. Credentials can be shared or stolen, and unauthenticated requests have no authenticated identity to count against.
  • Multiple independent controls: For login and account-recovery flows, assess separate controls for source IP and account or username. This helps address both high-volume source behavior and repeated attempts against one account; relying only on a composite IP-plus-username counter can leave gaps.

API keys can help identify or meter clients, but possession of a key does not establish that a request is safe. OWASP warns: “Do not rely exclusively on API keys to protect sensitive, critical or high-value resources.” See the OWASP REST Security Cheat Sheet and its guidance on testing for brute force.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where should different limits apply?

A single global request count can miss the endpoints that create the most risk. OWASP highlights authentication and account recovery as well as expensive search, export, and bulk operations as areas to assess. Apply policies according to route risk and the identity information available at that point.

  • Authentication and recovery: Consider independent limits by source and by account or username to address different abuse patterns.
  • Expensive operations: Use a stricter request allowance where search, export, or bulk work can consume substantial resources. Consider concurrency caps or workload limits when the number of simultaneous operations matters more than request frequency.
  • General API capacity: A broader gateway limit can help protect shared backend capacity, while narrower route or customer policies manage localized risk.

Per-request constraints matter alongside frequency controls: cap input sizes, result counts, execution time, and other workload dimensions as appropriate. A client that sends few but exceptionally costly requests can still exhaust resources.

What should I return when an API rate limit is exceeded?

Return HTTP 429 Too Many Requests when rejecting a request because it exceeded a rate limit. Document how clients should recover, while avoiding disclosure of sensitive internal enforcement details.

A Retry-After header can tell a client when to try again if the service can provide a useful delay. Header formats are not universal: Cloudflare, for example, documents Ratelimit, Ratelimit-Policy, and retry-after headers, with its Retry-After value expressed in seconds and rounded up until more capacity is available (Cloudflare API limits). Do not assume another API uses the same headers or semantics.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should clients handle 429 responses?

Clients should not immediately repeat a rejected request in a tight loop: retries add traffic while the service is already limiting requests. Honor a supplied retry delay when applicable. If throttling continues, increase the wait between attempts rather than retrying at a fixed rapid interval. AWS Well-Architected guidance recommends controlling retry calls and increasing backoff intervals after repeated failures (AWS: Limit retries).

Can an API gateway enforce rate limits?

Yes. Depending on the product and configuration, a gateway can enforce throttling at scopes such as an account, stage, route, method, or client. AWS API Gateway documents token-bucket throttling and configurable limits; its REST API documentation also describes usage plans and method-level targets. These are useful enforcement mechanisms, but AWS says throttles are best-effort targets, not guaranteed ceilings (AWS: Throttle API requests for better throughput).

Before relying on a gateway policy, check the provider’s documented scopes, burst behavior, quotas, and enforcement guarantees. A gateway may not know application-specific cost or identity details, so it may need to complement controls in the service itself.

Does a rate limit stop denial-of-service attacks?

No single API rate limiter should be treated as a complete denial-of-service defense. It can contain traffic at the point where it is enforced, but distributed sources, traffic that reaches infrastructure before the limiter, or a small number of unusually expensive requests may still threaten availability. Use rate limits as one layer alongside authentication, workload controls, and suitable upstream availability protections. OWASP notes that API keys may reduce denial-of-service impact, but should not be the sole protection for sensitive resources (OWASP REST Security Cheat Sheet).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.