Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

API rate limiting controls how many requests a defined client, route, or service can make over time. The right design matches the protected backend capacity, the burst it can absorb, and the way requests are distributed across gateway instances. Token buckets suit controlled bursts; window-based counters suit interval policies; concurrency limits address simultaneous work. None of those labels alone tells you whether excess requests will be rejected, delayed, or queued.

What an API rate limit controls

A rate-limit policy combines an allowance, a time period, and a scope or identity—for example, requests per second for a consumer on a route. It can protect an upstream service from overload, distribute capacity among clients, or enforce an aggregate boundary. A limit is not automatically a universal requests-per-second quota: its value and behavior depend on the API and the configured policy.

Rate limits are also different from longer-period quotas and concurrency limits. A quota caps use over a longer accounting period; a concurrency limit caps simultaneous in-flight operations, regardless of how quickly those operations began. Apache APISIX documents concurrency control as a separate plugin from its rate-limiting plugins in its overview of API gateway rate-limiting algorithms.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the main rate-limiting algorithms differ

Approach Burst and boundary behavior What happens to excess traffic Implementation considerations
Token bucket Allows bursts up to bucket capacity while limiting sustained demand to the refill rate. Often rejects requests without available tokens; handling is implementation-specific. Tracks token state. AWS API Gateway uses this model, but documents configured throttling as best-effort targets, not guaranteed ceilings. AWS HTTP API throttling
Leaky bucket / traffic shaping Smooths bursts toward a more regular output rate. May delay or queue requests, or reject them; the algorithm name does not settle which. Behavior and supported options depend on the gateway and plugin. APISIX maps its limit-req plugin to leaky bucket.
Fixed window Counts requests in discrete intervals; a caller can spend allowance just before a reset and again just after. Usually rejects requests after the current window’s count is exhausted. Simple interval counting, but the reset boundary can permit a brief surge. Kong’s window-type guide illustrates this effect.
Sliding window Measures usage over a moving interval, reducing the fixed-window reset-boundary effect. Usually rejects requests above the moving-window allowance; exact counting and retry behavior vary. Storage, approximation, and rejected-request accounting depend on implementation. Kong documents sliding-window support.
Concurrency limit Constrains simultaneous in-flight work rather than requests within a time interval. May reject or otherwise defer new work, depending on the product. Useful when operations occupy backend resources for different lengths of time; pair with a request-rate policy when both problems matter. APISIX documents a distinct concurrency plugin.

Token bucket: separate sustained rate from burst size

A bucket replenishes tokens at a configured rate up to a maximum capacity; each accepted request consumes tokens. The refill rate controls continuing demand, while capacity determines how much short-lived traffic can arrive together. This makes the model useful when legitimate clients need occasional bursts but sustained traffic still needs a bound. AWS API Gateway uses token-bucket throttling for HTTP APIs and cautions that its configured rate and burst values are best-effort targets: exceeding them can lead to HTTP 429 responses, but they are not hard request ceilings. See AWS’s HTTP API throttling documentation.

Leaky bucket: verify whether the gateway smooths or rejects

Traffic shaping can spread bursts over time, but products do not all implement excess work the same way. Apache APISIX describes limit-req as leaky-bucket based. Kong documents delayed-and-retried throttling as an optional capability of its advanced rate-limiting plugin, rather than an automatic consequence of choosing a bucket algorithm. Check the specific product version and configuration before promising queueing or delay. Kong’s gateway rate-limiting documentation describes its plugin options.

Fixed and sliding windows: choose how much boundary smoothing matters

A fixed window is easy to reason about when a policy is explicitly tied to discrete periods, but its reset can concentrate traffic on either side of the boundary. A sliding window reduces that particular effect by evaluating a moving interval. It does not guarantee identical results across gateways: counting, storage, treatment of rejected requests, and retry behavior remain implementation details. Kong describes both window types and their boundary behavior in its window-types guide.

Concurrency limits: protect capacity consumed by work in progress

Two operations may each count as one request but hold database connections or workers for very different durations. A request-per-time limit does not directly cap how many such operations are active at once. A concurrency cap addresses that load shape; it complements rather than replaces a rate policy when both arrival rate and in-flight work need control.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the scope before setting a number

The identity used for counting determines who shares an allowance. Common scopes include account, API key or credential, authenticated consumer, IP address, route, service, or combinations of these. IP-only policies can group unrelated people behind a shared address; unauthenticated endpoints may still need IP- or network-level controls. Kong documents consumer, credential, IP, service, and route scopes in its gateway rate-limiting guide.

Layer policies so fairness controls do not become the only protection for shared infrastructure. Per-consumer or per-route limits can constrain individual demand, while broader service or regional limits protect aggregate capacity. For API Gateway REST APIs, AWS documents usage-plan client, stage/method, account, and regional throttling layers, along with their precedence in its REST API throttling guide.

Set and deploy a policy safely

  1. Define the capacity objective. Identify the upstream resource to protect and the routes with materially different costs. Establish safe capacity using representative service measurements rather than selecting a public requests-per-second figure by convention.
  2. Select the counting key and layers. Decide whether each policy applies to a consumer, credential, IP, route, service, account, or a combination. Add an aggregate safeguard when the goal includes protecting shared backend capacity.
  3. Set sustained rate and burst independently. For a token bucket, set the refill rate for continuing demand and capacity for the allowed short-lived burst. Relate burst size to queue depth, downstream concurrency, and latency budgets.
  4. Choose state semantics for replicas. A local counter is fast and avoids coordination, but independently enforced allowances can add up across gateway replicas. Shared counters can make decisions more consistent across replicas, at the cost of coordination latency and dependence on the state store. Kong documents Redis support for its rate-limiting plugins; that does not establish a universal consistency guarantee. See Kong’s documentation.
  5. Specify behavior when limiter state is unavailable. Decide whether requests fail open or fail closed, and define timeouts, fallback limits, and monitoring. These choices are product- and architecture-specific, not standardized across gateways.
  6. Test the configured behavior. Exercise normal load, bursts, window boundaries where relevant, multiple replicas, and state-store failure. Verify whether excess requests are rejected, delayed, or queued, and whether the resulting backend latency stays within budget.
  7. Observe and tune. Monitor allowed and rejected requests, key cardinality, saturation, backend latency, and state-store health. Treat configured targets according to the gateway’s documented enforcement semantics, not as guarantees unless the product explicitly provides one.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Handle HTTP 429 and retries correctly

Use 429 Too Many Requests when rejecting a request because a rate limit has been exceeded. Include Retry-After when the server can give a meaningful retry time. Slack documents that its HTTP APIs return 429 and a Retry-After header containing seconds until retry; its example value of 30 is illustrative, not a general wait interval or universal API limit. Slack also notes that method tiers can change. See Slack’s rate-limit documentation.

  • If a response includes Retry-After, honor it rather than immediately retrying.
  • When many clients might retry together, add jitter and cap retry attempts to avoid creating another burst.
  • Before replaying a request that can cause side effects, check whether it is safe to retry and use idempotency protections where appropriate. A 429 response alone does not make every operation safe to repeat.

Why gateway-specific behavior matters

Algorithm names are useful shorthand, not a cross-vendor behavior contract. Gateways can differ in counting, reset timing, rejected-request accounting, shared-state options, failure handling, and whether excess work is queued or rejected. AWS’s HTTP API throttles are explicitly best-effort targets, while AWS REST APIs document multiple throttling scopes and precedence. Kong’s standard and advanced plugins have different supported algorithms and Redis options. APISIX’s plugin-to-algorithm mapping is specific to APISIX. Check the documentation for the exact gateway, plugin, version, and configuration you deploy rather than transferring assumptions from another product.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.