A reliable Node.js speech-to-text client should not retry every HTTP 429 the same way. First identify what limit was hit, then bound retries, control how work enters the queue, and make delays and queue health visible. The limits and queue behavior differ by provider, API generation, region, and request mode.
1. Classify the 429 before retrying
A 429 response is a signal to inspect, not a diagnosis. Read the provider’s error code and response body, and distinguish temporary throttling from a limit that waiting cannot fix.
- Temporary rate limiting: The provider may accept a later attempt. Apply its documented retry guidance.
- Credits or spend limit: Restore eligible credits or adjust the account’s usage limit; a retry loop will not resolve the account condition.
- Concurrency limit: Reduce active requests or streams and admit new work only as capacity becomes available.
- Hard session or request limit: Change how the audio is divided or which request mode is used. Retrying the same completed or oversized operation cannot remove the limit.
OpenAI’s rate-limit guidance says 429s can represent temporary rate limiting, exhausted prepaid credit, or a spend/usage limit. Inspect the error details before deciding to retry. OpenAI’s 429 troubleshooting guidance also recommends honoring a valid Retry-After; if it is absent or invalid, use exponential backoff with jitter and cap both attempts and total retry time. Unsuccessful attempts count toward per-minute limits.
For Amazon Transcribe streaming, LimitExceededException is an HTTP 429 and commonly indicates that the concurrent-stream quota was exceeded. AWS also identifies maximum session duration and a rapid increase in concurrency as possible causes. Check the specific failure before treating it as a transient stream-capacity error. See the streaming API reference and streaming guide.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
2. Check the quota’s scope and unit
Before setting a queue size or retry policy, record which service and exact API path the client uses. A quota might count requests, audio-processing volume, concurrent streams, payload size, or session duration. Those units are not interchangeable.
For each provider integration, capture:
- Provider and API generation or version.
- Project and region, where applicable.
- Request mode: synchronous, long-running/asynchronous, or streaming.
- The constrained unit, such as requests per interval, processing hours per day, active streams, audio length, or payload size.
- Whether the limit is an account quota, a project quota, a regional limit, or a hard per-request/session limit.
Google Cloud Speech-to-Text illustrates why the API generation matters. Its v1 quota page lists 900 recognition requests per 60 seconds and 480 hours of audio processing per day, with quotas shared by applications and IP addresses using a developer project. Those figures are specific to the v1 page, were marked last updated 2026-09-30 UTC, and can change; verify the current project quota before relying on them. Google’s current quota page describes separate regional limits and request modes, so do not carry v1 values over to a newer API generation without checking.
Rank #2
Google documents synchronous recognition for audio of one minute or less, asynchronous/long-running recognition for audio up to 480 minutes in its overview, and streaming recognition for real-time audio. The current quota page has mode-specific size, session, and request limits. Confirm the relevant API generation and region rather than treating the overview’s audio lengths as a complete quota specification. See the v1 request modes and Speech-to-Text overview.
3. Make retry timing bounded and non-amplifying
Use this order for a retry decision:
- Inspect the error body and status to confirm that the failure is plausibly transient.
- If the provider returns a valid
Retry-After, wait at least as long as it specifies. - Otherwise, use exponential backoff with random jitter so clients do not all retry together.
- Set a maximum attempt count and a total elapsed-time budget. When either is exhausted, surface a terminal failure or move the job to an explicitly managed later-retry state.
OpenAI’s official SDKs already retry eligible errors and honor Retry-After. Check the SDK and its configuration before adding an application retry loop: nested retries can multiply attempts and latency. OpenAI also notes that unsuccessful attempts count toward per-minute limits, so aggressive retries can worsen throttling.
Rank #3
Google Cloud’s SLA specifies a back-off requirement for its SLA context: at least one second after the first error, growing exponentially up to 32 seconds. That is contractual language for that SLA context, not a universal rule for every provider or every request. See the Google Cloud Speech-to-Text SLA.
4. Control admission, concurrency, and queue semantics
An application-managed queue should prevent a burst of waiting jobs from becoming a burst of simultaneous retries. Limit active work to a capacity your client can defend, release capacity when an operation finishes, and schedule retries independently rather than immediately promoting every delayed job.
Rank #4
Provider-managed queueing is a separate feature with provider-specific behavior. Amazon Transcribe’s optional job queue defers jobs that exceed the concurrent processing limit and processes them FIFO. AWS documents a maximum of 10,000 queued jobs and a default queue processing bandwidth ratio of 0.9; the documented defaults may be increased on request. These are AWS job-queue settings, not guarantees for a Node.js queue or other transcription providers. Details are in Amazon Transcribe job queueing.
For any application queue, decide explicitly what happens when a job cannot proceed: stays queued, is retried after a delay, is rejected to the caller, or is dropped after a defined expiry. Also consider duplicate submissions: a client timeout may leave uncertainty about whether the provider accepted the original request. Do not assume a universal idempotency guarantee across speech-to-text APIs; use provider-specific documented behavior and retain a stable application job identifier for reconciliation.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute5. Make retries and queue health observable
These are recommended application metrics, not fields mandated by the providers. Emit enough context to tell whether a job is waiting, actively consuming capacity, retrying, or finished:
- Provider, API generation, region/project where relevant, and request mode.
- HTTP status and provider error code, plus attempt number.
- Chosen retry delay and whether it came from
Retry-Afteror local backoff. - Retry-budget exhaustion and final job outcome.
- Queue depth, oldest-job age, and active concurrency.
- Enqueue-to-completion time, including the terminal result.
Log response metadata and error codes, but never API credentials or sensitive audio or transcript content. Pair metrics with alerts for sustained growth in queue depth or oldest-job age, repeated retry-budget exhaustion, and concurrency remaining pinned near its configured ceiling. Those signals help distinguish a short throttling episode from a stuck worker pool or a quota that needs to be addressed.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

