Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no meaningful single ranking of speech-to-text API rate limits: OpenAI publishes model-tier RPM and TPM, Google Cloud publishes project- and region-scoped request limits, Azure Speech sets resource-scoped quotas, and Amazon Transcribe separates per-operation TPS from concurrency. Compare the limit that matches your workload—short file requests, batch submissions, or live streams—and verify the quota for your own account and region before sizing capacity.

What do RPM, TPM, TPS, and concurrency mean?

  • RPM means requests per minute. It limits how many API requests can be sent during a minute, not how much audio those requests contain.
  • TPM means tokens per minute. It is a model usage measure, not a count of transcription jobs or audio minutes.
  • TPS means transactions per second for a named API operation.
  • Concurrency is the number of requests, jobs, or sessions that can be active at the same time. A concurrent stream can stay open, while a submitted batch job may process a much longer recording.

These measures describe different constraints. None, on its own, tells you how many minutes of audio a service will transcribe per minute; that also depends on the request type and content limits.

What are the published limits for each provider?

The figures below are documentation values, not universal capacity guarantees. Their scopes differ, so compare only matching workload modes and quota units. Provider limits can change; check the live quota page, console, or support channel for your own configuration.

Provider and workload Published request or usage limit Published concurrency Scope and qualification
OpenAI GPT-Transcribe Tier 1: 500 RPM and 200,000 TPM
Tier 2: 5,000 RPM and 2,000,000 TPM
Tier 3: 5,000 RPM and 4,000,000 TPM
Tier 4: 10,000 RPM and 10,000,000 TPM
Tier 5: 30,000 RPM and 150,000,000 TPM
No concurrency figure is listed in the model limits table. Model limits depend on usage tier; Free is unsupported for this model. OpenAI says tiers increase automatically as requests and spend increase. These figures do not establish audio-job throughput. See OpenAI’s GPT-Transcribe documentation.
Google Cloud Speech-to-Text synchronous recognition 300 requests per 60 seconds per region Not stated for this request class. Limits apply per developer project and are shared by applications and IP addresses using that project.
Google Cloud Speech-to-Text batch recognition 150 requests per 60 seconds per region Not stated for this request class. Limits apply per developer project and are shared by applications and IP addresses using that project.
Google Cloud Speech-to-Text resource and operation requests 100 resource requests per 60 seconds per region; 150 operation requests per 60 seconds per region Not stated for these request classes. Project-scoped limits shared across applications and IP addresses. These request categories are distinct from the streaming session concurrency cap.
Google Cloud Speech-to-Text streaming 3,000 requests per minute shared across streaming sessions; initial session configuration does not count toward this request quota 300 concurrent sessions Per developer project and region; apps and IPs using a project share its quota. Google’s quota page says values may change and was last updated 2026-09-30 UTC. See Google Cloud’s quota documentation.
Azure Speech real-time speech-to-text, Standard S0 No RPM figure stated here for real-time requests. Default 100 for the base model endpoint and 100 for a custom endpoint Per Speech resource. Real-time speech-to-text and speech translation concurrency are combined. The existing concurrency value is not visible in the portal, CLI, or API; contact support to verify it.
Azure Speech real-time speech-to-text, Free F0 No RPM figure stated here. 1 Per Speech resource. See Microsoft’s Azure Speech quotas and limits for current quota details.
Azure Speech fast and batch transcription, Standard S0 600 requests per minute shared by fast transcription and batch transcription Not stated as a separate concurrency value here. Per Speech resource. Azure says the shared fast/batch rate can be adjusted; other batch constraints are not adjustable.
Amazon Transcribe job submission: StartTranscriptionJob 25 transactions per second 250 concurrent transcription jobs TPS applies in each supported Region; the concurrent-job quota is separate. Check Service Quotas for the relevant AWS account and region to confirm whether the quota can be adjusted.
Amazon Transcribe streaming: StartStreamTranscription 25 transactions per second 25 concurrent HTTP/2 and WebSocket streams TPS applies in each supported Region; stream concurrency is a separate quota. Check Service Quotas for the relevant AWS account and region. See AWS’s Amazon Transcribe endpoints and quotas.

Why request limits are not the same as audio capacity

A request quota governs submissions or API operations; concurrency governs active work. A batch submission might represent one request even when its audio is long, whereas a real-time stream occupies a session while it remains open. Payload and duration limits are a separate constraint again.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
JOUNIVO USB Microphone, 360 Degree Adjustable Gooseneck Design, Mute Button & LED Indicator, Noise-Canceling Technology, Plug & Play, Compatible with Windows & MacOS
  • 360 Degree Position Adjustable Gooseneck Design --Plug and play USB microphone Pick up the sound from 360-degree with high sensitivity, in the best possible location for sound to your PC gaming, dragon voice dictation, and talk to Cortana
  • Mute Button & LED Indicator --One-click to mute/unmute your microphone for pc, Build-in LED indicator tells you the working status at any time
  • Intelligent Noise-Canceling Tech --Premium omnidirectional condenser microphone with noise-canceling technology can pick up your clear voice and reduce background noise and echo
  • USB Plug&Play(1.8/6ft USB Cable) -- No driver required. Just need to plug & play for the microphone to start recording, well compatible with Windows(7, 8, 10 and 11) and macOS. (NOT compatible with Xbox/Raspberry Pi/Android)
  • Solid Construction--Adopting premium metal pipe and heavy-duty ABS stand to make sure that you will be satisfied with our computer mic quality

Google’s documented content limits illustrate the distinction: synchronous recognition accepts up to 10 MB or one minute of audio; a streaming session can remain open for five minutes with audio sent near real time; and batch recognition accepts up to five files per request, with each file up to eight hours. These request-shape limits do not replace Google’s request-rate or concurrent-session quotas. See Google Cloud’s quota and limit documentation.

How should you size a transcription workload?

  1. Choose the workload mode. Separate short synchronous calls, batch job submissions, and live streams. Do not use a streaming concurrency figure to estimate batch submission capacity, or vice versa.
  2. Identify the exact quota scope. Check whether the value applies to a model tier, developer project, Speech resource, AWS account, or supported region. Shared project or resource limits may be consumed by more than one application.
  3. Check the operation and unit. Match the request to the specific model, endpoint, or API operation. Treat RPM, TPM, TPS, and active sessions or jobs as separate ceilings.
  4. Verify content constraints. Check file size, duration, files-per-request, and stream-duration limits as well as the rate quota. A workload can fit under the request cap and still exceed a content limit.
  5. Confirm the live quota and adjustability. Consult the provider’s quota page or account console for the project, resource, account, and region you will use. For Azure concurrency, contact support because the existing value is not exposed in the portal, CLI, or API. For AWS, inspect the relevant Service Quotas entry.
  6. Plan for bursts and shared usage. Keep a queue for work that cannot start immediately, limit concurrent workers to the applicable cap, and use backoff when a request is throttled. Validate the plan with a representative workload; documentation quotas alone do not establish sustained audio throughput.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Can these quotas identify the fastest or most accurate provider?

No. The published limits describe operational ceilings, not a controlled comparison of transcription speed or accuracy. A higher RPM, TPS, or concurrency number alone does not prove that a provider processes more audio per minute or produces better transcripts.

Best Value
Sound Tech GN-USB-2 18 Inch Professional Uni-Direction Noise Canceling Gooseneck Stereo Microphone with 10 FT USB Cord
  • The GN-USB-2 gooseneck is specially designed for professional voice communications. The GN-USB-2 is compatible for applications such as Hands-free dictation, PC recording software, voice recognition and internet chat.
  • Features: Plug n Play, Noise cancelling, On/Off LED indicator, Detachable USB A~B cable, 16 inch adjustable neck, Weight base with non-skid rubber mounts
  • Specifications: Element: fixed-charge back plate, permanently polarized condenser, Polar Pattern: Hypercardioid, Sensitivity: -40 +/- 2dB(0dB=1V/Pa at 1KHz), Frequency Response: 40Hz~16KHz, Output Impedance: 75-Ohm +/- 30% Max Input S.P.L.: 138dB, Signal/Noise Ratio: 65dB, Output Connector: USB A~B. Power Supply: Phantom Power 3V DC
  • Operating Systems: Microsoft Windows 2000, Windows XP, Windows 7 and Windows 8 , Apple Mac Os9 and all OX X variations
Rank #4
Sale
Philips SpeechMike Premium Touch Dictation USB Microphone, Push-Button
  • Microphone grille with optimized structure
  • Integrated pop filter
  • International products have separate terms, are sold from abroad and may differ from local products, including fit, age ratings, and language of product, labeling or instructions.
Rank #3
Sale
FIFINE K669B USB Microphone, Condenser Recording Mic for Vocals, Meeting
  • [Convenient Setup] Plug and play recording USB microphone for PC, with 5.9-Foot USB cable included for computer PC laptop, is connected directly to USB-A port for recording music, computer singing or podcast. The office condenser microphone for computer is easy to use and install. (NOT compatible with Xbox and Phones)
  • [Durable Metal Design] Solid sturdy metal construction design, the computer microphone for Zoom meetings with stable tripod stand is convenient when you are doing voice overs or livestreams on YouTube. Durable material extends the service life of the voice-over microphone.
  • [Mic Volume Knob] Gaming condenser USB mic compatible for PS4 with additional volume knob itself has a louder or quieter adjustment and is more sensitive. Your voice would be heard well enough through the zoom microphone USB when gaming, skyping or voice recording. Also, you can adjust your volume to zero and protect your privacy.
  • [Widely Use] USB-powered design, the condenser microphone for recording no need the 48v Phantom power supply, works well with Cortana, Discord, voice chat and voice recognition. The podcast microphone for Mac, with USB-B to USB-A/C cable, is compatible with desktop, laptop or PS4/PS5, which meets most of your daily recording needs.
  • [Clear Output Voice] Cardioid condenser microphone for PC captures your voice properly, producing clear smooth and crisp sound. Great computer recording mic for gamers/streamers/youtubers focus on the main source and reduces background noise. The streaming microphone does the job well for broadcast ,OBS and teamspeak.
Rank #2
TKGOU USB Microphone, 360 Degree Adjustable Gooseneck Design
  • 【HIGH DEFINITION AUDIO 】 This microphone embeds a patented audio filter in order to record only your voice. Good for home studio, Chatting, Skype,Discord, Yahoo Recording, YouTube Recording, Google Voice Search and Steam.
  • 【PLUG & PLAY 】 You just need to plug the microphone and it will work ! No software to install. A single button to turn it on or off. Compatible with every operating system - Mac OS X Windows Linux - and every PC brand.
  • 【SMOOTH AND CLEAR】 Noise cancellation and isolates the main sound source, This USB Microphone is perfect for videoconferencing, Skype, dictation or voice recognition. The audio filter will give you a clear and confident voice. Anti-pop filter included !
  • 【MUTE BUTTON & LED INDICATOR 】One click to mute/unmute your microphone,Build-in LED indicator tells you the working status at any time.Built with a mix of metal and heavy duty plastic, it's solid as a tank. It is very stable thanks to its weight.360 Degree Position Adjustable Gooseneck Design --Adopting the design of metal gooseneck pipe pickup the sound from 360-degree with high sensitivity
  • 【SATISFACTORY SERIVCE】- 30 days unconditional return. TKGOU Customer service 2 years, We are committed to ensuring that you are 100% satisfied, If you have any questions, please contact us directly.We will provide you with a more friendly and satisfactory service.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.