Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Voice activity detection (VAD) does not mute tab audio by itself, and it does not automatically lower speech-to-text (STT) charges. VAD makes a speech-or-silence decision. An application can use that decision to split audio into chunks for transcription, and a vendor’s billing rules then determine whether any of that saves money. If tab audio goes quiet once VAD is switched on, the cause is almost always in how the application routes, enables, or plays its media tracks. The fix is found by tracing those paths, not by tuning VAD.

Four separate layers the title combines

The apparent contradiction comes from treating four different layers as one. Keeping them apart makes the symptom easier to diagnose.

  • Speech detection. The VAD model decides whether a stretch of audio contains speech. It produces an event or a flag. It does not, on its own, change what anyone hears.
  • Audio chunking. The transcription pipeline uses that decision, or a silence rule, to decide where one segment ends and the next begins.
  • Billing. The vendor meters usage in its own unit, under its own API mode and plan. Whether fewer seconds are submitted, and whether those seconds are what gets billed, is a vendor question.
  • Media-track routing and playback. The application decides which audio track is captured, whether that track is enabled, and which element plays remote or tab audio back to the user.

A problem in the fourth layer can look like a VAD problem because the two happen at the same moment. Nothing in the first three layers requires playback to stop.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What VAD does in OpenAI’s Realtime API

OpenAI’s Realtime server VAD detects speech turns and chunks audio based on silence. In supported configurations it exposes three settings: a threshold, prefix padding, and silence duration. The threshold sets how loud audio must be to count as speech, prefix padding keeps a little audio before detected speech starts, and silence duration sets how long a pause must last before a turn is considered finished.

#1 Best Overall
MOSWAG USB to 3.5mm Jack Audio Adapter External Stereo Sound Card
  • 【USB external sound card audio adapter】This USB to aux adapter supports listening and speaking,Easily adds a 3.5mm TRRS aux port integrated microphone-in and audio out interface to your devices
  • 【High Quality Sound】 Equipped with an advanced built-in DAC chip, this USB sound card supports both CTIA and OMIP standard headphones. This USB to Aux adapter delivers stable 16-bit/48kHz audio output and effective noise reduction, faithfully reproducing and enhancing the original sound quality. Note: The 3.5mm male microphone jack does not support TS or TRS connectors
  • 【Wide Compatibility】USB to 3.5mm Jack Audio Adapter support TRRS headsets and microphones.USB male wide compatibility with Windows 10/9/8/7/Vista/XP,Linux,Mac OS X google Chromebook,Raspberry Pi, PS4,PS5 and Windows Surface 3 etc
  • 【Plug and Play】USB Sound Adapter no driver required,USB headset adapter plug and play;the durable nylon braided cable of the USB audio adapter ensures stable transmission and allows you to use your 3.5mm headphones more conveniently.USB to 3.5 mm port will be automatically recognized by system in seconds
  • 【Portable and Durable】USB to audio jack adapter is equipped with an aluminum shell.The nylon braided of the USB to 3.5mm jack audio adapter is more durable,smaller and lighter than other plastic shells and PVC cable USB audio adapter,ensuring a much longer lasting life

OpenAI’s Realtime VAD guide draws the boundary directly: “In transcription mode, VAD only controls how audio is chunked.” That sentence describes what VAD is responsible for in transcription sessions. It says nothing about muting, volume, or playback. The guide is at https://developers.openai.com/api/docs/guides/realtime-vad.

Support also depends on the model. According to the same guide, gpt-live-transcribe and gpt-realtime-whisper require turn detection to be omitted or set to null, with turns finished by sending input_audio_buffer.commit. Check the current model-specific documentation before implementing anything, because these requirements can change.

The protocol-level picture is similar. RFC 6464 defines a voice activity bit that signals audio-level information in a header extension. The RFC states that this “vad extension attribute only controls the semantics of this header extension attribute, and does not make any statement about whether the sender is using any other voice activity detection features, such as discontinuous transmission, comfort noise, or silence suppression.” In other words, detecting voice and suppressing audio are different things, even in a standard that defines the detection signal. The RFC is at https://www.rfc-editor.org/rfc/rfc6464.html.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
SABRENT USB External Stereo Sound Card Adapter, Plug & Play (AU-MMSA)
  • PLUG IN AND HEAR SOUND IN SECONDS - USB Type-A connector with a 3.5mm stereo headphone output and a separate 3.5mm mono microphone input. No drivers, no software, no external power - the adapter is USB bus-powered and is recognized as a standard USB audio device.
  • WORKS ON WINDOWS, MAC AND LINUX - Driverless on Windows 98SE/ME/2000/XP/Server 2003/Vista/7/8, Linux and Mac OSX, and compliant with the USB Audio Device Class 1.0 specification, so any system that supports class-compliant USB audio will see it. Select it as the sound output and input device after plugging it in.
  • TWO JACKS, TWO JOBS - The green jack is stereo OUT for headphones or powered speakers; the pink jack is mono microphone IN for a 3.5mm mic. It does NOT support 4-pole headsets on a single combo plug, it does NOT power passive speakers, and it does NOT add surround sound - it is a stereo 2-channel adapter.
  • FOR LAPTOPS AND DESKTOPS THAT NEED AN AUDIO PORT BACK - Adds a headphone and mic port to a laptop, desktop, or mini PC whose onboard jack has failed or was never there. Managed and work-issued computers can block new USB audio devices by policy - check with your IT department before ordering for a company machine.
  • SABRENT SUPPORT AND WARRANTY - What is in the box: one USB audio sound adapter. Backed by a 1-year limited warranty, extended to 2 years when you register within 90 days on the manufacturer's website.

Why chunking does not prove lower billed minutes

File transcription has its own chunking control. In OpenAI’s Audio API transcription reference, the chunking_strategy parameter accepts "auto" or a manual server_vad configuration. If chunking is not set, the endpoint transcribes the audio as a single block.

Chunk boundaries describe how audio is segmented. They do not show how many seconds are billed. The table below separates what each option controls from what it does not establish.

Option What it controls What it does not establish
chunking_strategy: "auto" Segmentation chosen automatically by the endpoint Any billed-duration reduction. Segment count is not a savings measure.
Manual server_vad configuration Boundaries set by silence-based rules you choose Lower charges. Boundary control is not billing evidence.
Chunking unset Audio transcribed as a single block Whether a single block costs more or less than chunked input. Not stated in the reference.

The phrase “saves STT billed minutes” is therefore conditional. A VAD front end can avoid submitting nonspeech intervals in some designs, and API VAD can split an input into chunks. Neither fact establishes a universal billing formula. Actual charges depend on the vendor, the API mode, what audio is submitted, and the vendor’s current billing unit. OpenAI’s pages cited here give no savings percentage and no minutes-saved figure, so any cost outcome has to be checked against your own vendor and plan before you rely on it.

Rank #3
Sale
UGREEN UGREEN USB Sound Card Adapter with TRS 3.5mm Headphone & Mics
  • Upgrade the Sound Quality: UGREEN Aux to USB adapter is the perfect solution for upgrading the sound quality of your laptop or desktop computer. With its high-resolution DAC chip, this adapter offers stunning audio quality that will completely transform your listening experience
  • Crystal-Clear Sound: Experience high-fidelity audio like never before! With a built-in DAC chip, this USB audio adapter delivers rich and immersive audio. The USB Aux adapter facilitates high-resolution audio output and noise reduction up to 16bit/48kHz to enhance the original sound quality of your devices
  • Plug and Play: Simply connect this sound card to your device and you're ready to go - no drivers or external power sources required. Whether you're using it for gaming, recording music, or watching movies, this adapter is sure to impress
  • Wide Compatibility: The USB to audio jack is Compatible with Windows 11/10/98SE/ME/2000/XP/Server 2003/Vista/7/8/Linux/Mac OSX/PS5/PS4/Google Chromebook/Windows Surface Pro 3/Raspberry Pi. So no matter what you're using, this adapter is sure to work seamlessly with your setup. (*Note: NOT compatible with PS3.)
  • Compact and Portable: UGREEN Aux to USB adapter is constructed with durable ABS material that makes it easy to take on the go. Don't miss out on this opportunity to elevate your audio experience - get your hands on the UGREEN Aux to USB adapter today

Comparing the two Realtime VAD modes

Two VAD approaches are documented for supported OpenAI Realtime sessions. They differ in how they decide that a speaker has finished.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Attribute server_vad semantic_vad
Basis for ending a turn Periods of silence Estimated turn completion from the words spoken
Interruption and chunking behavior Chunks at silence boundaries, so short pauses can split an utterance Per the guide, may wait longer when an utterance trails off, which can reduce premature chunking
Configuration Threshold, prefix padding, and silence duration Not stated in the guide’s settings list reviewed for this article
Latency trade-off Shorter silence settings respond faster but risk cutting speakers off Gives the speaker more time to finish, at the cost of a longer wait before a turn closes
Effect on browser playback Not stated. The guide describes chunking only. Not stated. The guide describes chunking only.

Neither mode is documented as affecting playback. If a change in VAD mode seems to coincide with lost tab audio, the change is a timing cue for where to look, not a mechanism.

Input and playback are separate paths

OpenAI’s WebRTC guide for browsers shows the usual structure. Microphone input travels on a negotiated media track, while remote generated speech is attached to an audio element for playback. Capture and playback therefore have distinct paths, and each can be inspected on its own.

Rank #4
USB to 3.5mm Jack Audio Adapter (2-Pack), External Sound Card Converter,STSAUX USB-A to 3.5mm TRRS 4-Pole Mic Female with PC Windows, Laptop, Mac, Desktops, Linux, Plug and Play No Drivers Needed
  • 【STSAUX USB to 3.5mm Audio Adapter with Advanced DAC Chip】 - Experience crystal clear stereo sound and microphone functionality through an external USB sound card. Equipped with a high-performance DAC smart chip, it delivers 24-bit/96KHz high-resolution audio output with enhanced noise reduction, providing immersive experiences for gaming, music, and communication. Serves as an ideal replacement for damaged audio ports on laptops and desktop computers.
  • 【Plug and Play Instant Setup】 - No drivers or software required! The STSAUX USB-A to 3.5mm TRRS adapter automatically recognizes devices within seconds. Fully compatible with Windows 10/11, Mac OS, Linux, Chrome OS, , and Surface Pro. Important note: Not compatible with TVs, cars, or iOS microphone functionality (audio output only).
  • 【Full Compatibility with TRRS CTIA Standards】 - Connect any 3.5mm headphones, microphones, or speakers to USB-enabled devices. Supports CTIA/OMTP standards and works with Android headsets, gaming headphones, and PC speakers. Explicitly incompatible with due to USB audio output limitations.
  • 【Dual-function Microphone + Audio Support】 - Unlike basic audio-only adapters, STSAUX design enables simultaneous listening and speaking through a single TRRS 4-pole interface. Essential for Discord, Zoom, in-game voice chat.
  • 【18-month Warranty + Lifetime Technical Support】 - Each package includes 2 adapters backed by an 18-month warranty. STSAUX provides 24/7 customer service and troubleshooting. Ideal as backup solutions for home office setups, gaming stations, or tech toolkits.

That example illustrates the separation but does not describe how any particular application routes captured tab audio. Your application may use a different capture method, a mixing step, or a virtual audio device. Those are the places to verify.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Other platforms gate work on speech presence

Apple’s SpeechDetector is a documented example of gating transcription on speech presence, so the system does not spend power trying to transcribe likely silence. It shows that speech-presence gating is a common design. It does not show billing savings for any STT service, and it is not evidence that the gate mutes other audio on the device.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Diagnosing missing tab audio after enabling VAD

Work through these checks in order. Each is a diagnostic step, not a confirmed cause.

Best Value
UGREEN USB to 3.5mm Jack Audio Adapter Headphone DAC 24bit 96kHz
  • 2 in 1 USB Audio Adapter: UGREEN USB to 3.5mm jack audio adapter supports both listening and talking. In-line control and microphone function allows you to chat while gaming. This USB Aux adapter enables you to connect headphones, headset, speakers with 3.5mm jack to your PC through the USB A port. [*Note: The 3.5mm adapter can support the connection of a single microphone for use. Please note that the 3.5mm male microphone's attribute requirement is TRRS]
  • 24bit 96KHz High Quality Sound: Built-in advanced smart chip, USB sound card supports CTIA and OMIP standard headsets. The USB to Aux adapter facilitates high-resolution audio output and noise reduction up to 24bit/96kHz (others are 16bit/48kHz)to enhance the original sound quality of your PC devices. Note: Microphone does not support 24bit 96KHz, only headphones can
  • Built to Last: With aluminium alloy shell, the Aux to USB adapter is more durable and portable than other plastic USB audio adapter. Nylon braided design of USB to audio jack adapter ensures stable audio transmission. 10000+ bend tests prove that USB headset adapter can prevent damage caused by tangling and scratching
  • Broad Compatibility: This USB external sound card is compatible with Windows 11/10/98SE/ME/2000/XP/Server 2003/Vista/7/8, Linux//Raspberry Pi, MacBook Pro 2019, MacBook Air 2018, PS5/ PS4, Switch 2, Google Chromebook, Windows Surface Pro 3. *Note: Does not support Apple headphone line control
  • Plug and Play: No driver is required, just plug and play. USB bus-powered, no external power is required for this convenient sound card. *Note:Supports regular headphones with a resistance of less than 50 ohms, but does not support high-resistance headphones
  1. Identify the stream that feeds transcription. In your code, find the media track passed to the transcription connection. Confirm whether it is the tab-audio track or a separate microphone track.
  2. Determine what the VAD output does. Search your handlers for VAD events. If they only record speech or silence, VAD is not gating playback. If any handler changes a track or element, that handler is the candidate.
  3. Check whether the track’s state changes when VAD fires. Log track.enabled and track.readyState at each VAD event. A track that flips to enabled = false, or whose source is replaced, would explain silence on the playback side.
  4. Confirm the playback element still has the intended stream. In browser developer tools, inspect the audio element and check srcObject, muted, paused, and volume. Confirm the element is attached to the stream you expect.
  5. Look for a separate processing or routing step. Audio processing, noise suppression, mixing, or a virtual audio device between capture and playback can remove or mute tab audio. Disable one step at a time to see which one restores audio.

If step 3 or 4 shows no change, the VAD settings are probably not the problem, and changing threshold or silence duration will only move chunk boundaries. If a handler changes track state, fix that handler so playback is decoupled from the transcription decision.

When a service comparison is needed, verify the current terms and billing unit directly with the vendor before drawing conclusions about cost.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.