Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a Gemini Live voice agent as a streaming loop: capture microphone audio, send small PCM chunks through a persistent bidirectional session, play returned audio as it arrives, and handle interruptions and function calls in your application. Google’s GenAI SDK is a practical starting point for a server-side prototype; a browser that connects directly in production should use ephemeral tokens rather than expose a long-lived API key.

Choose how your application will connect

The Gemini Live API maintains a stateful WebSocket session. It accepts audio, video, and text inputs and can return native audio, text, or function-call requests. Choose the connection path based on where you want media handling, credentials, and application logic to live.

Approach What you manage When it fits
Google GenAI SDK The SDK wraps the WebSocket in an asynchronous interface. Your application still captures and prepares audio, consumes events, plays audio, and handles tools. A first server-side implementation or a custom agent where SDK convenience is useful.
Direct WebSocket You manage the protocol messages, session setup, event parsing, media lifecycle, and credentials. A client or service that needs direct protocol control and is prepared to own the connection lifecycle.
Agent framework integration The framework provides its own abstractions and lifecycle; verify its current Live API support and terms. An application that already uses a real-time framework or needs a broader audio, video, or telephony stack.

Google’s overview points agent developers to the Agent Development Kit Streaming route and lists integrations including LiveKit Agents, Pipecat by Daily, Fishjam by Software Mansion, Vision Agents by Stream, Voximplant, Agora, and Firebase AI SDK. Treat these as options to evaluate, not as a guarantee of identical features across frameworks. Google says client-to-server media connections can avoid an extra proxy hop; a backend-mediated design, meanwhile, can keep credentials and tool execution within application-controlled infrastructure.

Verify the model before you build around it

Google’s capabilities guide, checked on 2026-10-05, recommends gemini-3.8-live as its default for most low-latency voice-agent experiences and gemini-3.8-live-extended-thinking when more background reasoning is needed. It describes Gemini 3.1 Flash Live Preview as legacy. Model names, availability, and preview labels can change, so confirm them in Google’s current documentation before shipping or copying a model ID into configuration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
AIRHUG USB Microphone No Speaker, Desktop Computer Mic for Laptop
  • Without Built in Speaker- Please note that AIRHUG 21 microphone for pc does not have a speaker function. Built in an excellent 360° omnidirectional microphone pick up your voice within adius 6 ft. You don't have to loudly speak up to the computer or laptop
  • Be Hear Your Clear Voice - With an advanced AIRHUG noise-canceling technology, better than traditional microphone technology. The sampling rate of the pc microphone is 48k hz. When at the online calls, the other side hear your clear and real voice
  • AI Noise Reduction Mode - AIRHUG 21 USB microphone is with AI Noise Reduction Mode,eliminating background noise such as fans noise, keyboard clicks, and general background noise.Provide clear and crisp online calls for you.Great for your online learning,podcasting,conferencing and gaming. For a natural, realistic sound that captures your true voice with high fidelity, we recommend switching to Original Mode (Green Light)
  • Smart Memory& Mute Function& LED Indicator - Every restart, the computer microphone starts in recording mode (not muted), so you never miss sound by accident. It also remembers your last sound mode (noise reduction or original). No need to adjust every time. Every recording starts the way you like, easy and simple. You can direct operate mute mode for this pc microphone. The built-in indicator light of mic informs the status(Blue: AI Noise Reduction; Green: Original Mode; Red: Muted)
  • Widely Compatible Feature - AIRHUG 21 external microphone for laptop is great for small conference with 1-3 participants. The conference microphone is compatible with Zoom,Skype,Microsoft,Teams,Google meeting,Webex,Facetime, and most of the online meeting apps. It is a great choice for anyone who needs to make video meeting, online education,seminars, remote training, business negotiations,etc

Keep the trade-off explicit: a voice agent optimized for a lean interaction loop and one asked to do more background reasoning may call for different model choices. Do not assume one model or tool capability applies to every Live model.

Build the streaming audio loop

Google documents raw little-endian 16-bit PCM at 16 kHz for native audio input and raw 16-bit PCM at 24 kHz for output. For an input blob, include the sample rate in its MIME type, such as audio/pcm;rate=16000. Many microphones provide audio at another rate, commonly 44.1 or 48 kHz; resample it before sending rather than labeling samples with a rate they do not have. A particular microphone is not required—the important part is the audio stream your application sends.

Rank #2
TONOR Conference USB Microphone with AI Noise Canceling for PC, G11 Pro
  • Built-in AI Noise Reduction: Compared to the base model, G11 pro upgraded AI noise cancellation, effectively eliminates distractions like fan noise, keyboard clicks. It delivers clear, crisp teleconferencing experiences, making it perfect for conference calls, online learning and chatting
  • Omnidirectional Conference Mic: Features omnidirectional pickup pattern with a pickup distance of 11.5 ft, making it easy to capture sounds from 360° directions. Highly sensitive pickup ensures participants hear everything clearly. Tips: This is not a speaker
  • Effortless Control: Physical volume and monitoring control buttons are built into the microphone body, allowing you to effortlessly adjust both microphone and monitoring volume. Click to adjust volume between 4 levels
  • Mute & Monitor: Quickly mute/unmute your microphone by one tap. Built-in 3.5mm jack allows connection of headphones for monitoring. Long press for 3 seconds to enable/disable: Blue-Mic mode, Red-Mute, Purple-Monitoring. Note: Do not connect the 3.5mm jack to external speakers, as this may cause feedback interference
  • Plug & Play: Compatible with all operating systems,both Windows and macOS. No additional drivers needed . If there is no response after inserting the mic, please go to the microphone setting of your computer and select the mic as the INPUT device
  1. Open a Live session. With the GenAI SDK, the documented connection entry points are Python client.aio.live.connect(...) and JavaScript ai.live.connect(...). Configure audio as a response modality, and supply system instructions, voice configuration, and function declarations if the agent needs them. With a raw WebSocket, send the setup message first; it configures the model and session options.
  2. Capture and normalize microphone audio. Convert captured audio to raw little-endian 16-bit PCM at 16 kHz. Google says the API can resample input, but its best-practices guide recommends application-side resampling of typical microphone input.
  3. Send short chunks continuously. Google recommends chunks of 20–40 milliseconds and cautions against waiting to buffer a full second of microphone audio. Smaller regular chunks help avoid adding unnecessary buffering delay.
  4. Receive model output while the turn is running. Keep a receive loop active and inspect incoming server content for audio chunks, text, completion, interruption, and function-call events. Do not wait for an entire long response before starting playback.
  5. Play the output stream. Feed each returned audio chunk into a playback path configured for the documented 24 kHz raw PCM output. Keep playback buffering under application control so it can be stopped promptly.
  6. Flush a paused input stream when needed. For continuous audio, voice activity detection is enabled by default. If microphone input pauses for more than about a second, Google’s capabilities guide says to send an audioStreamEnd event to flush cached audio.

The SDK guide demonstrates the asynchronous connection, realtime input methods, and receive loop; the raw WebSocket reference describes the underlying setup-and-message protocol. In a raw WebSocket implementation, configuration is not generally changeable while the connection is open, apart from supported pause/resume mechanisms, so choose session configuration before connecting.

Handle interruption and turn completion

Voice activity detection lets the user interrupt naturally, but canceling generation on the server does not remove audio already queued in the client. When an incoming server event reports serverContent.interrupted as true, stop current playback and clear locally queued audio. Otherwise, the agent may keep speaking from its local buffer after the server has canceled its turn.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Movo WebMic USB Microphone for AI Coding, Voice Prompts & Dictation
  • BUILT FOR DICTATION & VIBE CODING – Talk to your AI assistant, dictate code, or draft documents by voice. The Movo WebMic's clear, close-up capture means fewer transcription errors so your words land right the first time.
  • CARDIOID PICKUP FOR CLEAN VOICE-TO-TEXT – The directional cardioid capsule focuses on your voice and rejects noise from behind, giving speech-to-text engines and AI prompts the clean input they need to stay accurate.
  • HANDS-ON CONTROLS, ONE-TOUCH MUTE – Built-in knobs adjust mic gain and headphone monitoring level, a 3.5mm headphone jack lets you hear yourself live, and one-touch mute keeps you in control during calls and long coding sessions.
  • PLUG AND PLAY ON PC & MAC – Connect over USB with no drivers or extra hardware. Works instantly with your dictation app, AI coding tools, and voice typing — the LED glows to show you're connected and turns red when muted.
  • DESKTOP STAND + 1-YEAR WARRANTY – Includes a desktop stand that keeps the mic at talking distance on your desk, backed by friendly US-based support and a 1-year warranty.

Also track generation completion separately from interruption. Completion means the model turn has finished; interruption means the current response was cut off. Use those events to update playback and interface state rather than assuming that every response ends normally.

Add function calls without giving the model direct authority

Function calling is a request-and-response cycle, not automatic execution. Declare only the functions the agent should be able to request. When Live API returns a function call, application code—not the model—must decide whether the operation is permitted, validate its arguments, perform the work, and send a function response containing the function name, call ID, and result.

Rank #4
Sale
NEEWER UM04 USB Microphone AI Cancelling Desktop Omnidirectional Condenser
  • 【Plug & Play Microphone】 Directly connect to a computer/laptop and use—no drivers needed. Compatible with macOS Windows PC iPhone Android for video conference, online teaching, Zoom calls, gaming, and podcast. Note: Set UM04 as the default input device on a PC if multiple audio devices are connected. Some phones may require OTG activation
  • 【Mute/AI Noise Cancellation/RGB】 Built with the DSP chip. Tap once to mute (red light on); tap twice to enable AI noise cancellation (green light on); tap and hold for 3s to turn dynamic RGB light effects on or off
  • 【Omnidirectional Pickup Pattern】 360° omnidirectional pattern evenly captures sound from all directions—portable mic and professional microphone for group online meetings or use by multiple persons in conference room. Optimal pickup distance: 4.9ft/1.5m
  • 【3.5mm TRS Headphone Jack】 Plug monitoring headphones into the 3.5mm jack to monitor audio in real time or in playback. Only supports 3.5mm TRS headphone output. Note: It is a microphone, not a speaker or speakerphone
  • 【10 Volume Adjustment Levels】 Supports 10 adjustable volume levels and mic gain control (2dB increments). Intuitive light effects, dynamic during volume adjustment, solid at max or min level, allow you to know the status at a glance
  1. Declare a narrow tool. Describe the function and its expected arguments in session configuration. Give it only the scope required for the task.
  2. Validate every request. Check argument types, ranges, user authorization, and any required confirmation before executing an action.
  3. Execute in trusted application code. Do not treat a model-generated request as proof of permission, and do not expose credentials or unrestricted backend operations through a tool.
  4. Return the result. Send the function response with the matching function name and call ID, then continue receiving the model turn.
  5. Handle failure explicitly. Return a controlled error or ask the user for clarification when a tool fails, arguments are invalid, or an action needs confirmation.

Google’s tools guide lists function calling and Google Search support, but tool availability depends on the model. Its table lists Gemini 3.1 Flash Live Preview with Search and synchronous function calling, and Gemini 2.5 Flash Live Preview with Search and synchronous or asynchronous function calling. The same table does not list Google Maps, code execution, or URL context as supported. Recheck the current table for the model you actually select.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Keep browser credentials and session lifecycle under control

For a production browser-to-Gemini connection, use ephemeral tokens rather than placing a standard long-lived API key in frontend code. A backend can issue a short-lived token while keeping the durable credential server-side. If your architecture routes media through a backend instead, account for the additional media hop and make the backend responsible for protecting credentials and enforcing tool permissions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
BOYA mini 2 Wireless Lavalier Microphone for iPhone 15/16/17 & Android
  • Thumb-Sized Mic: Weighing only 5 grams—the BOYA mini 2 lavalier microphone is the lightest microphone you can get. Its streamlined design seamlessly blends with your clothing for complete concealment and all-day comfort.
  • Adaptive AI Noise Cancellation: Instantly suppresses noise from clicks to roars. Activate Strong mode (-40 dB) for loud environments, or Light mode (-15 dB) to maintain a natural sound atmosphere.
  • 48kHz/24Bit Richer Sound: BOYA mini 2 microphone for iphone captures pristine audio with 48kHz/24-bit resolution for exceptional clarity. An 80dB signal-to-noise ratio ensures a pure recording, while a high 120dB SPL handles loud sounds without distortion.
  • Smart App Control: Unlock the full potential of your BOYA mini 2 clip on microphone with the free BOYA Central app. This app gives you quick access to key settings like volume, noise cancellation, and EQ—all from your phone.
  • Limiter & Safety Track: BOYA mini 2 lapel microphone wireless uses an limiter to prevent distortion by adjusting volume in real-time. A -12 dB safety track further guards against clipping, ensuring every recording is protected.

Plan for sessions to end and for conversational context to grow. Google’s capabilities guide lists 15 minutes for audio-only sessions and 2 minutes for audio-plus-video sessions without session-extension techniques. It lists context windows of 128k tokens for native-audio-output models and 32k tokens for other Live API models. These are documented limits, not a promise that a given session will last that long; verify the current values and supported extension methods before relying on them.

  • Session resumption: implement the documented resumption flow if a conversation needs to continue across a connection break.
  • Server shutdown notice: handle a server GoAway event and arrange a controlled reconnect or session transition instead of treating the socket as permanent.
  • Context growth: consider context-window compression for long conversations. Google’s best-practices guide estimates audio at approximately 25 tokens per second and recommends compression and session resumption for longer-running conversations.
  • Cost tracking: billing is token-based, and session context accumulates, so later turns can include more context than earlier ones. Check current official pricing and estimate usage for your own interaction pattern before setting a budget; a universal per-turn price cannot be inferred from the API format alone.

Test the agent against real interaction failures

A successful first reply is not enough to validate a streaming agent. Test the full path from microphone capture through playback, including cases that expose buffering, session, and authorization mistakes.

  • Speak while the agent is responding; confirm playback stops immediately on interruption and that queued audio is discarded.
  • Test microphone input at the device’s normal sample rate and verify that the application resamples to 16 kHz PCM rather than merely changing the MIME label.
  • Test short, continuous chunks and a pause longer than about a second; confirm the audio-stream-end behavior is handled.
  • Make a tool call with valid, invalid, and unauthorized arguments; confirm only explicitly permitted actions run and each call receives a response.
  • End or interrupt a turn, disconnect the network, and simulate a server shutdown notice; verify that the interface does not leave audio playing or imply the session is still active.
  • Run a longer conversation and monitor context growth, resumption behavior, and token usage before enabling extended sessions for users.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.