Free tools Windows power users keep installed
One-click scans. No signup required.
ElevenLabs AI Voice Isolator is a cloud tool that extracts human speech from recordings containing noise, music, echo, wind, room ambience, or competing voices. It works in the ElevenLabs web app and through the POST /v1/audio-isolation API. It costs 1,000 credits per minute and is not a dedicated vocal-stem separator.
That distinction determines whether Voice Isolator is the right tool. It is designed around intelligible dialogue, so it can be useful for interviews, podcasts, videos, meetings, and field recordings. It is a weaker choice when the goal is to produce a clean acapella, isolate instruments, perform detailed offline restoration, or preserve every nuance of a damaged recording.
Key takeaways
- ElevenLabs Voice Isolator separates human speech from background noise, music, ambience, echo, feedback, and other interference.
- Voice Isolator processes supported audio and video files up to 500 MB or one hour, whichever limit is reached first.
- ElevenLabs documents a usage rate of 1,000 credits per minute, so a 30-minute recording uses approximately 30,000 credits under the published rate.
- The tool may suppress music around dialogue, but ElevenLabs says Voice Isolator is not specifically optimized for isolating vocals from music.
- The web app offers upload, recording, preview, and download workflows, while the API returns audio from
POST /v1/audio-isolation.
What is ElevenLabs AI Voice Isolator?
ElevenLabs AI Voice Isolator is a speech-isolation service that attempts to extract intelligible human voice from a mixed recording. The input may contain traffic, wind, office noise, crowd chatter, room ambience, background music, microphone feedback, reverberation, or other competing sounds. ElevenLabs describes the product and its intended use cases on its official Voice Isolator page.
Voice Isolator is closer to an AI dialogue-extraction tool than to a conventional noise gate or a simple noise-reduction filter. Traditional noise reduction often estimates a noise profile and attenuates similar frequencies. Voice isolation instead tries to distinguish speech from the rest of a complex audio mixture. The result can be more useful than ordinary denoising when the background changes over time, but the model can also alter the voice or create processing artifacts.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
- [Convenient Setup] Plug and play recording USB microphone for PC, with 5.9-Foot USB cable included for computer PC laptop, is connected directly to USB-A port for recording music, computer singing or podcast. The office condenser microphone for computer is easy to use and install. (NOT compatible with Xbox and Phones)
- [Durable Metal Design] Solid sturdy metal construction design, the computer microphone for Zoom meetings with stable tripod stand is convenient when you are doing voice overs or livestreams on YouTube. Durable material extends the service life of the voice-over microphone.
- [Mic Volume Knob] Gaming condenser USB mic compatible for PS4 with additional volume knob itself has a louder or quieter adjustment and is more sensitive. Your voice would be heard well enough through the zoom microphone USB when gaming, skyping or voice recording. Also, you can adjust your volume to zero and protect your privacy.
- [Widely Use] USB-powered design, the condenser microphone for recording no need the 48v Phantom power supply, works well with Cortana, Discord, voice chat and voice recognition. The podcast microphone for Mac, with USB-B to USB-A/C cable, is compatible with desktop, laptop or PS4/PS5, which meets most of your daily recording needs.
- [Clear Output Voice] Cardioid condenser microphone for PC captures your voice properly, producing clear smooth and crisp sound. Great computer recording mic for gamers/streamers/youtubers focus on the main source and reduces background noise. The streaming microphone does the job well for broadcast ,OBS and teamspeak.
“Cleaner” does not always mean “more natural.” Difficult recordings may develop metallic or watery speech, burbling sustained vowels, missing consonants, distorted breaths, unnatural pauses, or residual background sound. Always compare the processed file with the original rather than overwriting the source.
Does Voice Isolator remove background noise and music?
Voice Isolator can suppress many sounds around spoken dialogue, including background chatter, traffic, wind, office ambience, echo, microphone feedback, and music beneath speech. Performance depends heavily on the recording: speech that is loud and distinct from the background is generally an easier separation problem than speech mixed at a similar level with dense music or another voice.
Music claims require particular caution. ElevenLabs’ product material presents removing background music as a use case, but the technical Voice Isolator documentation says the system is not specifically optimized for isolating vocals from music. Voice Isolator may produce more intelligible speech from a video or interview with music underneath; that does not make it a reliable tool for creating a clean vocal stem or instrumental track from a song.
| Recording problem | What Voice Isolator may do | Important limitation |
|---|---|---|
| Traffic, HVAC, or office ambience | Reduce changing environmental interference around speech | Residual noise or voice artifacts may remain |
| Wind | Suppress some wind competing with dialogue | Severe low-frequency distortion can remain difficult |
| Background music | Reduce music so spoken words are easier to hear | Not a guaranteed vocal or instrumental stem extractor |
| Echo and reverberation | Attempt to reduce room reflections | Strong reverb is part of the recorded speech environment and may cause artifacts |
| Overlapping speakers | Attempt to preserve a speech signal from a mixture | The model may damage or confuse voices speaking simultaneously |
| Clipped audio | Process the distorted recording | Isolation cannot reliably reconstruct waveform information lost to clipping |
Who should use ElevenLabs Voice Isolator?
Voice Isolator is a sensible first test when the important signal is spoken dialogue and the source was recorded in an uncontrolled environment. Typical users include:
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →- Podcasters and interviewers: Clean a conversation recorded in a noisy office, café, street, or event venue.
- Video editors: Improve dialogue from footage before replacing or synchronizing the audio track in an editor.
- Journalists: Make an interview easier to understand, while retaining the original and reviewing the result rather than treating the output as unquestionable evidence.
- Creators: Rescue social-media or YouTube recordings made without ideal microphone placement.
- Meeting and lecture users: Improve speech audibility in recordings with room noise or distant microphones.
- Developers: Add hosted speech cleanup to an application through the API.
- Voice-library users: Clean a speech clip before using it to search ElevenLabs’ public Voice Library for a possible matching voice. A match does not establish ownership or prove a speaker’s identity; the official Voice Library guidance describes this as a search workflow.
What files and limits does Voice Isolator support?
According to the Voice Isolator capability documentation checked on August 18, 2026, the service lists the following formats and accepts files up to 500 MB or one hour, whichever limit is reached first. Format availability can change between the web interface and other ElevenLabs product surfaces, so the live interface should take precedence if it differs.
| Category | Documented formats | Limit |
|---|---|---|
| Audio | AAC, AIFF, OGG, MP3, OPUS, WAV, FLAC, M4A | Up to 500 MB or one hour |
| Video | MP4, AVI, MKV, MOV, WMV, FLV, WEBM, MPEG, 3GPP | Up to 500 MB or one hour |
If a file exceeds a limit, trim irrelevant silence, export a more efficient encoded copy, or split the recording into sections. Keep the original file and use stable boundaries or timecode notes so processed sections can be aligned accurately later.
Rank #2
- [USB Output] Enables simple setup. USB studio recording microphone kit provides a direct convenient plug-and-play connection to pc and laptop without any additional hardware or drivers for recording vocals, podcasts and Skype. Studio microphone for recording vocals is never been easier to get high-quality sound for your voice and computer-based audio recordings. (Incompatible with Xbox)
- [Excellent Sound Quality] With rugged construction for durable performance, the vocal recording microphone, USB condenser mic for PC,offers a wide frequency response and handles high SPLs with ease. Ideal for project/home-studio applications. The cardioid condenser capsule captures crystal-clear audio from the front and avoid ambient noise when communicating/creating/recording. Comes ready to go with a desktop mic boom arm stand and 8.2ft USB cable, you're guaranteed to get great-sounding results.
- [Durable Arm Set] The podcast microphone bundle with versatile and sturdy broadcast suspension boom scissor arm with 180° up and down rotation, 135° forward and backward extension for optimal adjustment, for capturing your voice in podcast or voiceover. The double pop filter attached on the music recording microphone provides two layers of dissipation, removes the rush of air, minimize the popping sounds or cancel noise that can compromise your recording, great for studio as well as home use.
- [Easy to Attach] The streaming microphone for PC includes adjustable boom studio scissor arm stand that features a heavy-duty combo mount consisting of a sturdy C-clamp and a detachable desktop mount. With 13" fixed horizontal arm and offers a 30" reach, the low-profile, table-hugging design of audio recording microphone allows on-air talent to perform without facial obstruction to record in podcasting or make dubbing sounds for videos, use voice chat in Discord or online conference on Zoom or Skype.
- [The Accessory Package Includes] The studio microphone music recording comes with practical accessories for you to use in most of recording. The scissor arm stand is made out of all steel construction, sturdy and durable, a studio-grade shock mount, a double pop filter, premium 8.2' USB-B to USB-A/C cable, a podcast PC gaming microphone, a user manual and friendly Technical Support.
How do you use Voice Isolator in the ElevenLabs web app?
The documented web-app route is ElevenLabs → Audio Tools → Voice Isolator. Menu names and placement can change; the following sequence reflects the documented workflow observed on August 18, 2026.
- Sign in to ElevenLabs.
- Open Voice Isolator under Audio Tools.
- Upload or drag and drop a supported audio or video file. The interface may also let you record directly with the device microphone.
- Select Isolate voice.
- Wait for processing to finish.
- Preview the isolated result and compare difficult passages with the original.
- Download the cleaned audio file.
For video work, plan for a possible audio replacement step. The documentation says supported video files can be processed, but the API is explicitly audio-oriented, and the current web interface may return cleaned audio rather than a remuxed video. Import the result into your video editor, replace the dialogue track, and check synchronization from the beginning, middle, and end of the clip.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
How do you use the Voice Isolator API?
The API uses POST https://api.elevenlabs.io/v1/audio-isolation with multipart/form-data. The required multipart field is audio, and the response should be treated as binary audio rather than assumed to be JSON or text. The official API reference documents the endpoint, parameters, and a possible HTTP 422 response for an unprocessable request.
Minimal cURL request
curl -X POST "https://api.elevenlabs.io/v1/audio-isolation"
-H "xi-api-key: $ELEVENLABS_API_KEY"
-H "Content-Type: multipart/form-data"
-F "audio=@input.mp3"
--output isolated_audio.mp3
Store the API key in an environment variable or server-side secret manager. Do not place the key in browser JavaScript, a public mobile application, a downloadable script, or source control.
Optional output format
The API accepts an optional file_format parameter. The documented other value is the default for encoded audio. The pcm_s16le_16 value requests 16-bit PCM, 16 kHz, mono, little-endian audio, which ElevenLabs says can provide lower latency in suitable workflows.
curl -X POST "https://api.elevenlabs.io/v1/audio-isolation"
-H "xi-api-key: $ELEVENLABS_API_KEY"
-F "audio=@input.wav"
-F "file_format=pcm_s16le_16"
--output isolated_audio.pcm
Use the PCM option only when the receiving application expects those exact raw-audio parameters. A raw PCM response may not contain a normal file header, so an audio editor or playback tool needs the correct sample rate, bit depth, channel count, and byte order.
Rank #3
- [Natural Audio Clarity] Operated with frequency response of 50Hz-16KHz, the podcasting XLR mic delivers balanced audio range, likely to resonate with your audience. Directional cardioid dynamic microphone corded will not exaggerate your voice, while rejects unwanted off-axis noise for vocal originality and intelligibility during your PS5 gaming streaming video recording. (Tips: Keep the top of end-addressing XLR dynamic microphone AM8 facing audio source, and suggested recording range is 2 to 6 in.)
- [XLR Connection Upgrade-Ability] To use XLR connection, connect the podcast microphone to an audio interface (or mixer) using a separate XLR cable (NOT Included) . Well-connected and smooth operation improves audio flexibility to make you explore various types of music recording singing. The streaming mic isolates the pristine and accurate sound from ambient noise with greater no interference and fidelity. (RGB and function key on mic are INACTIVE when using XLR connection.)
- [USB Connection with Handy Mute] Skip the hassle of setting something up and plug the cable to play the dynamic USB microphone directly, which suits for beginner creators or daily podcast. You can quickly control the gamer mic with tap-to-mute that is independent of computer/Macbook programs to keep privacy when live streaming. LED mute reminder helps you get rid of forgetting to cancel the mute. (RGB and function key are only available for USB connection, but NOT for XLR connection)
- [Soothing Controllable RGB] RGB ring on the desktop gaming microphone for PC, with 3 modes and more than 10 light colors collection, matches your PC gears accessories for gaming synergy even in dim room. You can control the RGB key button of the dynamic microphone USB directly for game color scheme gaming or live streaming. Configured memory function, the streaming microphone RGB no need to repeated selections after turnning off and brings itself alive when power on. (Only available for USB connection)
- [More Function Keys] Computer microphone with headphones jack upgrades your rhythm game experience and gets feedback whether the real-time voice your audience hear as expected. Get the desired level via monitoring volume control when gaming recording. Smooth mic gain knob on the PC microphone gaming has some resistance to the point, easily for audio attenuation or boost presence to less post-production audio. (Only available for USB connection)
Official Python SDK pattern
The ElevenLabs API quickstart shows the SDK pattern below. The example downloads an input file, sends a file-like object to audio_isolation.convert, and writes the binary response to disk.
import os
import requests
from io import BytesIO
from elevenlabs.client import ElevenLabs
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
response = requests.get(
"https://storage.googleapis.com/eleven-public-cdn/"
"documentation_assets/audio/voice_with_background.mp3"
)
response.raise_for_status()
audio_data = BytesIO(response.content)
audio_stream = client.audio_isolation.convert(audio=audio_data)
with open("isolated_audio.mp3", "wb") as f:
f.write(audio_stream.read())
The official quickstart also uses the elevenlabs and python-dotenv packages. Your production code should additionally check upload errors, response status, timeouts, file size, and output validity before saving or publishing the result.
How much does ElevenLabs Voice Isolator cost?
According to ElevenLabs’ Voice Isolator cost guidance, Voice Isolator consumes 1,000 credits per minute of audio. The following estimates apply the published per-minute rate; actual partial-minute rounding should be confirmed in the account before treating a short-file estimate as a billing guarantee.
| Audio duration | Estimated credits at 1,000 per minute |
|---|---|
| 30 seconds | Approximately 500 |
| 1 minute | 1,000 |
| 5 minutes | 5,000 |
| 10 minutes | 10,000 |
| 30 minutes | Approximately 30,000 |
| 60 minutes | 60,000 |
Audio credits can also be used by other ElevenLabs products, including Voice Changer, sound effects, and Dubbing Studio. The ElevenLabs credits explanation says self-serve credits may roll over subject to plan rules. Trim silence and irrelevant material before processing, but retain enough context to avoid awkward edits or alignment problems.
As an observed signal on the usage-based billing page checked for this article, additional credits were listed at $0.30 per 1,000 credits for Creator, $0.24 for Pro, $0.18 for Scale, and $0.12 for Business. These are not timeless prices: plan structures, included credits, taxes, rollover terms, and usage-based rates can change. Check the live ElevenLabs pricing page before subscribing or purchasing extra credits.
ElevenLabs’ Voice Isolator product page currently advertises 10 minutes of isolated audio free for new or eligible users. Eligibility, geography, promotions, and signup terms can change, and the exact free-plan credit allowance has appeared differently across recent pricing-page snapshots. Treat the offer as current advertising rather than a permanent entitlement.
Rank #4
- Custom three-capsule array: This professional USB mic produces clear, powerful, broadcast-quality sound for YouTube videos, Twitch game streaming, podcasting, Zoom meetings, music recording and more
- Blue VO!CE software: Elevate your streamings and recordings with clear broadcast vocal sound and entertain your audience with enhanced effects, advanced modulation and HD audio samples
- Four pickup patterns: Flexible cardioid, omni, bidirectional, and stereo pickup patterns allow you to record in ways that would normally require multiple mics, for vocals, instruments and podcasts
- Onboard audio controls: Headphone volume, pattern selection, instant mute, and mic gain put you in charge of every level of the audio recording and streaming process
- Positionable design: Pivot the mic in relation to the sound source to optimize your sound quality thanks to the adjustable desktop stand and track your voice in real time with no-latency monitoring
What should you expect from difficult recordings?
Voice Isolator can improve intelligibility without producing a faithful restoration. Listen for artifacts and decide whether the result is better for the intended use, not merely more aggressively processed.
- Changing background noise: Traffic, HVAC, room ambience, and office noise are reasonable test cases for speech isolation. A low-noise source may need little processing, while a changing background may be difficult to remove uniformly.
- Wind: Wind can mask speech and overload a microphone’s low frequencies. Isolation may help, but severe wind distortion can be embedded in the speech signal.
- Echo: Room reflections overlap the voice rather than sitting in a perfectly separate track. Strong reverberation may remain or produce watery artifacts.
- Music: Dialogue over quiet instrumental music may be usable after processing. Dense arrangements, sung backing vocals, or music occupying the same frequencies as speech are harder cases.
- Overlapping voices: When speakers talk simultaneously, the model must decide which vocal content to preserve. Review the result carefully before using it for journalism, legal work, or evidence.
- Clipping: If the original recording is clipped, missing waveform information cannot reliably be reconstructed by isolation.
- Non-speech foreground sounds: Voice Isolator is optimized around human speech, not reliable extraction of individual instruments, animal sounds, sound effects, or arbitrary desired foreground sounds.
If no controlled hands-on test has been performed on your specific material, do not interpret ElevenLabs’ “studio-quality” or “studio-grade” marketing language as an independent performance measurement. Test a representative excerpt first and preserve both versions.
What should you do after isolation?
- Keep the original: Duplicate the source before uploading or editing.
- Inspect the result: Listen for missing consonants, burbling vowels, metallic tone, distorted breaths, residual music, and unnatural pauses.
- Use moderate downstream processing: Apply editing, EQ, de-essing, compression, or loudness normalization only after checking what Voice Isolator already changed.
- Replace dialogue carefully: In a video editor, align the processed audio with the original clip and check sync at multiple points.
- Export for the destination: Use the audio format and loudness requirements of the podcast host, video platform, archive, or application.
- Retain provenance: For interviews, journalism, legal matters, and archives, preserve the unprocessed file and document that the published or reviewed version was AI-processed.
How does Voice Isolator compare with alternatives?
No single alternative is universally better. Choose according to whether the central problem is speech cleanup, podcast mastering, live noise suppression, transcription-led editing, or music stem separation.
| Tool or category | Best fit | How it differs from Voice Isolator | Key qualification |
|---|---|---|---|
| ElevenLabs Voice Isolator | Cloud speech isolation, dialogue cleanup, and API workflows | Combines a browser tool with a documented audio-isolation API and a broader AI-audio platform | 1,000 credits per minute; not specifically optimized for vocal-from-music separation |
| Adobe Podcast Enhance Speech | Browser-based podcast and creator enhancement | Adobe advertises noise and echo removal, video support, bulk enhancement, and adjustable speech, music, and ambience levels | Adobe’s pricing page advertises Premium processing up to four hours per day and files up to 1 GB; limits and plans can change |
| Auphonic | Automated podcast post-production, leveling, loudness normalization, and batch workflows | More oriented toward production automation and mastering beyond isolation | Current pricing was not verified for this article |
| Descript | Transcription-led audio/video editing | Provides an editing environment with transcription, filler-word removal, screen recording, and dialogue production | Current pricing was not verified for this article |
| Krisp | Real-time microphone noise cancellation during calls and meetings | Designed primarily to prevent or reduce noise during live communication rather than clean uploaded files in post-production | Poorer fit for finished recordings, music separation, and offline editing |
| Desktop editors and restoration tools | Manual spectral repair, multitrack editing, detailed mixing, or offline work | May offer greater manual control through applications such as Audacity, DaVinci Resolve, Adobe Audition, or iZotope RX | Performance, format support, pricing, and stem-separation capability vary by product and should be tested separately |
Adobe Podcast is worth considering when a browser-based podcast workflow matters more than an ElevenLabs API. Adobe’s official pricing and features page advertises background-noise and echo removal, video support, bulk enhancement, adjustable speech/music/ambience levels, a Premium limit of up to four hours per day, files up to 1 GB, and a 30-day free trial. Verify the live terms before relying on those limits.
Consider Auphonic when leveling, loudness normalization, and batch podcast production are as important as cleaning speech. Consider Descript when transcription-based editing and video production belong in the same workspace. Consider Krisp when the problem is live call noise rather than post-production restoration. For clean acapellas or instrumental tracks, use a dedicated music stem-separation category rather than assuming a speech-isolation service will deliver equivalent results.
What are the privacy and compliance considerations?
Voice Isolator is a cloud service, so confidential recordings should be reviewed under your organization’s security, privacy, and data-processing requirements before upload. ElevenLabs advertises encryption, SOC 2, HIPAA and GDPR compliance, EU data residency, and Zero Retention modes on its product page. Those are vendor claims that may depend on the plan, region, contract, configuration, and applicable data-processing terms; verify the specific requirement with ElevenLabs before making a compliance decision.
Recommended Free Tools
Best Value
- Cardioid Pick-up: Cardioid pickup pattern that captures clear and crisp voice in front of the mic and suppresses unwanted background noise. Design for chatting, teleconferencing, recording, podcast
- For Podcast: Equipped with a non-slip stand that adds stability while occupying a small desktop area. One-click mute and volume control for easy operation during the recording. The shock mount and pop filter can prevent recordings from being disturbed by vibration
- Strong Compatibility: TC-777 is multi-device and program compatible, you can use it on Windows, MAC, PS4 and 5. It can also be quickly recognized by Zoom, Skype, Discord, allowing you to start creating or communicating immediately. (Not compatible with Xbox)
- Plug & Play: With a USB 2.0 data port, the TC-777 is plug and play, with no additional drivers or assembly process required. The angle of both microhone and pop filter can be adjusted as needed to achieve the best audio effect
- What's In the Box: 1 x Microphone with Power Cord(1.9m), 1 x Foldable Mic Tripod, 1 x Mini Shock Mount, 1 x Pop Filter and 1 x Manual
Platform licensing also does not grant rights to the recording, the people speaking in it, the background music, or any voice used in the source. Confirm that your organization has permission to process and publish the material, especially for interviews, customer calls, copyrighted music, and biometric or health-related information.
Common API and upload failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Authentication error | Missing, invalid, revoked, or incorrectly named API key | Send the key in xi-api-key, verify the environment variable, and keep the key server-side |
| Unprocessable request or HTTP 422 | Malformed or unsupported file, missing multipart field, or invalid parameter | Check the API reference, use the required audio field, and validate the file before upload |
| Upload rejected | File exceeds 500 MB or one hour | Trim, recompress, or split the file while retaining alignment notes |
| Unreadable output | Binary audio was handled as text or JSON | Save the response as bytes and inspect the returned format |
| PCM playback failure | Output is raw 16-bit, 16 kHz, mono, little-endian PCM and the player lacks those parameters | Configure the player or editor with the documented PCM properties, or use the default encoded output |
| Video has no cleaned picture track | The API is audio-oriented | Import the returned audio into a video editor and remux or replace the original track yourself |
Is ElevenLabs Voice Isolator worth using?
ElevenLabs Voice Isolator is a strong first test for cloud-based cleanup when the desired signal is spoken dialogue and the recording contains unpredictable background interference. The browser workflow is straightforward, the service accepts both audio and video inputs within documented limits, and developers can call the capability through an API.
Use a short representative sample before processing a long recording. The 1,000-credits-per-minute rate makes a one-hour file consume 60,000 credits, and difficult material may still contain artifacts or require manual editing. Voice Isolator is not the default choice for music stem separation, offline restoration, severe clipping, or simultaneous speakers whose exact words must be preserved.
| Choose this when you need… | Most appropriate direction |
|---|---|
| Fast browser-based cleanup of spoken dialogue | ElevenLabs Voice Isolator |
| Speech cleanup through a hosted developer API | ElevenLabs Voice Isolator API |
| Podcast enhancement with bulk and loudness-oriented workflow | Adobe Podcast or Auphonic |
| Transcription-led editing and video production | Descript |
| Live microphone noise cancellation | Krisp |
| Clean instrumental or vocal music stems | Dedicated music stem-separation software or service |
| Offline, highly controlled restoration | Desktop audio editor or restoration suite |
Frequently Asked Questions
Is ElevenLabs Voice Isolator free?
ElevenLabs currently advertises 10 free minutes of isolated audio for new or eligible users, but eligibility, geography, promotions, and signup terms can change. Voice Isolator otherwise consumes 1,000 ElevenLabs credits per minute, so check the live pricing page and your account before processing a long file.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteCan ElevenLabs Voice Isolator remove vocals from a song?
ElevenLabs Voice Isolator is designed to isolate human speech, not to reliably create vocal or instrumental stems from music. It may suppress music around spoken dialogue, but ElevenLabs says the tool is not specifically optimized for isolating vocals from music.
Can Voice Isolator process video?
Yes. ElevenLabs documentation lists multiple supported video formats, including MP4, AVI, MKV, MOV, WMV, FLV, WEBM, MPEG, and 3GPP, with a maximum of 500 MB or one hour. The API is audio-oriented, so developers should expect to handle the cleaned audio separately.
How many credits does a 30-minute recording use in Voice Isolator?
A 30-minute recording uses approximately 30,000 credits at ElevenLabs’ documented rate of 1,000 credits per minute. Actual billing behavior for partial minutes and plan-specific rules should be confirmed in the account before processing.
The Bottom Line
Bottom line: ElevenLabs Voice Isolator is best understood as a cloud speech-cleanup and dialogue-isolation tool. It is convenient for interviews, podcasts, videos, and API applications, but it is not a guaranteed music stem separator or a replacement for careful offline restoration. Test a short excerpt, retain the original, inspect artifacts, and calculate credit usage before sending a long recording.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →

