Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Microsoft’s VALL-E 2 is a research-only AI speech generator, not a tool the public can currently try or buy. Microsoft Research says it has no plans to put the model in a product or expand public access. The researchers reported “human parity” on two speech benchmarks, while also warning that voice impersonation and spoofing could be misused.

What VALL-E 2 does—and what “human parity” means

VALL-E 2 is a zero-shot text-to-speech system: it can generate speech in the style of a speaker using a speech prompt, rather than relying on a conventional model trained separately for each speaker at inference time. The paper, posted June 8, 2024, represents speech with discrete codes from a neural audio codec and uses a codec language-model approach. Its methods include repetition-aware sampling to reduce unstable, repetitive output and grouped code modeling to shorten code sequences and improve long-sequence modeling efficiency. Read the VALL-E 2 paper.

The authors evaluated the system on LibriSpeech and VCTK, reporting improvements over earlier systems in robustness, naturalness, and speaker similarity. Their “human parity” result refers to specified metrics in those benchmark experiments; it does not establish that every generated voice sounds human to every listener or will pass as genuine in every setting. Microsoft Research likewise limits the parity claim to the LibriSpeech and VCTK experiments. Microsoft Research’s announcement describes the project and its results.

The paper’s subjective evaluations included 40 test cases on the LibriSpeech test-clean set and 60 test cases from 60 distinct speakers on VCTK. Its benchmark comparisons draw on results reported in prior papers, despite differences in model architectures and training data. Those comparisons provide evidence about performance under the study’s conditions—not proof of universal equivalence to human speech or that synthetic speech cannot be detected.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Voice Recorder 5000mAh 128GB AI Intelligent Triple Noise Reduction, Long Battery Life 40 Days, Voice Activated, Magnetic Digital Audio Recorder for Meetings, Interviews, Lectures, Classroom
  • AI-Triple Noise Reduction Technology: The voice recorder utilizes AI intelligence, featuring a triple noise reduction system that intelligently detects and models noise. Through DSP chips, it effectively reduces noise, enhancing audio quality for a clearer and purer sound experience
  • 40 Days Continuous Recording Capability: The audio recorder is equipped with a 5000mAh large-capacity battery, capable of supporting continuous recording for up to 35 days or 1000 hours. With just one charge, it meets the usage demands of various scenarios
  • Dual Powerful Magnetic Design: The recording device features a dual powerful magnetic suction design, ensuring a firm and reliable attachment to any ferrous surface, freeing up your hands for added convenience
  • One-Touch Operation System: This mini recorder device is equipped with one-touch power-on and save functions, allowing you to easily start the device and provide protection measures to ensure safe operation. Additionally, the one-touch voice activation feature enables you to enjoy a convenient hands-free experience without the hassle of complicated operations
  • Large Storage Capacity: The digital voice recorder is equipped with a 128GB large-capacity storage card, providing up to 460 days of standby time, supporting continuous recording for up to 1000 hours, and capable of storing up to 9500 hours of files

Does VALL-E 2 clone a voice from three seconds of audio?

Three seconds is associated with the earlier VALL-E system: the VALL-E 2 paper describes its predecessor as using a three-second recording. VALL-E 2’s experiments also include short prompts, including three-second prompts, but that is not a guarantee that any voice can be perfectly cloned from exactly three seconds. Prompt length and quality, background noise, the speaker, and evaluation conditions all affect results.

Can the public try or buy VALL-E 2?

No public access is described in Microsoft Research’s published position. The research team stated: “VALL-E 2 is purely a research project. Currently, we have no plans to incorporate VALL-E 2 into a product or expand access to the public.” That is a statement of current plans, not a promise of a permanent ban or a formal determination that the model can never be released.

Rank #2
64GB Magnetic Voice Activated Recorder - 40 Hours Continuous Recording Device with AI-Intelligent Triple Noise Reduction - Portable Audio Recorder Device for Lectures Meetings Interviews
  • [Smart Phone Connectivity for File Management]: L810 Voice Recorder supports direct connection to smartphones via an OTG adapter. This innovative feature allows you to manage your audio files on the go. You can easily rename, forward, or delete files directly from your smartphone.This is perfect for busy professionals, students, and journalists who need to quickly access and share their recordings
  • [Efficient Voice Activation Function]: With the voice activation feature, L810 recorder only starts recording when it detects sound above 45dB . This means you can save storage space and time by avoiding recording silent periods. The 60° wide-angle recording capability ensures that all sounds are captured clearly, making it perfect for large classrooms, conference rooms, or interview settings
  • [Crystal Clear Sound Quality]: Equipped with advanced microphones and AI noise reduction technology, this audio recorder effectively filters out background noise, ensuring you capture crystal-clear audio. Whether you're recording lectures, meetings, interviews, or daily conversations, the high-quality sound makes it easy to understand every word
  • [Convenient Recording and Playback]: One-click operation, VA mode for voice activated recording, ON mode for regular recording, OFF to save recording. Equipped with a headphone adapter to support volume adjustment, track switching and playback speed
  • [64GB Storage Capacity]: This portable recorder offers a generous 64GB of storage, capable of holding up to 768 hours of audio files at 192kbps quality . A quick 2-hour charge provides up to 28 hours of continuous recording, and it can even record while charging. Plus, it automatically saves your recordings when the battery is low, ensuring you never lose important audio

Microsoft offers separate Azure speech services, including text-to-speech and custom neural voice capabilities. They are not VALL-E 2 access. Microsoft’s Azure guidance recommends disclosing when speech or avatars are synthetic, choosing voice types carefully for the use case, obtaining explicit written permission for custom neural voices, and providing human support when a transactional interaction is ambiguous. Azure text-to-speech documentation and custom neural voice documentation describe those separate services and guidance.

Why does Microsoft identify misuse risks?

A convincing generated voice could be used to impersonate a specific person or spoof voice identification. Microsoft’s researchers say their experiments assumed users had consent to be the target speaker. They add that real-world generalization to unseen speakers should include a speaker-approval protocol and a synthesized-speech detection model. These are safeguards the researchers say would be needed; the announcement does not say VALL-E 2 has been made publicly available with those protections.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
RECOLX AI Voice Recorder, AI Transcriber with GPT-5.2, Pearl Gray
  • GPT-5.2 AI Transcription & Summary Turn hours of audio into clear text and concise key-point summaries with GPT-4o/5/5.2/0SS-120b, 03-mini,Gemini-3-Pro,Claude-Sonnet-4.5 powered AI. Perfect for meetings, lectures, interviews and brainstorming sessions when you don’t want to take notes by hand.
  • Language Speech-to-Text Support Record in up to 112 languages and accents and convert speech to text with high accuracy. Ideal for international teams, bilingual students, researchers and anyone working across multiple languages.
  • Long-Lasting, All-Day Recording Up to 30 hours of continuous recording on a full charge keeps you covered across business days, conferences or back-to-back classes without worrying about battery.
  • Clear Audio with Noise Reduction High-sensitivity microphone and intelligent noise reduction help capture your voice clearly, even in busy offices, classrooms or cafés, so transcripts stay accurate and easy to read.
  • Portable, Easy Workflow Anywhere Slim, pocket-friendly design goes with you to meetings, lectures, interviews and trips. Connect via USB-C to quickly export audio and text files to your laptop or cloud tools for easy organizing and sharing.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How can you tell whether a voice is AI-generated?

The VALL-E 2 benchmark results do not establish a dependable way for listeners to identify synthetic speech, nor do they prove that detection is impossible. A natural-sounding voice alone is not reliable proof of identity. For consequential requests—especially requests to transfer money, reveal credentials, or bypass normal approval—verify through a separate trusted channel rather than relying on a voice call or clip. Organizations deploying synthetic speech should disclose its use and use suitable detection and consent controls; Microsoft’s Azure guidance recommends disclosure and permission for its separate services.

How to compare voice-generation claims

A broad claim that a system “sounds human” is hard to interpret without the conditions behind it. When comparing speech generators, check:

  • Evaluation: Which dataset, metrics, and listening method were used, and were comparisons conducted under equivalent conditions?
  • Speech quality: How did the system perform on robustness, naturalness, and speaker similarity, including difficult or repetitive text?
  • Prompt conditions: What prompt length and recording quality were tested, and how did noise or speaker differences affect results?
  • Safeguards: Are consent, voice-ownership approval, synthetic-speech disclosure, and detection addressed?
  • Access: Is the specific system actually available to the intended user, or is it a research project?

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.