Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parakeet-TDT-0.6B-v3 is a compact, locally runnable speech-to-text model that supports 25 European languages and reports strong results on selected speech-recognition benchmarks. Its measured accuracy is promising, but it is not a universal guarantee: results depend on language, recording conditions, dataset, and decoding setup. Here’s what it does, how it performs, and what you need to run it.

What is Parakeet-TDT?

Parakeet-TDT-0.6B-v3 is NVIDIA’s multilingual automatic speech-recognition (ASR) model: software that converts spoken audio into text. NVIDIA describes it as a 600-million-parameter model designed for high-throughput transcription. Its architecture combines a FastConformer encoder with a Token-and-Duration Transducer (TDT) decoder. The v3 release extends the earlier English-only v2 model with multilingual coverage and automatic input-language detection. NVIDIA model card

For practical transcription, the model card documents punctuation and capitalization, word- and segment-level timestamps, and options for handling long audio. Input examples use 16 kHz mono audio in WAV or FLAC format. Those features make it more useful than a bare speech-to-text output when you need timed text or readable transcripts.

How accurate is Parakeet-TDT?

NVIDIA’s 2025 model card reports a 6.34% average word error rate (WER) on its listed Open ASR Leaderboard evaluation. It also reports 1.93% WER on LibriSpeech test-clean and 3.59% on LibriSpeech test-other. Lower WER means fewer word substitutions, deletions, and insertions relative to the reference transcript. The model card’s benchmark results

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
  • AI Performance: 767 AI TOPS
  • OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode)
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • A 2.5-slot design maximizes compatibility and cooling efficiency for superior performance in small chassis

These scores describe particular evaluations, not every real-world recording. Benchmark WER excludes punctuation and capitalization errors, and results can differ across languages, datasets, accents, background noise, and decoding settings. Treat the scores as useful reference points rather than a promise for your audio—especially for noisy recordings or high-stakes medical and legal transcription, where human review may be essential.

Speed is a separate question from accuracy. The official card does not establish one universal throughput or real-time factor; performance depends on the GPU, runtime, batch size, quantization, and audio length. A result from one setup should not be assumed for another.

Which languages does it support?

Version 3 supports 25 European languages and automatically detects the language of the input audio, according to NVIDIA. The model card is the appropriate reference for the complete language list and any usage details. This broader coverage is the key difference from the English-only v2 release. See the supported languages in NVIDIA’s model card

Can you run Parakeet-TDT locally?

Yes. NVIDIA documents local deployment through NeMo-Speech.cpp, NVIDIA NeMo, and Transformers. The model is optimized for NVIDIA GPU-accelerated systems, but the card also states a minimum of 2 GB RAM to load it. That is a loading minimum, not a recommendation for fast or high-throughput production use; actual hardware needs and speed depend on the runtime and workload.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NeMo-Speech.cpp

The model card’s local command-line route uses a GGUF model. After installing and configuring NeMo-Speech.cpp and obtaining the model file, transcribe a WAV file with:

nemo-speech transcribe audio.wav

NVIDIA NeMo

In a NeMo environment, the card shows loading the model with ASRModel.from_pretrained using the identifier nvidia/parakeet-tdt-0.6b-v3, then transcribing audio and requesting timestamps. Consult NVIDIA’s current model-card examples for the exact code and environment requirements. NeMo usage examples

Rank #2
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4. System Requirements: Minimum 850W PSU with 16-pin 12V-2x6 (12VHPWR) connector required. Verify before purchasing.
  • Military-grade components deliver rock-solid power and longer lifespan for ultimate durability. Compatibility: 348mm (13.7") length, 3.6 slots, 4.3 lbs. Confirm case clearance and slot spacing. GPU bracket included.
  • Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
  • 3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans
  • Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads

Transformers

The documented Transformers route uses AutoModelForTDT and AutoProcessor. NVIDIA’s model card notes that official Transformers support may require installing Transformers from source as of the card’s publication. Check the current implementation and package compatibility before building a deployment around this route.

Long recordings

NVIDIA documents local-attention settings for extending processing beyond full-attention limits. Long-audio behavior depends on configuration, so check the model-card guidance and test with recordings representative of your intended workload.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Parakeet-TDT vs. Whisper

There is no single meaningful winner without specifying the Whisper model, language, dataset, hardware, and decoding configuration. Parakeet-TDT v3’s clearest distinguishing points are NVIDIA’s published 25-language European coverage, automatic language detection, local deployment paths, and its reported benchmark WER. For a fair comparison, transcribe the same representative audio with the chosen versions and settings, then compare WER, latency, resource use, and the quality of timestamps or other features you need. Do not compare headline scores from different datasets as if they were a direct head-to-head test.

What GPU do you need?

NVIDIA positions Parakeet-TDT for GPU-accelerated systems, but its model card does not specify one required GPU model or a universal VRAM recommendation. The 2 GB RAM figure is a minimum to load the model, not an inference-performance target. For local inference, a CUDA-capable NVIDIA graphics card is the most directly aligned hardware choice; confirm compatibility with your selected runtime and test the workload at the batch size and audio duration you expect to use.

License and training context

NVIDIA releases Parakeet-TDT-0.6B-v3 under CC BY 4.0 and describes it as ready for commercial and non-commercial use, subject to compliance with the license. Review the license terms for your intended use. The model card reports training for 150,000 steps on 128 A100 GPUs and fine-tuning for 5,000 steps on 4 A100 GPUs; these figures describe the published training setup, not hardware required to run transcription. License and model details

Quick Recap

SaleBestseller No. 1
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
AI Performance: 767 AI TOPS; OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode); Powered by the NVIDIA Blackwell architecture and DLSS 4
$790.71
Bestseller No. 2
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans; Auto-Extreme precision automated manufacturing helps ensure higher reliability
$1,831.31

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.