To evaluate word error rate (WER) in a brain-to-text system, calculate substitutions, deletions and insertions against a reference transcript, then divide their total by the number of words in that reference. The result is meaningful only alongside the task, participants, test split, vocabulary, decoder and scoring protocol that produced it. WER alone cannot tell you whether a system supports fast, useful communication.
How do you calculate word error rate?
WER counts the word edits required to turn the reference transcript into the system’s predicted transcript. Its conventional formula is:
WER = (S + D + I) / N
- S: substitutions, where a predicted word replaces a reference word.
- D: deletions, where a reference word is missing from the prediction.
- I: insertions, where the prediction contains an extra word.
- N: the number of words in the reference.
For example, if a 20-word reference requires two substitutions, one deletion and one insertion to match the output, WER is 4/20, or 20%. This does not mean that 80% of the words were “understood”: WER is an edit ratio, and insertions can make it exceed 100%. The foundational Brain-To-Text paper describes the measure as a way to assess decoded phrases: Frontiers, 2015.
Specify how multiple utterances are combined
A common corpus-level method pools all test-set edits and divides by the total number of reference words. This gives longer utterances more weight. Another approach calculates each sentence’s WER and averages those percentages, giving every sentence equal weight regardless of length. These methods can produce different scores, so identify which one you report rather than labeling both simply “WER.”
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
Also disclose the reference tokenization and text normalization: for example, how punctuation, capitalization, disfluencies and unfinished utterances are handled. There is no universal convention established across the cited brain-to-text work, so use the protocol actually applied in the study.
What must be reported with a WER?
A score needs enough context for readers to understand what was tested and reproduce the comparison. Report the following:
Rank #2
- Participants and population: number of participants, whether results are individual or aggregated, and relevant diagnosis or speech status when reported.
- Speech task: attempted, overt or imagined speech; prompted or conversational content; and whether the task was open-loop or closed-loop.
- Language and vocabulary: language, vocabulary size, prompt construction, and any language-model vocabulary constraints. State whether test text or prompts appeared during training.
- Test split and timing: what was held out—sentences, trials, sessions, days or participants—and what calibration data were available at test time.
- Recording and decoding pipeline: neural recording setup, intermediate representations such as phonemes or characters, decoder, vocabulary constraints, language-model rescoring or beam search, and final text output.
- Scoring details: reference tokenization, normalization, exclusions, pooled or sentence-averaged aggregation, and treatment of partial utterances.
- Sample size and uncertainty: number of test trials and reference words, point estimate, and confidence interval with its method.
- Practical outcomes: communication rate, latency, correction burden and error types where available.
These details separate performance of the neural decoder from performance of the complete system. A language model can change the final text, so a WER difference cannot automatically be attributed to neural decoding alone.
Can you compare WER across different brain-to-text studies?
Yes, but only as a qualified comparison. Before ranking scores, check whether the studies align on participant population, speech task, vocabulary, held-out data, calibration, full decoding pipeline and metric protocol. If they do not, describe the results as outcomes under different conditions, not direct head-to-head evidence that one system is better.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
Vocabulary illustrates the problem. A 2023 Nature neuroprosthesis paper reports 9.1% WER for a 50-word vocabulary and 23.8% for a 125,000-word vocabulary in its study (Nature, 2023). These figures belong to different vocabulary settings; they do not, by themselves, isolate vocabulary size as the cause of the difference.
Other published results are similarly tied to their evaluation:
Rank #4
| Reported result | What it describes |
|---|---|
| 0.44% WER over 50 evaluation sentences | A 2023 medRxiv report’s initial closed-loop session, following 213 training sentences in a 50-word-vocabulary setting; it is not evidence of broad-vocabulary or cross-participant performance. medRxiv, 2023 |
| 5.77% WER versus 8.93% | A 2025 PubMed-indexed article’s Brain-to-Text ’24 comparison: fine-tuned language model versus the leading benchmark method in that paper. It illustrates the role of language-model design in final WER, not a universal leaderboard ranking. PubMed, 2025 |
| 24.69% to 10.22% end-to-end WER | The improvement reported for BIT against a prior end-to-end method under the ICLR 2026 paper’s evaluation. Interpret it within that paper’s benchmark and task conditions, including its discussion of transfer between attempted and imagined speech. ICLR, 2026 |
These examples are useful for understanding what published systems have reported, but their tasks, protocols and pipelines differ. The evidence here does not establish the current official benchmark leader or a single authoritative benchmark normalization rule; consult the relevant challenge organizers’ documentation before making official leaderboard claims.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What should accompany WER?
WER treats every word edit equally. It does not indicate whether an error changes the intended meaning, whether common words account for most correct predictions, or how quickly a user can communicate. Complement it with measures that address the research question:
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
- Learn about your brainwaves, train your meditation, and develop your own applications with the mindwave mobile wireless headset.
- Bt/ble Dual mode module and support iOS, Android, PC, and Mac platform. Detects raw-brainwaves, eeg power spectrums (Alpha, beta, etc.), esense meters for attention, meditation, and future algorithms.
- More than 100 brain training games and educational apps available from the NeuroSky online store. Uses a single AAA battery (not included) for 8-hour battery run time
- Phoneme error rate (PER) and character error rate (CER): show performance at smaller sound or text units. They answer different questions from WER and should not be combined into a single score.
- Words per minute: gives a measure of communication throughput, which a low WER cannot convey on its own.
- Word-level error analysis: examine error types, word frequency and semantic impact when the goal is usable communication. A 2025 Interspeech study introduced refined word-level alignment and four additional metrics for exact correctness and semantic distance; it reported frequency-related performance disparities and greater semantic cost for errors on infrequent words. Interspeech, 2025
When reporting practical performance, pair these outcomes with latency and the amount of user correction required if those data are available. A lower WER does not necessarily mean faster or more useful communication.
A concise reporting checklist
For a results table or paper, include the WER and enough detail to interpret it:
- participant count and test-trial count;
- test-set reference word count;
- task, vocabulary and language context;
- held-out split and calibration conditions;
- recording setup and complete decoder-to-text pipeline;
- normalization, tokenization and aggregation method;
- uncertainty interval and how it was calculated;
- PER, CER, rate, latency, correction burden or word-level analysis when relevant.
For example, a 2026 bioRxiv preprint reports pooling errors across trials and dividing by total target words, with confidence intervals estimated using 10,000 bootstrap resamples of individual trials. That is a specific, reproducible choice—not a universal requirement for every study. bioRxiv, 2026
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

