iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Brașov, with ș (s with a comma below), is the correct Romanian spelling. Braşov uses a different character, ş (s with a cedilla). A 2026 benchmark by Daniel Butnaru reported that 15 language models generally produced the correct spelling from clean input but often copied the cedilla form after it appeared in the prompt. That points to context copying in this test—not proof that every model always confuses the characters.
Why Brașov uses ș, not ş
Romanian orthography uses a comma below for the letters ș and ț. Academia Română explicitly distinguishes the comma from the cedilla, noting “virgulița (și nu sedila, aşa cum se preciza în lucrări mai vechi)”—“the comma below (and not the cedilla, as was specified in older works).” Its guidance supports writing the city’s name as Brașov.
The difference is not merely visual. In Unicode, the correct Romanian ș is U+0219, while the cedilla ş is U+015F. Some fonts make them look nearly alike, but software can treat them as distinct characters. An exact text comparison or search for one form may not match the other.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallFor public-sector documents in Romania, Ministry Order 414/2006 requires public authorities, public institutions, and notaries to use the Romanian character set defined by the Academy’s orthographic guidance and gives encoding recommendations for electronic documents. That requirement should not be generalized to every private publisher.
#1 Best Overall
What the 15-model benchmark found
In his 2026 write-up, Daniel Butnaru describes testing 15 models through Kaggle’s model proxy using its defaults. The tasks included generating Romanian text from ASCII place names, echoing or summarizing correct and cedilla-typed text, proofreading with different instructions, comparing Unicode sequences, restoring diacritics from context, and writing a Romanian text-fixing function. The author says scores were computed from Unicode code points, without an LLM judge; the echo task was repeated three times per model.
The central contrast was between clean and contaminated input. Butnaru reports that all tested models retained the correct letters in 45 runs when the input was correctly spelled. With cedilla forms in the input, reported clean-answer rates varied from 44% to 85% across models. The results suggest that models could generate or preserve the right Romanian characters, yet often echoed the wrong form when it was already present in context.
Rank #2
- Used Book in Good Condition
Task type changed the result
For cedilla-typed input, the author reports clean answers in 97% of replies, 62% of summaries, and 45% of retrieval-style answers. In this benchmark, 1,105 of 1,186 incorrect words—93%—were copied verbatim from the input. That pattern makes the task important: a model asked to repeat or retrieve text may preserve its spelling rather than silently edit it.
Explicit proofreading instructions helped
Butnaru reports correction rates of 79% for “Proofread the following text,” 99% for “Correct the spelling and the diacritics,” and 97% when the prompt explicitly described the cedilla issue. In a separate follow-up, adding a Romanian orthography rule produced clean-answer rates ranging from 64% to 96%, with ten of the 15 models perfect on that task. These are outcomes from the author’s prompts and setup, not guaranteed performance in another product or deployment.
How broad are these results?
They describe one author’s benchmark, not a population-wide estimate of how language models behave. The model versions, proxy defaults, prompts, and run configuration all matter; model names and availability also change. The article names 15 models from Anthropic, Google, OpenAI, and other providers, but its dated list should not be read as a current-market ranking.
Butnaru also reports that 30 of 35 Romanian website front pages fetched on 30 September 2026 contained cedilla forms. That is a small, dated sample, not a representative estimate of Romanian websites or the wider web. The author’s own summary of the clean-input condition was: “When the guest message or news paragraph was typed correctly, every model kept every letter correct, in all 45 runs.” The scope is those tested runs, not every model or use case.
Rank #4
The article also notes that a character-comparison task needed an identical-string control to interpret its results. Taken together, the evidence supports a narrower conclusion: in this benchmark, cedilla spellings in context often carried through to model output, while clean input was handled correctly in the reported runs. It does not establish that models cannot distinguish the two Unicode characters.
How to avoid carrying the wrong form into your output
State the orthography you want
When asking a model to edit Romanian text, specify that it must correct both spelling and diacritics, and name the Romanian forms ș and ț with commas below. A generic request to proofread may be less reliable in this benchmark than an explicit instruction.
Normalize text carefully in software
If your application needs consistent Romanian text, normalize known legacy cedilla forms as part of a tested text-processing pipeline. Unicode NFC normalization can handle canonically equivalent encodings, but it does not by itself convert U+015F into U+0219; that requires an explicit character mapping. Test both precomposed and decomposed input. Avoid stripping all diacritics or replacing every cedilla indiscriminately: marks can distinguish words, and cedillas are valid in other languages.
Check exact matches and search behavior
If a name search or equality check misses Brașov, inspect the actual code points rather than relying on how the glyph looks. A system that needs to match legacy text can compare after a deliberate normalization-and-mapping step while retaining the original text for display where appropriate.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →

