Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In one India-based experiment, switching a buying question from English to romanized Hinglish changed the brands returned by AI search systems—but the size of the difference varied sharply. Perplexity showed the largest language-related change, Gemini showed a substantial one, and ChatGPT’s smaller result was fragile. The experiment measured outputs; it did not uncover a universal recommendation formula or explain why the systems behaved differently.
What the experiment tested
Depra Research’s IndicGEO Study 01 compared how three answer systems responded to equivalent shopping questions in English and romanized Hinglish. One example was “best face wash for oily skin under ₹500”; its Hinglish counterpart was “India mein oily skin ke liye 500 rupees ke andar konsa face wash kharidna chahiye?” These are examples from the study, not a representative survey of user questions.
The researchers used 10 unbranded buying prompts—five skincare and five fashion—and independently reviewed the English-Hinglish pairs for equivalence. Three prompt pairs were rewritten after review and before analysis. Each prompt was run eight times in each language on ChatGPT with web search, Gemini, and Perplexity Sonar with live web search. All requests were India-geolocated.
Collection took place in one randomized, interleaved 57-minute window on 14 August 2026, producing 480 responses. ChatGPT and Gemini were captured from their consumer web apps; Perplexity was queried through the Sonar API answer model. The results therefore describe those products and capture conditions, not every way someone might use the services.
#1 Best Overall
How the study distinguished language effects from answer variation
Repeated runs of the same prompt can produce different answers even when the wording does not change. To account for that, the study compared the overlap between brand lists for English-Hinglish prompt pairs with the overlap between repeated runs in the same language. It used Jaccard overlap: the number of shared brands divided by the number of brands appearing in either list.
This comparison asks whether the language-pair difference exceeded ordinary same-language rerun variation in this dataset. It does not identify the internal cause. Retrieval, available sources, model generation, and other factors could all contribute, but the design does not isolate them.
Which systems changed their brand lists most?
| System | Reported language-related difference beyond rerun variation | One-sided permutation-test p-value | How to read the result |
|---|---|---|---|
| ChatGPT | 2.2 percentage points | .046 | Fragile; the two-sided p-value was .092, four of 10 prompt-level deltas were negative, and the effect was carried by skincare. The authors characterize ChatGPT as approximately language-invariant. |
| Gemini | 7.8 percentage points | .0098 | A larger measured language-related difference than ChatGPT in this set. |
| Perplexity Sonar | 23.4 percentage points | .0010 | The largest measured language-related difference in this set. |
All three one-sided tests cleared the study’s pre-set test, but that does not make the three estimates equally robust. In particular, the small ChatGPT figure depends on a one-sided test and is qualified by the study’s own sensitivity checks. The percentages describe the study’s overlap comparison, not the share of users who would see different recommendations.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsPerplexity also had the highest same-language rerun overlap among the three systems. That means its answers were comparatively consistent across repeated runs of the same-language prompts, while the English-Hinglish lists still differed more than rerun variation. Stability on reruns is not the same as similarity to another engine.
What changed in citations and brand mentions
Gemini linked fewer Hinglish answers to citations
Gemini included at least one citation in 72 of 80 English responses (90.0%) and 33 of 80 Hinglish responses (41.3%). Its average number of citations per answer fell from 9.4 in English to 5.2 in Hinglish. The study records the difference but does not establish its cause: as the report puts it, “The study records the drop. It does not establish the cause, because the design cannot separate how Gemini searches from how it writes.” Displayed citations identify sources linked with responses; they do not explain why a particular brand was recommended.
One Perplexity brand had a notable swing
The Derma Co appeared in 2 of 80 English Perplexity answers and 16 of 80 Hinglish answers, an increase of 17.5 percentage points in mention share. This is an example from the tested prompts, not a general measure of the brand’s performance or visibility.
Rank #4
Citation rankings depend on what is counted
Across the corpus, YouTube had 534 citation appearances, Reddit had 191, and Nykaa had 176. Counting answers that cited a domain at least once gives a different ordering: Reddit appeared in 106 of 480 answers, while YouTube appeared in 100 of 480. A citation appearance can occur more than once in an answer, so appearances and answers citing a domain are distinct measures.
What this says—and does not say—about AI recommendations
The results support a limited conclusion: for these 10 India-focused skincare and fashion questions, equivalent English and Hinglish wording could produce different brand-list overlap, and the measured gap varied by system. They do not show that a platform follows a fixed language-based ranking rule, that one language is always favored, or that the same differences apply to other categories, regions, prompts, or dates.
Best Value
- The study covered two product categories, 10 prompts, India-geolocated requests, and one collection period.
- It used logged-out usage; results for signed-in users or other settings are not established.
- Perplexity was queried through the Sonar API, whose behavior may differ from the consumer app.
- The study did not isolate whether differences came from search and retrieval, source availability, answer generation, or another mechanism.
- Depra, the study publisher, sells a related visibility-monitoring service, a relevant commercial interest when weighing its findings.
Depra also published a re-analysis of the same corpus about differences between engines’ cited websites and brand lists. It is a vendor analysis of the same data, not an independent replication.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to check brand visibility without overreading one answer
A single output is not enough to establish a trend. As Ifham Baig, Depra AI’s co-founder and CTO, advises in his DEV Community article, “A single manual check tells you almost nothing.” That is the author’s recommendation, rather than an independently tested finding.
For a useful visibility check, repeat the specific questions and conditions that matter to the audience you are trying to understand:
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →- Match the audience’s language and intent. Test equivalent prompts in the languages people use, without changing the product need or constraints between versions.
- Repeat each prompt. Multiple runs help reveal whether a brand-list change is larger than the system’s ordinary answer variation.
- Keep engines and access methods distinct. Record whether a response came from a consumer app or an API; do not treat those capture methods as interchangeable.
- Track more than one outcome. Compare brand-list overlap, same-engine rerun consistency, citation frequency and source mix, and how often a brand appears first.
- Record the conditions and date. Geography, language, prompt wording, product category, and collection period define what the observation can support.
These measures answer different questions. A high rerun overlap indicates consistency within one system; it does not establish that its recommendations match another system. Frequent citations indicate linked sources, not the reason a brand was selected. A first-brand repeat rate describes position in the sampled answers, not a universal rank.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

