Generative AI can feel uncanny when an image or interaction invites us to read it as humanlike but does not quite hold together. That is one possible explanation, not a universal rule: studies have measured different responses to different images, faces and chatbots, and realism does not produce the same reaction in every test.
What does “uncanny valley” mean for generative AI?
The uncanny valley is a proposed pattern in which a nearly humanlike image or character feels less familiar or more unsettling than something clearly artificial. Applied to generative AI, the phrase can describe the “almost real, but somehow off” response to an image—or the discomfort of talking with a system that seems humanlike in some ways but not others.
It is more useful to treat uncanniness as a response measured through ratings such as eeriness, familiarity, liking or trust than as a fixed property of an image or model. A study that asks whether people can identify synthetic faces is measuring something different from one that asks how strange an image feels.
What do studies of AI-generated images and faces show?
The evidence does not describe one consistent path from obviously artificial to photorealistic. The studies below differ in their images, methods, sample sizes and outcomes, so their findings should not be read as interchangeable.
#1 Best Overall
| Study and material | What participants did | What the study found |
|---|---|---|
| Rapp and colleagues, 2025; 20 Stable Diffusion images | Explored participants’ perceptions, appraisals and emotions about generated images. | Participants described images in terms of technical quality and fidelity; reactions ranged from seeing them as prototypical to finding them strange. This qualitative study does not estimate how often people generally find AI images uncanny. Read the study. |
| Kishnani, MIT master’s thesis, February 2025; Stable Diffusion XL images | In a separate image task, 56 participants assessed outputs spanning different levels of realism. | Highly realistic or clearly stylized images raised fewer concerns than intermediate-realism images in this selected set. The small sample and chosen model limit how broadly to apply the result. Read the thesis. |
| Nightingale and Farid, PNAS, 2022; StyleGAN2 faces | One experiment tested classification of real versus synthetic faces; another asked a separate group to rate trustworthiness. | In the classification experiment, 315 participants averaged 48.2% accuracy, close to the 50% chance level. In the trust-rating experiment, 223 participants gave real faces an average 4.48 and synthetic faces 4.82 on a seven-point scale; the authors reported the synthetic faces were rated 7.7% more trustworthy in that experiment. Read the paper. |
The PNAS result does not establish that AI faces are generally more trustworthy, or that every viewer mistakes every synthetic face for a real person. It describes selected StyleGAN2 faces and particular tasks. The authors wrote of those experiments: “Synthetically generated faces are not just highly photorealistic, they are nearly indistinguishable from real faces and are judged more trustworthy.”
That finding can coexist with discomfort toward some other images. A face that is difficult to classify is not necessarily an image that a viewer finds eerie; difficulty identifying its source and emotional unease are separate outcomes.
Rank #2
Does the uncanny valley apply to AI chatbots?
There is preliminary evidence that conversational behavior can feel less human when its humanlike signals seem inconsistent. In a February 2025 MIT master’s thesis, 60 participants interacted across three text-agent conditions. The prompt-engineered “Uncanny-Valley Bot” received the lowest ratings for anthropomorphism, animacy, likeability and perceived intelligence among those conditions. The result concerns one engineered chatbot setup and short interactions, not all AI assistants or longer-term use. The thesis describes the experiment and its limitations.
For a chatbot, the cues are not visual realism but conversational ones: whether its language, apparent personality and responses sustain the impression of a coherent humanlike agent. The thesis offers one controlled test of that idea; it does not establish a universal chatbot “valley” or identify one behavior that reliably causes it.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteWhy can a realistic image still feel off?
One candidate explanation is a mismatch among realism cues. An image may look convincing in some respects while other details seem inconsistent, making it harder to interpret as a coherent person or animal. A 2015 Cognition experiment found that reducing consistency across selected visual realism features increased eeriness and coldness for human and animal depictions. In that experiment, increasing category uncertainty did not produce the predicted effect. This supports feature mismatch as one possible factor, not a complete explanation of reactions to current image generators. Read the study.
Viewers’ appraisals can also involve more than whether a picture looks human. In their exploration of Stable Diffusion outputs, Rapp and colleagues found that participants considered technical quality and fidelity, and that reactions could extend from an unsettling image to perceptions of the AI behind it. Participants also expressed awareness of societal bias. Those observations help describe how people interpret generated images; they do not show how common any particular reaction is.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Why do findings about realism seem to conflict?
“Realism” is not a single experimental condition. Intermediate-realism Stable Diffusion XL images in a small thesis experiment, selected photorealistic StyleGAN2 faces in a face-classification task, and other generated pictures differ in model, stimulus selection, visual cues and what participants were asked to judge. A result about eeriness cannot be substituted for one about trust or identification accuracy.
Stimulus selection matters even outside generative AI. In six studies involving 1,343 participants, Palomäki and colleagues found that the uncanny-valley effect did not replicate with some non-photorealistic CGI morph stimuli, but was prominent with pre-evaluated photorealistic robot pictures. Their findings show why an effect seen with one set of images may not carry over to another. Read the replication study.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
The wider uncanny-valley literature is itself methodologically varied. A 2021 meta-analysis included 72 studies from 468 identified and examined 247 effect sizes, reporting a pooled Hedges’ g of 1.01 with a 95% interval of 0.80 to 1.22. That synthesis covers the included uncanny-valley literature broadly—not generative AI alone—and its authors noted substantial variety in stimulus techniques and outcome measures, with no consensus on theory and methodology. Read the meta-analysis.
Quick Recap
What can readers reasonably conclude?
- Some generated images or interactions can provoke an uncanny response, but the response depends on what is shown and what viewers are asked to judge.
- Visual mismatch is a plausible source of unease, while a highly realistic synthetic face can also be hard to distinguish from a real face in a particular experiment.
- Evidence directly about generative AI remains limited: the cited work includes selected model outputs, small samples, qualitative exploration and task-specific experiments.
- Claims about a single uncanny-valley curve for all images, models, chatbots or viewers go beyond what these studies establish.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

