Meta’s OpenEQA benchmark found a sharp gap between people and the tested AI system: in results announced on April 11, 2024, GPT-4V scored 48.5%, compared with 85.9% for humans. Meta called vision-language models “nearly blind” specifically on questions requiring spatial understanding—not across every task involving images. The finding is a 2024 benchmark result, not a measure of the strongest models available in 2026.
What OpenEQA tests
OpenEQA—short for Open Embodied Question Answering—is a benchmark developed by researchers at Fundamental AI Research (FAIR), Meta. It tests whether an AI agent can answer open-ended questions about a particular physical environment and its contents. Unlike a general-knowledge question, a prompt such as “Where did I leave my badge?” depends on information about a specific place.
The benchmark contains more than 1,600 human-generated question-answer pairs drawn from more than 180 real-world environments. Meta says different human annotators checked whether questions could be answered and whether the supplied answers were correct. The authors describe OpenEQA as the first open-vocabulary EQA benchmark to support both episodic memory and active exploration.
How episodic-memory and active EQA differ
The two settings distinguish between recalling earlier observations and seeking new information. Meta uses smart glasses and mobile robots to illustrate the respective contexts.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
| Setting | Where answer context comes from | Example device context | Capability being probed |
|---|---|---|---|
| Episodic-memory EQA | The agent’s memory of earlier observations | Smart glasses | Remembering and retrieving facts about an observed place |
| Active EQA | New information gathered through exploration or action | Mobile or home robot | Choosing how to gather missing information, then answering |
For example, remembering where a person left a badge fits the first setting if the agent previously observed it. A robot that moves around to check where the badge is fits the second. The benchmark therefore probes not only what a model can recognize, but whether it can use observations over time or gather information it still needs.
What “nearly blind” means in Meta’s finding
Meta’s FAIR research team wrote that “for questions that require spatial understanding, even the best VLMs are nearly ‘blind’”. The phrase refers to a specific weakness in the tested systems: on spatial questions, vision-language models did not perform much better than text-only models. Meta suggested that they may rely on language-based expectations instead of extracting useful spatial relationships from visual observations.
Meta illustrated the issue with: “I’m sitting on the living room couch watching TV. Which room is directly behind me?” The announcement says model guesses varied essentially at random. The point is that recognizing objects in an image is not the same as reliably understanding how rooms or objects are arranged relative to one another.
The claim is not that images never helped. The OpenEQA project page says multimodal models consistently outperformed text-only baselines on episodic-memory EQA and did particularly well on object localization and recognition; visual inputs also helped on some world-knowledge questions. Other categories remained closer to the blind GPT-4 baseline. Thus, the authors’ “nearly blind” characterization applies to spatial understanding and reflects a broader grounding problem in some benchmark categories—not a universal absence of visual benefit.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRank #2
How to read the 48.5% and 85.9% scores
Meta reported 48.5% for GPT-4V and 85.9% for human performance in its April 11, 2024 announcement. Those figures show a substantial gap on the benchmark as evaluated and reported at that time. They should not be presented as current 2026 scores or as a ranking of today’s models: the cited announcement and project page do not establish a fresh 2026 rerun or comparison.
Because OpenEQA answers are open-ended, a response can be correct in multiple phrasings. The authors use an LLM-powered evaluation protocol called LLM-Match to judge correctness. Meta says blind user studies found its correlation with people comparable to agreement between two humans. That is the authors’ reported validation of the scoring method, not an independently established guarantee that every automatic judgment matches human assessment.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Where to find the paper and benchmark
The work, titled OpenEQA: Embodied Question Answering in the Era of Foundation Models, lists Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, and collaborators as authors and is cited as a CVPR 2024 paper. The OpenEQA project page links to the paper, code, and benchmark. Meta’s original explanation and reported figures are in its April 11, 2024 announcement.
The Facebook Research repository documents dataset files, baselines, and a GPT-4 evaluation script. GitHub marks the repository archived on November 1, 2025, so it is read-only. The repository identifies the release as MIT-licensed.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

