Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Meta’s OpenEQA benchmark found a sharp gap between people and the tested AI system: in results announced on April 11, 2024, GPT-4V scored 48.5%, compared with 85.9% for humans. Meta called vision-language models “nearly blind” specifically on questions requiring spatial understanding—not across every task involving images. The finding is a 2024 benchmark result, not a measure of the strongest models available in 2026.

What OpenEQA tests

OpenEQA—short for Open Embodied Question Answering—is a benchmark developed by researchers at Fundamental AI Research (FAIR), Meta. It tests whether an AI agent can answer open-ended questions about a particular physical environment and its contents. Unlike a general-knowledge question, a prompt such as “Where did I leave my badge?” depends on information about a specific place.

The benchmark contains more than 1,600 human-generated question-answer pairs drawn from more than 180 real-world environments. Meta says different human annotators checked whether questions could be answered and whether the supplied answers were correct. The authors describe OpenEQA as the first open-vocabulary EQA benchmark to support both episodic memory and active exploration.

How episodic-memory and active EQA differ

The two settings distinguish between recalling earlier observations and seeking new information. Meta uses smart glasses and mobile robots to illustrate the respective contexts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Setting Where answer context comes from Example device context Capability being probed
Episodic-memory EQA The agent’s memory of earlier observations Smart glasses Remembering and retrieving facts about an observed place
Active EQA New information gathered through exploration or action Mobile or home robot Choosing how to gather missing information, then answering

For example, remembering where a person left a badge fits the first setting if the agent previously observed it. A robot that moves around to check where the badge is fits the second. The benchmark therefore probes not only what a model can recognize, but whether it can use observations over time or gather information it still needs.

What “nearly blind” means in Meta’s finding

Meta’s FAIR research team wrote that “for questions that require spatial understanding, even the best VLMs are nearly ‘blind’”. The phrase refers to a specific weakness in the tested systems: on spatial questions, vision-language models did not perform much better than text-only models. Meta suggested that they may rely on language-based expectations instead of extracting useful spatial relationships from visual observations.

Meta illustrated the issue with: “I’m sitting on the living room couch watching TV. Which room is directly behind me?” The announcement says model guesses varied essentially at random. The point is that recognizing objects in an image is not the same as reliably understanding how rooms or objects are arranged relative to one another.

The claim is not that images never helped. The OpenEQA project page says multimodal models consistently outperformed text-only baselines on episodic-memory EQA and did particularly well on object localization and recognition; visual inputs also helped on some world-knowledge questions. Other categories remained closer to the blind GPT-4 baseline. Thus, the authors’ “nearly blind” characterization applies to spatial understanding and reflects a broader grounding problem in some benchmark categories—not a universal absence of visual benefit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to read the 48.5% and 85.9% scores

Meta reported 48.5% for GPT-4V and 85.9% for human performance in its April 11, 2024 announcement. Those figures show a substantial gap on the benchmark as evaluated and reported at that time. They should not be presented as current 2026 scores or as a ranking of today’s models: the cited announcement and project page do not establish a fresh 2026 rerun or comparison.

Because OpenEQA answers are open-ended, a response can be correct in multiple phrasings. The authors use an LLM-powered evaluation protocol called LLM-Match to judge correctness. Meta says blind user studies found its correlation with people comparable to agreement between two humans. That is the authors’ reported validation of the scoring method, not an independently established guarantee that every automatic judgment matches human assessment.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where to find the paper and benchmark

The work, titled OpenEQA: Embodied Question Answering in the Era of Foundation Models, lists Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, and collaborators as authors and is cited as a CVPR 2024 paper. The OpenEQA project page links to the paper, code, and benchmark. Meta’s original explanation and reported figures are in its April 11, 2024 announcement.

The Facebook Research repository documents dataset files, baselines, and a GPT-4 evaluation script. GitHub marks the repository archived on November 1, 2025, so it is read-only. The repository identifies the release as MIT-licensed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.