OpenEQA
1,600 questions about real spaces — top VLMs were 'nearly blind'
Latest: OpenEQA (Apr 2024)
FAIR's open-vocabulary Embodied Question Answering benchmark (CVPR 2024): over 1,600 human-authored questions about 180+ real homes and offices, testing whether an agent that has 'seen' a space can answer questions about it — from episodic memory or by actively exploring. Includes an automatic LLM-based scorer validated against humans.
Why it matters
OpenEQA exposed how little vision-language models actually see: on questions about real homes and offices, GPT-4V scored 48.5% against 85.9% for humans, and on spatial questions barely beat text-only guessing. Its 1,600+ human-written questions became the reference test for the memory-equipped smart-glasses assistants Meta is building toward.
Facts
- Headline result: on spatial questions, top VLMs were 'nearly blind' — access to visual input barely beat language-only guessing, while humans scored far higher.
- Framed by Meta as a milestone toward smart-glasses assistants that remember your world.
Try it yourself
Project page ↗ Dataset and code on GitHub ↗ Read the paper (PDF) ↗
Sources
Meta AI blog ↗open-eqa.github.io ↗