metaai·lightalo unofficial · independent
Universe / World Models & Embodied AI / OpenEQA
World Models & Embodied AI · 2024

OpenEQA

1,600 questions about real spaces — top VLMs were 'nearly blind'

open source

Latest: OpenEQA (Apr 2024)

FAIR's open-vocabulary Embodied Question Answering benchmark (CVPR 2024): over 1,600 human-authored questions about 180+ real homes and offices, testing whether an agent that has 'seen' a space can answer questions about it — from episodic memory or by actively exploring. Includes an automatic LLM-based scorer validated against humans.

Why it matters

OpenEQA exposed how little vision-language models actually see: on questions about real homes and offices, GPT-4V scored 48.5% against 85.9% for humans, and on spatial questions barely beat text-only guessing. Its 1,600+ human-written questions became the reference test for the memory-equipped smart-glasses assistants Meta is building toward.

Facts

Try it yourself

Sources

More in World Models & Embodied AI

V-JEPAPARTNRMeta MotivoV-JEPA 2I-JEPAVL-JEPA

Read the World Models & Embodied AI story on the sky →

✦ Open on the map Explore World Models & Embodied AI Quiz me