Ego-Exo4D
Cooking, bike repair, bouldering — seen first- and third-person at once
Latest: Ego-Exo4D v2 (2024; ~1,300 hours)
A multimodal dataset pairing time-synced first-person (Aria glasses) and third-person video of skilled activities — cooking, bike repair, soccer, bouldering, dance, music — from 800+ participants in 13 cities. Built by FAIR, Project Aria, and 15 university partners; v2 spans about 1,300 hours across 5,035 captures with expert commentary annotations.
Why it matters
Ego-Exo4D pairs synchronized first-person Aria and third-person video of skilled activities — cooking, bike repair, bouldering, music — from 800+ participants in 13 cities, with gaze, 7-channel audio and expert coach commentary. It is the dataset for AI that could one day teach physical skills through AR glasses.
Facts
- Every capture includes eye gaze, 7-channel audio, and expert 'coach' commentary — the dataset was explicitly designed for AI that could one day teach humans physical skills through AR glasses.
Try it yourself
Dataset site ↗ Documentation and download ↗ Read the paper ↗
Lineage
Sources
arXiv ↗Meta AI blog ↗ego-exo4d-data.org ↗