Meta Locate 3D
'The flower vase near the TV console' — located in 3D
Latest: Locate 3D (Apr 2025)
An end-to-end model (April 2025) that localizes objects in 3D scenes from natural-language queries like 'the flower vase near the TV console' — operating directly on point clouds from RGB-D sensor streams via a 3D-JEPA encoder and language-conditioned decoder that outputs 3D bounding boxes and masks, ready for real robots.
Why it matters
Locate 3D turns a sentence — 'the flower vase near the TV console' — into a 3D box and mask directly from RGB-D point clouds, using a self-supervised 3D-JEPA encoder, so robots can ground language in real scenes. Its 130,000-annotation dataset roughly doubled the world's supply of 3D referring expressions.
Facts
- Shipped with a new 130,000-annotation referring-expression dataset across ARKitScenes, ScanNet, and ScanNet++ (1,346 scenes) — roughly doubling the world's supply of such 3D language annotations.
Try it yourself
Official Locate 3D demo ↗ locate-3d on Hugging Face ↗ Code on GitHub ↗