PARTNR
100,000 household chores that showed robot planners slow humans down
Latest: PARTNR (Nov 2024)
The largest benchmark for human-robot collaboration: 100,000 natural-language household tasks across 60 simulated houses and 5,800+ objects, built on Habitat 3.0. Released with a dataset of human demonstrations, it evaluates whether LLM-driven robot planners can coordinate with people on chores like tidying and cooking.
Why it matters
PARTNR is the largest human-robot collaboration benchmark — 100,000 language-specified household tasks across 60 houses and 5,800+ objects — and its finding was humbling: state-of-the-art LLM planners coordinated so poorly that they often slowed the human down. It gave embodied-AI research a concrete target for genuinely helpful robots.
Facts
- Findings were humbling: state-of-the-art LLM planners coordinated poorly with human partners, often slowing the human down rather than helping — exactly the gap the benchmark was built to expose.
Try it yourself
Code on GitHub ↗ Episodes dataset on Hugging Face ↗ Read the paper ↗
Lineage
Sources
arXiv ↗GitHub · partnr-planner ↗