GAIA Benchmark
466 questions: humans scored 92%, GPT-4 with plugins scored 15%
Latest: GAIA (November 2023); public leaderboard maintained on Hugging Face
A benchmark for General AI Assistants, co-created by Meta FAIR (Grégoire Mialon, Yann LeCun, Thomas Scialom) with Hugging Face. Its 466 real-world questions require reasoning, web browsing, multi-modality and tool use — conceptually simple for humans, brutal for AIs — and it became the standard yardstick for agentic systems.
Why it matters
The standard yardstick for agentic AI: 466 real-world questions requiring reasoning, browsing, multi-modality and tool use that are simple for humans and hard for models. Co-created by FAIR and Hugging Face, its public leaderboard became the benchmark agent frameworks report, from early AutoGPT-style systems to today's frontier agents.
Facts
- At release, human respondents scored 92% while GPT-4 with plugins managed 15% — the gap that launched a thousand agent frameworks.
- Climbing the GAIA leaderboard became a rite of passage for 2024-2026 agent startups.
Try it yourself
Public leaderboard (Hugging Face Space) ↗ Dataset on Hugging Face ↗ Read the paper ↗