metaai·lightalo unofficial · independent
Universe / Safety & Trust / GAIA Benchmark
Safety & Trust · 2023

GAIA Benchmark

466 questions: humans scored 92%, GPT-4 with plugins scored 15%

open source

Latest: GAIA (November 2023); public leaderboard maintained on Hugging Face

A benchmark for General AI Assistants, co-created by Meta FAIR (Grégoire Mialon, Yann LeCun, Thomas Scialom) with Hugging Face. Its 466 real-world questions require reasoning, web browsing, multi-modality and tool use — conceptually simple for humans, brutal for AIs — and it became the standard yardstick for agentic systems.

Why it matters

The standard yardstick for agentic AI: 466 real-world questions requiring reasoning, browsing, multi-modality and tool use that are simple for humans and hard for models. Co-created by FAIR and Hugging Face, its public leaderboard became the benchmark agent frameworks report, from early AutoGPT-style systems to today's frontier agents.

Facts

Try it yourself

Sources

More in Safety & Trust

Purple LlamaCyberSecEvalLlamaFirewall

Read the Safety & Trust story on the sky →

✦ Open on the map Explore Safety & Trust Quiz me