{
 "generated": "2026-09-03",
 "verified": "2026-09-03",
 "categories": {
  "muse": {
   "label": "The Muse Era",
   "hue": 275
  },
  "language": {
   "label": "Language & LLMs",
   "hue": 210
  },
  "vision": {
   "label": "Vision",
   "hue": 175
  },
  "world": {
   "label": "World Models & Embodied AI",
   "hue": 140
  },
  "speech": {
   "label": "Speech & Sound",
   "hue": 25
  },
  "genmedia": {
   "label": "Generative Media",
   "hue": 330
  },
  "science": {
   "label": "AI for Science",
   "hue": 95
  },
  "oss": {
   "label": "Open Source Infra",
   "hue": 45
  },
  "games": {
   "label": "Games & Strategy",
   "hue": 355
  },
  "safety": {
   "label": "Safety & Trust",
   "hue": 250
  },
  "edge": {
   "label": "On-Device & Silicon",
   "hue": 190
  },
  "reality": {
   "label": "Reality Labs Research",
   "hue": 300
  }
 },
 "projects": [
  {
   "id": "msl",
   "name": "Meta Superintelligence Labs",
   "full_name": "Meta Superintelligence Labs (MSL)",
   "tag": "The 2025 reboot: a superintelligence lab built around a $14.3B talent bet",
   "year": 2025,
   "year_label": "2025",
   "cat": "muse",
   "size": 3,
   "status": "active",
   "open": false,
   "lineage": [],
   "desc": "Meta's umbrella AI organization formed June 30, 2025 after its $14.3B investment for 49% of Scale AI. Chief AI Officer Alexandr Wang leads frontier-model group TBD Lab; Nat Friedman leads Products and Applied Research; FAIR and MSL Infra complete the four divisions. MSL built the entire Muse model family that replaced Llama as Meta's frontier line.",
   "facts": [
    "Poached ChatGPT co-creator Shengjia Zhao (chief scientist, July 2025) and at least 8 OpenAI researchers with reported nine-figure offers.",
    "Four divisions set Aug 2025.",
    "Cut 600 AI jobs Oct 22, 2025.",
    "Yann LeCun announced his exit Nov 19, 2025 (leaving at year's end) to found his own Advanced Machine Intelligence startup.",
    "A ~6,500-person Applied AI engineering unit stood up in Mar 2026 under CTO Andrew Bosworth, partnering with MSL."
   ],
   "latest": "Applied AI engineering unit added under Maher Saba (Mar 2026)",
   "links": [
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs"
    },
    {
     "label": "Wikipedia",
     "url": "https://en.wikipedia.org/wiki/Meta_Superintelligence_Labs"
    },
    {
     "label": "Axios",
     "url": "https://www.axios.com/2025/10/22/meta-superintelligence-tbd-ai-reorg"
    },
    {
     "label": "CNBC",
     "url": "https://www.cnbc.com/2025/07/25/zuckerberg-shengjia-zhao-meta-ai-lab-chief-scientist-openai.html"
    }
   ],
   "why": "MSL is the organization behind Meta's 2025-26 pivot from open Llama to the closed Muse Spark line, assembled through the $14.3B Scale AI deal and nine-figure hires. Its research arm and Model API now decide what Meta ships to three billion users and to developers.",
   "try": [
    {
     "label": "MSL research blog",
     "url": "https://research.meta.ai/"
    },
    {
     "label": "Chat with MSL's Muse Spark at meta.ai",
     "url": "https://www.meta.ai"
    },
    {
     "label": "Read the Muse Spark launch post",
     "url": "https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs/"
    }
   ],
   "params": ""
  },
  {
   "id": "muse-spark",
   "name": "Muse Spark",
   "full_name": "Muse Spark",
   "tag": "Meta's first closed frontier model — and its comeback story",
   "year": 2026,
   "year_label": "2026",
   "cat": "muse",
   "size": 3,
   "status": "closed",
   "open": false,
   "lineage": [
    "llama4",
    "msl"
   ],
   "desc": "MSL's first frontier model and Meta's first closed-weight flagship, launched April 8, 2026 — a natively multimodal reasoning model with tool use, visual chain of thought, and multi-agent orchestration. Offers Instant, Thinking, and Contemplating modes (the last orchestrates parallel reasoning agents), and now powers Meta AI across apps reaching billions of users.",
   "facts": [
    "Contemplating mode scores 58% on Humanity's Last Exam and 38% on FrontierScience Research.",
    "Debuted 4th on Artificial Analysis (behind GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6) but led HealthBench Hard 42.8 vs Gemini 3.1 Pro's 20.6.",
    "Matches Llama 4 Maverick with over an order of magnitude less compute.",
    "Built with 1,000+ physician collaborators and trained on Meta's new AI superclusters.",
    "A 1.1 update shipped just three months later."
   ],
   "latest": "Muse Spark 1.2 (Aug 5, 2026); 1.1 added 1M-token context + agentic gains (Jul 9, 2026)",
   "links": [
    {
     "label": "Meta AI blog · introducing muse spark msl",
     "url": "https://ai.meta.com/blog/introducing-muse-spark-msl"
    },
    {
     "label": "Meta AI blog · introducing muse spark meta mo",
     "url": "https://ai.meta.com/blog/introducing-muse-spark-meta-model-api"
    },
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs"
    },
    {
     "label": "Wikipedia",
     "url": "https://en.wikipedia.org/wiki/Muse_Spark"
    },
    {
     "label": "TechCrunch",
     "url": "https://techcrunch.com/2026/04/08/meta-debuts-the-muse-spark-model-in-a-ground-up-overhaul-of-its-ai"
    },
    {
     "label": "CNBC",
     "url": "https://www.cnbc.com/2026/03/11/meta-ai-mtia-chip-data-center.html"
    }
   ],
   "why": "Muse Spark ended Meta's open-weight era and took the company from also-ran to credible frontier contender in one release, reportedly matching Llama 4 Maverick with over an order of magnitude less compute. Version 1.3 shipped September 2, 2026, using about 20% fewer tool calls and 25% fewer tokens on agentic coding.",
   "try": [
    {
     "label": "Use it via OpenRouter (Muse Spark 1.3)",
     "url": "https://openrouter.ai/meta/muse-spark-1.3"
    },
    {
     "label": "Meta Model API docs",
     "url": "https://dev.meta.ai/docs/"
    },
    {
     "label": "Chat at meta.ai",
     "url": "https://www.meta.ai"
    }
   ],
   "params": "undisclosed (closed weights)"
  },
  {
   "id": "muse-glimmer",
   "name": "Muse Glimmer",
   "full_name": "Muse Glimmer",
   "tag": "The open-source plot twist: a 30B agent that runs on your laptop",
   "year": 2026,
   "year_label": "2026",
   "cat": "muse",
   "size": 3,
   "status": "active",
   "open": true,
   "lineage": [
    "muse-spark"
   ],
   "desc": "A 30B-parameter open-weight multimodal agent model released August 10, 2026 under Apache 2.0 — Meta's return to open weights after Muse Spark went closed. Distilled from Muse Spark, it runs multi-step agent workflows (tool calls, coding, files, screenshots) entirely offline on a single consumer GPU or Mac, in 100+ languages.",
   "facts": [
    "Quantizes to under 20 GB.",
    "Speculative decoding (DFlash drafter) gives 1.5–3.1x faster generation.",
    "Optimized for MacBook M4-Max/M5-Max and RTX-5090 with AMD, Arm, Dell, Intel, NVIDIA as partners.",
    "Ships on Hugging Face, Ollama, LM Studio, llama.cpp, ExecuTorch, MLX, vLLM, SGLang.",
    "Launched alongside a public recommitment from Zuckerberg to open source, with an open-weight version of Muse Spark 1.2 promised next."
   ],
   "latest": "Muse Glimmer 30B (Aug 10, 2026)",
   "links": [
    {
     "label": "Hugging Face · Muse Glimmer 30B",
     "url": "https://huggingface.co/meta-models/Muse-Glimmer-30B"
    },
    {
     "label": "Hugging Face · muse glimmer",
     "url": "https://huggingface.co/blog/muse-glimmer"
    },
    {
     "label": "Meta Research",
     "url": "https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model"
    },
    {
     "label": "Meta for Developers",
     "url": "https://developer.meta.com/ai/models/muse-glimmer"
    },
    {
     "label": "CNBC",
     "url": "https://www.cnbc.com/2026/08/10/meta-muse-glimmer-open-weight-ai.html"
    },
    {
     "label": "TechCrunch",
     "url": "https://techcrunch.com/2026/08/10/metas-new-glimmer-ai-model-offers-a-hint-at-zuckerbergs-personal-intelligence-vision"
    },
    {
     "label": "Fortune",
     "url": "https://fortune.com/2026/07/09/meta-muse-spark-1-1-release-alexandr-wang-superintelligence-labs-mark-zuckerberg"
    }
   ],
   "why": "Glimmer is Meta's answer to the charge that it abandoned open source: a 30B Apache-2.0 agent model distilled from Muse Spark that runs offline on one consumer GPU and shipped day-one on Hugging Face, Ollama, llama.cpp and MLX. It re-established Meta as an open-weight player after Llama 4.",
   "try": [
    {
     "label": "Model on Hugging Face",
     "url": "https://huggingface.co/meta-models/Muse-Glimmer-30B"
    },
    {
     "label": "Run it locally with Ollama",
     "url": "https://ollama.com/library/muse-glimmer"
    },
    {
     "label": "Use it via OpenRouter",
     "url": "https://openrouter.ai/meta/muse-glimmer-30b"
    }
   ],
   "params": "30B (29.6B incl. a 1.8B ViT-G/14 perception encoder)"
  },
  {
   "id": "muse-image",
   "name": "Muse Image",
   "full_name": "Muse Image",
   "tag": "An image model that thinks before it draws",
   "year": 2026,
   "year_label": "2026",
   "cat": "muse",
   "size": 2,
   "status": "closed",
   "open": false,
   "lineage": [
    "emu",
    "muse-spark"
   ],
   "desc": "MSL's image generation and editing model, launched July 7, 2026. It behaves as an agent rather than a one-shot prompt-to-image mapper: it reasons over prompts, self-refines, composes from multiple reference photos, renders legible text, and can call web search and code execution. Available in Meta AI, meta.ai, Instagram Stories (US), and WhatsApp.",
   "facts": [
    "Ranked #2 on the Arena leaderboard for text-to-image, single-image editing, AND multi-image editing as of July 5, 2026.",
    "Self-refinement behavior emerged during RL training without being explicitly designed.",
    "Image quality improves log-linearly with test-time compute.",
    "Every output carries the 'Content Seal' invisible watermark that survives cropping, compression, resizing and screenshots (checker at meta.ai/identification)."
   ],
   "latest": "Muse Image (Jul 7, 2026)",
   "links": [
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/introducing-muse-image-muse-video-msl"
    },
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2026/07/introducing-muse-image-meta-ai"
    }
   ],
   "why": "Muse Image reframed image generation as an agent loop — reason, draw, critique, redraw — and its self-refinement emerged from RL rather than design, with quality scaling log-linearly with test-time compute. It reached #2 on Arena for text-to-image and both editing tracks and stamps a robust invisible watermark on every output.",
   "try": [
    {
     "label": "Check an image for the Content Seal watermark",
     "url": "https://www.meta.ai/identification"
    },
    {
     "label": "Generate via the Meta Model API ($0.01/image)",
     "url": "https://dev.meta.ai/docs/"
    },
    {
     "label": "Meta AI app on the App Store",
     "url": "https://apps.apple.com/us/app/meta-ai/id1558240027"
    }
   ],
   "params": "undisclosed (closed weights)"
  },
  {
   "id": "muse-video",
   "name": "Muse Video",
   "full_name": "Muse Video",
   "tag": "The Movie Gen lineage goes to production",
   "year": 2026,
   "year_label": "2026",
   "cat": "muse",
   "size": 2,
   "status": "preview",
   "open": false,
   "lineage": [
    "movie-gen",
    "muse-spark"
   ],
   "desc": "MSL's text-to-video model, revealed in early preview July 7, 2026 alongside Muse Image, with a full release promised for creators and Meta AI. It ranked #3 for text-to-video on the Arena leaderboard at announcement and is positioned to power AI video across Meta's apps, including the Vibes feed.",
   "facts": [
    "Part of the same July 2026 announcement wave as Muse Image, filling the slot where Vibes originally relied on partner models (Midjourney, Black Forest Labs) at its Sept 2025 launch."
   ],
   "latest": "Early preview (Jul 2026); full release pending as of Sept 2026",
   "links": [
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/introducing-muse-image-muse-video-msl"
    },
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2026/07/introducing-muse-image-meta-ai"
    }
   ],
   "why": "Muse Video is Meta's first in-house production video generator, replacing the Midjourney and Black Forest Labs models that Vibes launched with. Ranked #3 on Arena for text-to-video at its July 2026 preview, it is the model Meta intends to put behind AI video across Instagram, Facebook and Meta AI.",
   "try": [
    {
     "label": "Watch AI video in the Vibes feed",
     "url": "https://www.meta.ai/vibes"
    },
    {
     "label": "Read the preview announcement",
     "url": "https://ai.meta.com/blog/introducing-muse-image-muse-video-msl"
    },
    {
     "label": "Meta AI app on the App Store",
     "url": "https://apps.apple.com/us/app/meta-ai/id1558240027"
    }
   ],
   "params": "undisclosed (closed weights; early preview)"
  },
  {
   "id": "muse-code",
   "name": "Muse Code",
   "full_name": "Muse Code",
   "tag": "Meta's terminal coding agent, powered by Spark 1.2",
   "year": 2026,
   "year_label": "2026",
   "cat": "muse",
   "size": 2,
   "status": "closed",
   "open": false,
   "lineage": [
    "muse-spark",
    "cwm"
   ],
   "desc": "Meta's first coding agent, released in beta August 5, 2026 with Muse Spark 1.2. A terminal-based agent for macOS and Linux that plans changes, writes and debugs code across large repos, and keeps multiple specialized background subagents alive throughout a session, using isolated Git worktrees for parallel tasks and an append-only event log for crash-proof resume.",
   "facts": [
    "The 'Contributor' tier costs ~$0.10/$0.20 per M tokens — up to 20x cheaper than standard $1.25/$4.25 — in exchange for letting Meta train on your prompts and completions (60 rpm / 2.1M tokens-per-min cap vs 3,000 rpm / 4M standard).",
    "Ships /plan, /grill (plan stress-testing) and /goal skills.",
    "Installs via one curl command from dev.meta.ai.",
    "Muse Spark 1.2 is benchmarked on Terminal-Bench 2.1 and DeepSWE 1.1."
   ],
   "latest": "Muse Code beta (Aug 5, 2026), powered by Muse Spark 1.2",
   "links": [
    {
     "label": "Meta Research",
     "url": "https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2"
    },
    {
     "label": "Meta for Developers",
     "url": "https://developer.meta.com/ai/resources/blog/build-with-muse-code"
    },
    {
     "label": "VentureBeat",
     "url": "https://venturebeat.com/orchestration/meta-enters-the-ai-coding-wars-with-muse-spark-1-2-and-muse-code-with-persistent-async-background-agents"
    }
   ],
   "why": "Muse Code is Meta's entry into the terminal coding-agent market against Claude Code and Codex, notable for persistent background subagents, git-worktree isolation and a crash-safe event log. Its 'Contributor' tier — up to 20x cheaper in exchange for training on your prompts — is a new pricing model for the category.",
   "try": [
    {
     "label": "Install and run Muse Code (dev.meta.ai docs)",
     "url": "https://dev.meta.ai/docs/"
    },
    {
     "label": "Read the launch post",
     "url": "https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2"
    },
    {
     "label": "The model behind it on OpenRouter",
     "url": "https://openrouter.ai/meta/muse-spark-1.3"
    }
   ],
   "params": ""
  },
  {
   "id": "meta-model-api",
   "name": "Meta Model API",
   "full_name": "Meta Model API (dev.meta.ai)",
   "tag": "Muse Spark for developers",
   "year": 2026,
   "year_label": "2026",
   "cat": "muse",
   "size": 1,
   "status": "closed",
   "open": false,
   "lineage": [
    "muse-spark"
   ],
   "desc": "Meta's first-party API for the Muse model family, opened in public preview July 9, 2026 with Muse Spark 1.1 — Meta's first serious paid developer platform for frontier models, replacing the open-download Llama distribution model. Serves Muse Spark 1.1/1.2 with a 1M-token context window; Muse models are also available via OpenRouter.",
   "facts": [
    "Launch pricing: $1.25 per million input tokens, $4.25 per million output tokens, with $20 in free starter credits.",
    "Docs live at dev.meta.ai.",
    "Positioned head-on against GPT-5.5, Claude Opus 4.8 and Gemini 3.1 Pro APIs."
   ],
   "latest": "Public preview serving Muse Spark 1.2 with expanded global access (Aug 2026)",
   "links": [
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/introducing-muse-spark-meta-model-api"
    },
    {
     "label": "Axios",
     "url": "https://www.axios.com/2026/07/09/meta-ai-spark-model-update-developer"
    },
    {
     "label": "TechCrunch",
     "url": "https://techcrunch.com/2026/07/09/meta-enters-the-crowded-ai-coding-battle-with-muse-spark-1-1"
    }
   ],
   "why": "The Model API is the first time Meta has sold frontier inference directly, replacing the free Llama download model with a paid, OpenAI-SDK-compatible platform at $1.25/$4.25 per million tokens. It now serves Muse Spark 1.3, Muse Image and Muse Voice Transcribe with a 1M-token context, and absorbed the old Llama API in July 2026.",
   "try": [
    {
     "label": "API docs and quickstart",
     "url": "https://dev.meta.ai/docs/"
    },
    {
     "label": "Pricing and models overview",
     "url": "https://developer.meta.com/ai/products/meta-model-api/"
    },
    {
     "label": "Meta models on OpenRouter",
     "url": "https://openrouter.ai/meta"
    }
   ],
   "params": ""
  },
  {
   "id": "meta-ai-app",
   "name": "Meta AI app",
   "full_name": "Meta AI app",
   "tag": "The assistant in front of 3 billion people",
   "year": 2025,
   "year_label": "2025",
   "cat": "muse",
   "size": 2,
   "status": "closed",
   "open": false,
   "lineage": [
    "muse-spark"
   ],
   "desc": "Meta's standalone AI assistant app, launched April 29, 2025 at the first LlamaCon. Built for voice-first, personalized assistance with memory and a social Discover feed, it started on Llama 4 and was upgraded to Muse Spark in April 2026 — Meta AI now reaches over 3 billion users across the app, WhatsApp, Instagram, Facebook, Messenger and smart glasses.",
   "facts": [
    "Unveiled at Meta's inaugural LlamaCon developer conference.",
    "The 2026 Muse Spark upgrade added interruptible natural voice chat, physician-informed health answers, and multi-agent task handling."
   ],
   "latest": "Powered by Muse Spark (Apr 2026), with Muse Image generation (Jul 2026)",
   "links": [
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2025/04/introducing-meta-ai-app-new-way-access-ai-assistant"
    },
    {
     "label": "TechCrunch",
     "url": "https://techcrunch.com/2025/04/29/meta-launches-a-standalone-ai-app-to-compete-with-chatgpt"
    },
    {
     "label": "meta.ai",
     "url": "https://www.meta.ai"
    }
   ],
   "why": "The Meta AI app is the distribution endpoint for everything MSL builds, putting a standalone assistant, voice chat and the Discover and Vibes feeds in front of three billion users across Meta's apps. Its April 2026 switch from Llama 4 to Muse Spark made it the biggest deployment of a closed Meta model.",
   "try": [
    {
     "label": "Download on the App Store",
     "url": "https://apps.apple.com/us/app/meta-ai/id1558240027"
    },
    {
     "label": "Download on Google Play",
     "url": "https://play.google.com/store/apps/details?id=com.facebook.stella"
    },
    {
     "label": "Use it on the web",
     "url": "https://www.meta.ai"
    }
   ],
   "params": ""
  },
  {
   "id": "vibes",
   "name": "Vibes",
   "full_name": "Vibes",
   "tag": "A TikTok-style feed where every video is AI-made",
   "year": 2025,
   "year_label": "2025",
   "cat": "muse",
   "size": 1,
   "status": "closed",
   "open": false,
   "lineage": [
    "muse-video"
   ],
   "desc": "Meta's TikTok-style feed of AI-generated short videos, launched September 25, 2025 inside the Meta AI app and meta.ai. Users prompt, remix and restyle each other's AI videos and cross-post to Instagram and Facebook. Following early traction, Meta began testing a standalone Vibes app in February 2026, first in Brazil and Mexico.",
   "facts": [
    "Meta's answer to OpenAI's Sora app.",
    "At launch it leaned on partner models (Midjourney, Black Forest Labs) while Meta built its own — the slot Muse Video now fills.",
    "Expanded to Europe November 6, 2025.",
    "Critics memorably dubbed it an 'AI slop' feed, yet traction was strong enough to justify its own app.",
    "Critics instantly dubbed it 'the AI slop feed', yet Meta AI app downloads jumped 56% month-over-month to ~3.9M by mid-October 2025; every video displays its generating prompt for transparency."
   ],
   "latest": "Standalone Vibes app in test (Feb 2026, Brazil & Mexico)",
   "links": [
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2025/09/introducing-vibes-ai-videos"
    },
    {
     "label": "TechCrunch · meta tests a standalone app fo",
     "url": "https://www.techcrunch.com/2026/02/05/meta-tests-a-standalone-app-for-its-ai-generated-vibes-videos"
    },
    {
     "label": "TechCrunch · meta brings its short form vid",
     "url": "https://techcrunch.com/2025/11/06/meta-brings-its-short-form-video-feed-of-ai-slop-to-europe"
    }
   ],
   "why": "Vibes was Meta's bet that AI-generated short video could be a feed, not just a tool — and despite the 'AI slop' backlash, Meta AI app downloads jumped 56% month-over-month to about 3.9M by mid-October 2025, enough to justify a standalone app test in Brazil and Mexico in February 2026.",
   "try": [
    {
     "label": "Open the Vibes feed",
     "url": "https://www.meta.ai/vibes"
    },
    {
     "label": "Meta AI app on the App Store",
     "url": "https://apps.apple.com/us/app/meta-ai/id1558240027"
    },
    {
     "label": "Read the launch post",
     "url": "https://about.fb.com/news/2025/09/introducing-vibes-ai-videos/"
    }
   ],
   "params": ""
  },
  {
   "id": "llama",
   "name": "LLaMA",
   "full_name": "LLaMA (original)",
   "tag": "The leak that started the open-weight revolution",
   "year": 2023,
   "year_label": "2023",
   "cat": "language",
   "size": 3,
   "status": "superseded",
   "open": true,
   "lineage": [
    "opt"
   ],
   "desc": "FAIR's first large language model family (7B-65B), released Feb 24, 2023 to researchers under a non-commercial license. LLaMA-13B outperformed GPT-3 175B on many benchmarks, proving smaller models trained on more tokens could win — and kickstarting the open-weights era.",
   "facts": [
    "The paper that turned 'bigger is better' into 'trained-longer is better' — and launched a billion downloads.",
    "Weights leaked via a 4chan torrent about one week after announcement (early March 2023), spawning llama.cpp and the local-LLM movement.",
    "LLaMA-65B was trained on 1.4T tokens.",
    "The meta-llama/llama GitHub repo has ~59,600 stars.",
    "A 13B model beating a 175B model rewrote the industry's scaling assumptions overnight."
   ],
   "latest": "LLaMA 65B (Feb 2023; superseded by Llama 2+)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2302.13971"
    },
    {
     "label": "GitHub · llama",
     "url": "https://github.com/meta-llama/llama"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/large-language-model-llama-meta-ai"
    }
   ],
   "why": "LLaMA proved a 13B model trained on more tokens could beat GPT-3's 175B, overturning 'bigger is better' scaling assumptions. When its weights leaked a week after release, it seeded llama.cpp and the entire local-LLM movement — the open-weight era that Meta then made official strategy with Llama 2.",
   "try": [
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2302.13971"
    },
    {
     "label": "Original repo (59k stars)",
     "url": "https://github.com/meta-llama/llama"
    }
   ],
   "params": "7B / 13B / 33B / 65B"
  },
  {
   "id": "llama2",
   "name": "Llama 2",
   "full_name": "Llama 2",
   "tag": "Open weights you could build a business on",
   "year": 2023,
   "year_label": "2023",
   "cat": "language",
   "size": 3,
   "status": "superseded",
   "open": true,
   "lineage": [
    "llama"
   ],
   "desc": "Released July 18, 2023 with Microsoft as launch partner — the first Llama with a commercial-use license (7B/13B/70B, plus chat-tuned variants with RLHF). It made 'open weights you can build a business on' mainstream and seeded thousands of fine-tunes.",
   "facts": [
    "The moment 'open weights you can build a business on' became Meta's official AI strategy.",
    "Free for commercial use unless your product exceeded 700 million monthly active users — a clause aimed squarely at Big Tech rivals.",
    "Trained on 2 trillion tokens.",
    "Author count jumped from LLaMA-1's 14 to 68 — a sign of how strategic the program had become.",
    "Launched with Microsoft Azure as lead partner and a detailed 77-page safety-heavy report."
   ],
   "latest": "Llama 2 70B / Llama 2-Chat (Jul 2023)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2307.09288"
    },
    {
     "label": "GitHub · llama",
     "url": "https://github.com/meta-llama/llama"
    },
    {
     "label": "Meta AI",
     "url": "https://ai.meta.com/llama"
    },
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2023/07/llama-2"
    }
   ],
   "why": "Llama 2 made commercially usable open weights mainstream: 7B-70B base and RLHF chat models under a license free for anyone below 700 million monthly users, launched with Microsoft Azure. It seeded thousands of fine-tunes and set the template — open weights plus a permissive-but-not-OSI license — that every later Llama followed.",
   "try": [
    {
     "label": "Chat with Llama 2 in a Hugging Face Space",
     "url": "https://huggingface.co/spaces/huggingface-projects/llama-2-7b-chat"
    },
    {
     "label": "Model on Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-2-7b-chat-hf"
    },
    {
     "label": "Run it locally with Ollama",
     "url": "https://ollama.com/library/llama2"
    }
   ],
   "params": "7B / 13B / 70B"
  },
  {
   "id": "llama3",
   "name": "Llama 3 / 3.1 / 3.2 / 3.3",
   "full_name": "Llama 3 / 3.1 / 3.2 / 3.3",
   "tag": "405B: the first GPT-4-class model anyone could download",
   "year": 2024,
   "year_label": "2024",
   "cat": "language",
   "size": 3,
   "status": "superseded",
   "open": true,
   "lineage": [
    "llama2"
   ],
   "desc": "The 2024 herd: Llama 3 8B/70B (Apr 2024), Llama 3.1 with the 405B milestone (Jul 2024) — the first open frontier-class model — Llama 3.2 edge sizes (1B/3B) plus 11B/90B vision models (Sep 2024), and Llama 3.3 70B (Dec 2024) matching 405B-level quality at a fraction of the cost.",
   "facts": [
    "The first GPT-4-class model anyone could download.",
    "Llama 3.1 405B was trained on over 15T tokens using more than 16,000 H100 GPUs (~3.8x10^25 FLOPs).",
    "Llama 3 was trained on two custom 24,576-GPU clusters.",
    "The 'Herd of Models' paper lists hundreds of contributors.",
    "559 listed authors — one of the largest author lists in AI history."
   ],
   "latest": "Llama 3.3 70B (Dec 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2407.21783"
    },
    {
     "label": "GitHub · llama-models",
     "url": "https://github.com/meta-llama/llama-models"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-3.1-405B"
    },
    {
     "label": "Meta AI blog · meta llama 3",
     "url": "https://ai.meta.com/blog/meta-llama-3"
    },
    {
     "label": "Meta AI blog · meta llama 3 1",
     "url": "https://ai.meta.com/blog/meta-llama-3-1"
    }
   ],
   "why": "Llama 3.1 405B was the first openly downloadable model at GPT-4 class, trained on 15T+ tokens across 16,000 H100s, and Llama 3.3 70B then delivered near-405B quality at a fraction of the cost. The 3.2 1B/3B edge models and 11B/90B vision variants made Llama the default open stack from phones to clusters.",
   "try": [
    {
     "label": "Model on Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct"
    },
    {
     "label": "Chat with Llama 3.2 in a Hugging Face Space",
     "url": "https://huggingface.co/spaces/huggingface-projects/llama-3.2-3B-Instruct"
    },
    {
     "label": "Run it locally with Ollama",
     "url": "https://ollama.com/library/llama3.1"
    }
   ],
   "params": "1B / 3B / 8B / 11B / 70B / 90B / 405B"
  },
  {
   "id": "llama4",
   "name": "Llama 4",
   "full_name": "Llama 4 (Scout / Maverick / Behemoth)",
   "tag": "A 10-million-token context window — and a turning point",
   "year": 2025,
   "year_label": "2025",
   "cat": "language",
   "size": 2,
   "status": "superseded",
   "open": true,
   "lineage": [
    "llama3"
   ],
   "desc": "Meta's last open-weight frontier Llama generation, released April 5, 2025: natively multimodal mixture-of-experts models. Scout (17B active, 16 experts, 109B total) shipped with an unprecedented 10M-token context window; Maverick (17B active, 128 experts, 400B total) targeted flagship quality. Behemoth (~2T parameters, 288B active) was previewed but never released.",
   "facts": [
    "A context window big enough to read the entire Harry Potter series about six times over in one prompt.",
    "Scout's 10M-token context was the largest of any open-weight model at launch.",
    "Llama downloads passed 1 billion in March 2025 (~1.2B by LlamaCon that April).",
    "Behemoth was used internally as a teacher to co-distill Scout and Maverick but its weights never shipped — mid-training MoE-routing issues at 2T scale are the reported cause.",
    "The line was superseded by the closed Muse Spark in April 2026."
   ],
   "latest": "Scout & Maverick (Apr 5, 2025); Behemoth never shipped — no Llama 4.x/5 followed",
   "links": [
    {
     "label": "GitHub · llama-models",
     "url": "https://github.com/meta-llama/llama-models"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/meta-llama"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/llama-4-multimodal-intelligence"
    },
    {
     "label": "llama.com",
     "url": "https://www.llama.com"
    },
    {
     "label": "Axios",
     "url": "https://www.axios.com/2025/05/15/meta-behemoth-llama-scaling-delays"
    }
   ],
   "why": "Llama 4 was Meta's last open frontier generation and the moment the strategy cracked: Scout's 10M-token context and native multimodality were genuine firsts for open weights, but the 2T-parameter Behemoth never shipped amid reported MoE-routing problems, and within a year Meta replaced the line with the closed Muse Spark.",
   "try": [
    {
     "label": "Llama 4 Scout on Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct"
    },
    {
     "label": "Run it locally with Ollama",
     "url": "https://ollama.com/library/llama4"
    },
    {
     "label": "Use Maverick via OpenRouter",
     "url": "https://openrouter.ai/meta-llama/llama-4-maverick"
    }
   ],
   "params": "Scout 109B total / 17B active (16 experts); Maverick 400B total / 17B active (128 experts); Behemoth ~2T / 288B active (unreleased)"
  },
  {
   "id": "code-llama",
   "name": "Code Llama",
   "full_name": "Code Llama",
   "tag": "Llama learns to program",
   "year": 2023,
   "year_label": "2023",
   "cat": "language",
   "size": 2,
   "status": "superseded",
   "open": true,
   "lineage": [
    "llama2"
   ],
   "desc": "Llama 2 specialized for code (Aug 24, 2023): base, Python, and Instruct variants at 7B/13B/34B, with a 70B flagship added January 2024. Supported fill-in-the-middle completion and long 100K-token contexts, becoming the standard open coding model of its era.",
   "facts": [
    "Code Llama 70B Instruct scored 67.8% on HumanEval at release — the best open code model at the time.",
    "Same permissive community license as Llama 2."
   ],
   "latest": "Code Llama 70B (Jan 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2308.12950"
    },
    {
     "label": "GitHub · codellama",
     "url": "https://github.com/meta-llama/codellama"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/code-llama-large-language-model-coding"
    }
   ],
   "why": "Code Llama was the standard open coding model of 2023-24: fill-in-the-middle completion, 100K-token contexts and a 70B Instruct variant that hit 67.8% HumanEval, then the best open result. Released under the Llama 2 license, it became the base for countless open coding assistants before general models absorbed the niche.",
   "try": [
    {
     "label": "Try it in the Code Llama Playground",
     "url": "https://huggingface.co/spaces/codellama/codellama-playground"
    },
    {
     "label": "Model on Hugging Face",
     "url": "https://huggingface.co/codellama/CodeLlama-7b-Instruct-hf"
    },
    {
     "label": "Run it locally with Ollama",
     "url": "https://ollama.com/library/codellama"
    }
   ],
   "params": "7B / 13B / 34B / 70B"
  },
  {
   "id": "llama-stack",
   "name": "Llama Stack",
   "full_name": "Llama Stack",
   "tag": "Write the app once, run it local, cloud or on-prem",
   "year": 2024,
   "year_label": "2024",
   "cat": "language",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "llama3"
   ],
   "desc": "Open-source framework standardizing the APIs around Llama-based applications — inference, RAG, agents, tools, safety, evals, telemetry — with swappable providers so apps move between local, cloud, and on-prem unchanged. Launched September 2024 alongside Llama 3.2; now developed in its own llamastack GitHub org with partners like NVIDIA, IBM, Red Hat, and Dell.",
   "facts": [
    "The 0.5 release (Feb 2026) added OpenAI API conformance and an MCP-server Connectors API — Meta's open stack absorbing industry standards.",
    "Now community-governed at llamastack/llama-stack after moving out of the meta-llama org."
   ],
   "latest": "v0.7.2 (May 28, 2026)",
   "links": [
    {
     "label": "GitHub · llamastack/llama-stack · llama stack",
     "url": "https://github.com/llamastack/llama-stack"
    },
    {
     "label": "GitHub · llamastack/llama-stack · releases",
     "url": "https://github.com/llamastack/llama-stack/releases"
    },
    {
     "label": "pypi.org",
     "url": "https://pypi.org/project/llama-stack"
    }
   ],
   "why": "Llama Stack tried to standardize the whole app layer around Llama — inference, RAG, agents, safety, evals — with swappable local and cloud providers, and drew NVIDIA, IBM, Red Hat and Dell as partners. In 2026 the project outgrew its name: it is now OGX, a model-agnostic, OpenAI-compatible agentic API server outside the Meta org.",
   "try": [
    {
     "label": "pip install llama-stack",
     "url": "https://pypi.org/project/llama-stack"
    },
    {
     "label": "The project today: OGX on GitHub",
     "url": "https://github.com/ogx-ai/ogx"
    },
    {
     "label": "Why Llama Stack became OGX",
     "url": "https://ogx-ai.github.io/blog/from-llama-stack-to-ogx"
    }
   ],
   "params": ""
  },
  {
   "id": "llama-api",
   "name": "Llama API",
   "full_name": "Llama API",
   "tag": "Meta's hosted Llama, folded into the Model API in 2026",
   "year": 2025,
   "year_label": "2025",
   "cat": "language",
   "size": 1,
   "status": "archived",
   "open": false,
   "lineage": [
    "llama4"
   ],
   "desc": "Meta's first-party hosted inference service, previewed at LlamaCon (April 29, 2025) with OpenAI-SDK compatibility, one-click keys, and fine-tuning tools, plus fast-inference partnerships with Cerebras and Groq. It never left preview: in July 2026 it was folded into the newer Muse-era Meta Model API, with developers steered to third-party hosts for Llama.",
   "facts": [
    "Meta pledged 'we do not use your prompts or model responses to train our AI models.' Its 15-month life (preview to sunset) neatly brackets Meta's pivot from Llama to Muse."
   ],
   "latest": "Public preview (Apr 2025) — folded into the Meta Model API (Jul 2026)",
   "links": [
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/llamacon-llama-news"
    },
    {
     "label": "Meta for Developers",
     "url": "https://llama.developer.meta.com"
    }
   ],
   "why": "The Llama API was Meta's first attempt at hosted inference — OpenAI-SDK compatible, one-click keys, Cerebras and Groq fast paths — announced as Llama passed one billion downloads. It never left preview; its 15-month arc from LlamaCon to absorption into the Muse-era Model API traces Meta's pivot from open Llama to paid Muse.",
   "try": [
    {
     "label": "Read the LlamaCon announcement",
     "url": "https://ai.meta.com/blog/llamacon-llama-news"
    },
    {
     "label": "Its successor: Meta Model API docs",
     "url": "https://dev.meta.ai/docs/"
    },
    {
     "label": "Llama 4 hosted on OpenRouter today",
     "url": "https://openrouter.ai/meta-llama/llama-4-maverick"
    }
   ],
   "params": ""
  },
  {
   "id": "fairseq",
   "name": "fairseq",
   "full_name": "fairseq",
   "tag": "The toolkit that trained a decade of breakthroughs",
   "year": 2017,
   "year_label": "2017",
   "cat": "language",
   "size": 2,
   "status": "archived",
   "open": true,
   "lineage": [],
   "desc": "FAIR's sequence-to-sequence research toolkit, open-sourced in 2017 — the workbench on which a decade of Meta breakthroughs was built: convolutional seq2seq, RoBERTa, BART, XLM-R, wav2vec 2.0, NLLB, and MMS all shipped as fairseq examples. One of the most influential NLP codebases ever released.",
   "facts": [
    "~32,200 GitHub stars.",
    "Began as Lua Torch code for convolutional translation (pre-Transformer, May 2017) before the PyTorch rewrite.",
    "Its examples directory reads like a museum of modern NLP."
   ],
   "latest": "fairseq-py (actively archived-era; final major line 0.12.x)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/1904.01038"
    },
    {
     "label": "GitHub · fairseq",
     "url": "https://github.com/facebookresearch/fairseq"
    }
   ],
   "why": "fairseq was the workbench for a decade of Meta NLP: RoBERTa, BART, XLM-R, wav2vec 2.0, NLLB and MMS all shipped as fairseq examples, and its 32,000-star codebase trained much of the field's translation and speech research. The repository is now archived, but its examples directory remains a museum of modern NLP.",
   "try": [
    {
     "label": "pip install fairseq",
     "url": "https://pypi.org/project/fairseq"
    },
    {
     "label": "Archived repo on GitHub",
     "url": "https://github.com/facebookresearch/fairseq"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/1904.01038"
    }
   ],
   "params": ""
  },
  {
   "id": "roberta",
   "name": "RoBERTa",
   "full_name": "RoBERTa",
   "tag": "BERT, done right",
   "year": 2019,
   "year_label": "2019",
   "cat": "language",
   "size": 2,
   "status": "superseded",
   "open": true,
   "lineage": [
    "fairseq"
   ],
   "desc": "'A Robustly Optimized BERT Pretraining Approach' (July 2019): Meta showed BERT was severely undertrained, and that longer training on more data (160GB of text) with dynamic masking — no architecture change — topped the GLUE leaderboard. For years it was the default encoder for real-world NLP systems.",
   "facts": [
    "Beat BERT using the exact same architecture — the paper is a landmark argument that training recipes matter as much as architectures.",
    "Still among the most-downloaded models on Hugging Face years later."
   ],
   "latest": "RoBERTa base/large (Jul 2019)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/1907.11692"
    },
    {
     "label": "GitHub · fairseq",
     "url": "https://github.com/facebookresearch/fairseq/tree/main/examples/roberta"
    }
   ],
   "why": "RoBERTa showed BERT was badly undertrained: same architecture, more data, longer training and dynamic masking topped GLUE. It is the landmark argument that recipes matter as much as architectures, and years later roberta-base still logs nearly ten million monthly Hugging Face downloads as a default encoder for real-world NLP.",
   "try": [
    {
     "label": "roberta-base on Hugging Face",
     "url": "https://huggingface.co/FacebookAI/roberta-base"
    },
    {
     "label": "roberta-large on Hugging Face",
     "url": "https://huggingface.co/FacebookAI/roberta-large"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/1907.11692"
    }
   ],
   "params": "125M (base) / 355M (large)"
  },
  {
   "id": "xlmr",
   "name": "XLM-R",
   "full_name": "XLM-R (XLM-RoBERTa)",
   "tag": "100 languages in one encoder — the ancestor of NLLB",
   "year": 2019,
   "year_label": "2019",
   "cat": "language",
   "size": 1,
   "status": "superseded",
   "open": true,
   "lineage": [
    "roberta"
   ],
   "desc": "Cross-lingual RoBERTa trained on 2.5TB of filtered CommonCrawl covering 100 languages (November 2019). It delivered massive gains on low-resource languages and set the standard for multilingual understanding — a direct ancestor of Meta's translation moonshots like NLLB.",
   "facts": [
    "Trained on 2.5 terabytes of text across 100 languages; showed for the first time that one multilingual model could rival monolingual models on high-resource languages while lifting low-resource ones."
   ],
   "latest": "XLM-R base/large/XL/XXL (2019-2021)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/1911.02116"
    },
    {
     "label": "GitHub · fairseq",
     "url": "https://github.com/facebookresearch/fairseq/tree/main/examples/xlmr"
    }
   ],
   "why": "XLM-R proved one encoder trained on 2.5TB of CommonCrawl across 100 languages could match monolingual models on high-resource languages while lifting low-resource ones dramatically. It became the default multilingual backbone for years — xlm-roberta-base still sees over 20 million monthly Hugging Face downloads — and set up Meta's NLLB translation push.",
   "try": [
    {
     "label": "xlm-roberta-base on Hugging Face",
     "url": "https://huggingface.co/FacebookAI/xlm-roberta-base"
    },
    {
     "label": "xlm-roberta-large on Hugging Face",
     "url": "https://huggingface.co/FacebookAI/xlm-roberta-large"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/1911.02116"
    }
   ],
   "params": "~280M (base) / ~560M (large) / 3.5B (XL) / 10.7B (XXL)"
  },
  {
   "id": "bart",
   "name": "BART",
   "full_name": "BART",
   "tag": "Corrupt the text, learn to rebuild it — summarization's workhorse",
   "year": 2019,
   "year_label": "2019",
   "cat": "language",
   "size": 1,
   "status": "superseded",
   "open": true,
   "lineage": [
    "fairseq"
   ],
   "desc": "Denoising sequence-to-sequence pretraining (October 2019): corrupt text with noise (deletion, infilling, shuffling), train a full encoder-decoder Transformer to reconstruct it. BART became the go-to model for summarization and generation tasks, and its recipe influenced a generation of seq2seq models.",
   "facts": [
    "BART-large fine-tuned on CNN/DailyMail was the standard summarization baseline for years; distilled variants (DistilBART) still serve production summarizers."
   ],
   "latest": "BART base/large (Oct 2019)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/1910.13461"
    },
    {
     "label": "GitHub · fairseq",
     "url": "https://github.com/facebookresearch/fairseq/tree/main/examples/bart"
    }
   ],
   "why": "BART's corrupt-and-reconstruct pretraining made a full encoder-decoder Transformer the go-to for summarization and generation; bart-large-cnn was the standard summarization baseline for years and still handles over a million monthly Hugging Face downloads. Its denoising recipe shaped the seq2seq models that followed it.",
   "try": [
    {
     "label": "Summarize text with bart-large-cnn",
     "url": "https://huggingface.co/facebook/bart-large-cnn"
    },
    {
     "label": "bart-large on Hugging Face",
     "url": "https://huggingface.co/facebook/bart-large"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/1910.13461"
    }
   ],
   "params": "140M (base) / 400M (large)"
  },
  {
   "id": "opt",
   "name": "OPT-175B",
   "full_name": "OPT-175B",
   "tag": "The 175B model that published its 3 a.m. crash logs",
   "year": 2022,
   "year_label": "2022",
   "cat": "language",
   "size": 2,
   "status": "superseded",
   "open": true,
   "lineage": [
    "fairseq"
   ],
   "desc": "Open Pre-trained Transformers (May 3, 2022): a GPT-3-scale 175B model whose weights were shared with researchers, alongside the full codebase — and, famously, the raw 114-page logbook chronicling every crash, loss spike, and hardware failure of the training run. A radical act of transparency at frontier scale.",
   "facts": [
    "The first time a lab published its 3 a.m. on-call notes from training a 175-billion-parameter model — warts, crashes, and all.",
    "The public 'Chronicles' logbook documents 35+ manual restarts and cascades of GPU failures over ~2 months on 992 80GB A100s.",
    "Meta estimated the carbon footprint at ~75 tons CO2e versus an estimated 500 for GPT-3 — a 1/7th footprint at the same scale.",
    "The logbook records 35 training restarts and over 100 hosts cycled due to hardware failures — machines died almost daily on the 992 A100 GPUs.",
    "19 authors (Zhang et al.)."
   ],
   "latest": "OPT 125M-175B suite (May 2022)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2205.01068"
    },
    {
     "label": "GitHub · metaseq · metaseq",
     "url": "https://github.com/facebookresearch/metaseq"
    },
    {
     "label": "GitHub · metaseq · chronicles",
     "url": "https://github.com/facebookresearch/metaseq/tree/main/projects/OPT/chronicles"
    },
    {
     "label": "GitHub · metaseq · OPT175B_Logbook.pdf",
     "url": "https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf"
    }
   ],
   "why": "OPT-175B was the first GPT-3-scale model whose weights reached researchers, but its lasting contribution is transparency: Meta published the full codebase and a 114-page logbook of crashes, loss spikes and 35+ restarts across 992 A100s, at roughly one-seventh of GPT-3's estimated carbon footprint. It set the template LLaMA followed.",
   "try": [
    {
     "label": "OPT-1.3B on Hugging Face",
     "url": "https://huggingface.co/facebook/opt-1.3b"
    },
    {
     "label": "Read the OPT-175B training logbook (PDF)",
     "url": "https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2205.01068"
    }
   ],
   "params": "125M / 350M / 1.3B / 2.7B / 6.7B / 13B / 30B / 66B / 175B"
  },
  {
   "id": "galactica",
   "name": "Galactica",
   "full_name": "Galactica",
   "tag": "The science LLM that lasted three days in public",
   "year": 2022,
   "year_label": "2022",
   "cat": "language",
   "size": 2,
   "status": "archived",
   "open": true,
   "lineage": [
    "opt"
   ],
   "desc": "A 120B model trained on 48 million scientific papers, textbooks, and reference material to 'organize science' — write reviews, generate citations, solve equations. The public demo launched November 15, 2022 and was pulled three days later after criticism that it fluently hallucinated authoritative-sounding fake science. Model weights remain available.",
   "facts": [
    "The takedown happened just two weeks before ChatGPT launched — a legendary near-miss in AI history.",
    "Critics got it to generate a wiki article on 'the history of bears in space.' Yann LeCun publicly defended it as the demo came down."
   ],
   "latest": "Galactica 125M-120B (Nov 2022; demo withdrawn after 3 days)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2211.09085"
    },
    {
     "label": "GitHub · paperswithcode/galai",
     "url": "https://github.com/paperswithcode/galai"
    }
   ],
   "why": "Galactica is AI's most instructive failure: a 120B model trained on 48 million papers to 'organize science' whose demo was pulled after three days for fluently hallucinating fake citations and papers — two weeks before ChatGPT launched. It made hallucination a mainstream concern, and its weights remain available for research.",
   "try": [
    {
     "label": "galactica-1.3b on Hugging Face",
     "url": "https://huggingface.co/facebook/galactica-1.3b"
    },
    {
     "label": "pip install galai",
     "url": "https://pypi.org/project/galai"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2211.09085"
    }
   ],
   "params": "125M / 1.3B / 6.7B / 30B / 120B"
  },
  {
   "id": "blenderbot",
   "name": "BlenderBot",
   "full_name": "BlenderBot (1, 2, 3)",
   "tag": "A public chatbot that criticized its own CEO — months before ChatGPT",
   "year": 2020,
   "year_label": "2020-2022",
   "cat": "language",
   "size": 1,
   "status": "archived",
   "open": true,
   "lineage": [],
   "desc": "Meta's open-domain chatbot line: BlenderBot (2020) blended personality, empathy and knowledge; BlenderBot 2 (2021) added internet search and long-term memory; BlenderBot 3 (August 2022) scaled to 175B parameters with a live public US demo that learned from conversations — months before ChatGPT made chatbots a phenomenon.",
   "facts": [
    "Press had a field day when BB3 criticized its own CEO.",
    "All model weights, code, and even conversation data were released for research."
   ],
   "latest": "BlenderBot 3 175B (Aug 2022)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2208.03188"
    },
    {
     "label": "parl.ai",
     "url": "https://parl.ai/projects/bb3"
    }
   ],
   "why": "BlenderBot was Meta's public chatbot line before chatbots were a phenomenon: BB1 blended persona, empathy and knowledge, BB2 added search and long-term memory, and BB3 scaled to 175B with a live US demo that learned from users — months before ChatGPT. Everything, including conversation data, was released for research.",
   "try": [
    {
     "label": "blenderbot-400M-distill on Hugging Face",
     "url": "https://huggingface.co/facebook/blenderbot-400M-distill"
    },
    {
     "label": "BlenderBot 3 project page (ParlAI)",
     "url": "https://parl.ai/projects/bb3/"
    },
    {
     "label": "Read the BB3 paper",
     "url": "https://arxiv.org/abs/2208.03188"
    }
   ],
   "params": "BB1 90M / 2.7B / 9.4B; BB2 400M / 2.7B; BB3 3B / 30B / 175B"
  },
  {
   "id": "lcm",
   "name": "Large Concept Models",
   "full_name": "Large Concept Models (LCM)",
   "tag": "An AI that predicts the next idea, not the next word",
   "year": 2024,
   "year_label": "2024",
   "cat": "language",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR's bet that the token is the wrong unit of thought: LCMs predict the next sentence-level 'concept' in SONAR embedding space — a language- and modality-agnostic representation covering 200 languages — rather than the next word. Showed strong zero-shot cross-lingual generalization on summarization; training code open-sourced.",
   "facts": [
    "Reasons in an embedding space shared by 200 text languages and speech — so a model 'thinks' once and can surface it in any language.",
    "The repo picked up ~2,400 stars within a year.",
    "Because it reasons in a language-agnostic concept space, one trained model generalizes zero-shot to dozens of languages it was never tuned for.",
    "21 credited authors ('LCM team')."
   ],
   "latest": "LCM 1.6B/7B research release (Dec 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2412.08821"
    },
    {
     "label": "GitHub · large_concept_model",
     "url": "https://github.com/facebookresearch/large_concept_model"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/large-concept-models-language-modeling-in-a-sentence-representation-space"
    }
   ],
   "why": "LCM challenges the token as the unit of language modeling: it predicts the next sentence-level concept in SONAR's language- and modality-agnostic embedding space covering 200 languages, so one model reasons once and surfaces it in any language. It is FAIR's most explicit architectural bet against the dense next-token transformer.",
   "try": [
    {
     "label": "Training code on GitHub",
     "url": "https://github.com/facebookresearch/large_concept_model"
    },
    {
     "label": "pip install sonar-space (the embedding space it reasons in)",
     "url": "https://pypi.org/project/sonar-space"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2412.08821"
    }
   ],
   "params": "1.6B / 7B"
  },
  {
   "id": "blt",
   "name": "Byte Latent Transformer",
   "full_name": "Byte Latent Transformer (BLT)",
   "tag": "The architecture that killed the tokenizer",
   "year": 2024,
   "year_label": "2024",
   "cat": "language",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "A tokenizer-free LLM architecture that reads raw bytes, dynamically grouping them into patches sized by entropy — spending compute where text is hard, coasting where it's easy. First byte-level architecture to match tokenization-based models (Llama 3 class) at scale, with up to 50% fewer inference FLOPs; weights for 1B and 8B released 2025.",
   "facts": [
    "It killed the tokenizer — the one hand-engineered relic every modern LLM still depended on.",
    "Scaling study ran to 8B parameters and 8 trillion training bytes.",
    "Because there is no tokenizer, it is naturally robust to typos, weird spellings, and the classic 'how many r's in strawberry' failure class.",
    "Matches Llama 3 training performance with up to 50% fewer inference FLOPs, and is far more robust to typos and character-level noise because it reads raw bytes.",
    "14 authors (Pagnoni et al.)."
   ],
   "latest": "Dynamic BLT weights 1B & 8B (May 2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2412.09871"
    },
    {
     "label": "GitHub · blt",
     "url": "https://github.com/facebookresearch/blt"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/byte-latent-transformer-patches-scale-better-than-tokens"
    }
   ],
   "why": "BLT removed the tokenizer — the last hand-engineered component in modern LLMs — by reading raw bytes and grouping them into entropy-sized patches. It is the first byte-level architecture to match Llama 3-class tokenized models at scale, with up to 50% fewer inference FLOPs and native robustness to typos and character-level noise.",
   "try": [
    {
     "label": "BLT-1B weights on Hugging Face",
     "url": "https://huggingface.co/facebook/blt-1b"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/blt"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2412.09871"
    }
   ],
   "params": "1B / 7B released (paper scales to 8B)"
  },
  {
   "id": "coconut",
   "name": "Coconut",
   "full_name": "Coconut (Chain of Continuous Thought)",
   "tag": "Reasoning in latent space instead of words",
   "year": 2024,
   "year_label": "2024",
   "cat": "language",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR research letting LLMs reason in latent space instead of words: the model's hidden state is fed back as the next input embedding, so 'thoughts' never get flattened into tokens. Coconut can encode multiple candidate next steps simultaneously — an emergent breadth-first search — beating chain-of-thought on logic tasks that require planning and backtracking, with fewer thinking tokens.",
   "facts": [
    "Inspired by neuroscience: language areas of the human brain are largely quiet during hard reasoning.",
    "The continuous thought can hold several possible reasoning branches at once, like superposition."
   ],
   "latest": "Coconut (Dec 2024, code released)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2412.06769"
    },
    {
     "label": "GitHub · coconut",
     "url": "https://github.com/facebookresearch/coconut"
    }
   ],
   "why": "Coconut showed LLMs can reason in latent space instead of words: feeding the hidden state back as the next input lets a model hold several candidate reasoning branches at once, an emergent breadth-first search. It beat chain-of-thought on planning-heavy logic tasks with fewer thinking tokens, opening the latent-reasoning research line.",
   "try": [
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/coconut"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2412.06769"
    }
   ],
   "params": "research code trained on GPT-2 (124M)"
  },
  {
   "id": "multi-token",
   "name": "Multi-token prediction",
   "full_name": "Multi-token prediction",
   "tag": "Several tokens per step: stronger code models, up to 3x faster",
   "year": 2024,
   "year_label": "2024",
   "cat": "language",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "'Better & Faster Large Language Models via Multi-token Prediction' (April 2024): training models to predict several future tokens at once via parallel output heads yields better sample efficiency and stronger code models — 13B models solved 12-17% more HumanEval/MBPP problems — while enabling up to 3x faster inference via self-speculative decoding. 7B weights released for research.",
   "facts": [
    "The idea flowed into industry practice (speculative decoding heads); Meta released the 7B multi-token model on Hugging Face under a non-commercial license.",
    "Zero training-time overhead versus next-token prediction."
   ],
   "latest": "Paper Apr 2024; 7B code models released Jul 2024",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2404.19737"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/facebook/multi-token-prediction"
    }
   ],
   "why": "Multi-token prediction showed that training a model to predict several future tokens at once — with zero training overhead — yields better sample efficiency and stronger code models, with 13B models solving 12-17% more HumanEval/MBPP problems, plus up to 3x faster self-speculative decoding. The idea flowed straight into industry speculative-decoding heads.",
   "try": [
    {
     "label": "7B code models on Hugging Face",
     "url": "https://huggingface.co/facebook/multi-token-prediction"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2404.19737"
    }
   ],
   "params": "7B (released code models)"
  },
  {
   "id": "memory-layers",
   "name": "Memory Layers at Scale",
   "full_name": "Memory Layers at Scale",
   "tag": "Capacity without FLOPs: a lookup table the model learns to consult",
   "year": 2024,
   "year_label": "2024",
   "cat": "language",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR work scaling trainable key-value memory layers to add capacity without adding FLOPs: a 1.3B model with memory matched models trained on 2-4x more compute, and beat MoE models at equal budget on factual tasks. Part of the same December 2024 FAIR wave as LCM and BLT — three simultaneous bets against the standard dense transformer.",
   "facts": [
    "Scaled memory pools past 1 million keys and 128B memory parameters.",
    "Especially strong on factual QA — the memory acts like a built-in lookup table the model learns to consult."
   ],
   "latest": "Paper + code (Dec 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2412.09764"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/meta-fair-updates-agents-robustness-safety-architecture"
    }
   ],
   "why": "Memory layers add trainable key-value lookup capacity without adding FLOPs: a 1.3B model with memory matched dense models trained on 2-4x more compute and beat MoE at equal budget on factual tasks, scaling to 128B memory parameters. It is one of FAIR's three December 2024 bets against the standard dense transformer.",
   "try": [
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/memory"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2412.09764"
    }
   ],
   "params": "base models up to 8B, memory pools up to 128B parameters"
  },
  {
   "id": "cwm",
   "name": "Code World Model",
   "full_name": "Code World Model (CWM)",
   "tag": "A neural debugger that simulates code in its head",
   "year": 2025,
   "year_label": "2025",
   "cat": "language",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "code-llama"
   ],
   "desc": "A 32B dense open-weights research LLM from Meta FAIR (released Sept–Oct 2025) that learns a 'world model' of code execution: mid-trained on observation-action trajectories from Python interpreters and agentic Docker environments, then multi-task RL on verifiable coding, math, and software-engineering tasks, with a 131k-token context.",
   "facts": [
    "65.8% pass@1 on SWE-bench Verified (with test-time scaling), 68.6% on LiveCodeBench, 96.6% on Math-500, 76.0% on AIME 2024 — remarkable for a 32B dense model.",
    "It can simulate Python execution step by step, predicting variable states like a neural debugger.",
    "Released under a custom research license with three training-stage checkpoints for reproducibility.",
    "Weights come in three stages (pretrain, SFT, post-trained) so researchers can study the whole pipeline — rare transparency at 32B scale.",
    "131K-token context with alternating local/global sliding-window attention."
   ],
   "latest": "CWM 32B (Sept/Oct 2025), checkpoints at mid-train, SFT and RL stages",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2510.02387"
    },
    {
     "label": "GitHub · cwm",
     "url": "https://github.com/facebookresearch/cwm"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/cwm-an-open-weights-llm-for-research-on-code-generation-with-world-models"
    }
   ],
   "why": "CWM is the first open-weights LLM trained to model code execution itself — predicting variable states step by step like a neural debugger — and at 32B dense it reaches 65.8% on SWE-bench Verified. Releasing mid-train, SFT and RL checkpoints made the whole post-training pipeline studyable, rare transparency at that scale.",
   "try": [
    {
     "label": "Weights on Hugging Face",
     "url": "https://huggingface.co/facebook/cwm"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/cwm"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2510.02387"
    }
   ],
   "params": "32B dense"
  },
  {
   "id": "sam",
   "name": "Segment Anything",
   "full_name": "Segment Anything (SAM)",
   "tag": "Click anything, cut out anything — segmentation, solved",
   "year": 2023,
   "year_label": "2023",
   "cat": "vision",
   "size": 3,
   "status": "superseded",
   "open": true,
   "lineage": [
    "mask-rcnn"
   ],
   "desc": "The first promptable segmentation foundation model: click a point or draw a box and SAM returns a pixel-accurate mask for any object, zero-shot, in any image. Released April 2023 — model and code under Apache 2.0, alongside the research-licensed SA-1B, then the largest segmentation dataset ever built, it reset computer vision overnight.",
   "facts": [
    "One clickable model made 'cut out any object in any photo' a solved commodity, and it shipped with a browser demo on day one.",
    "SA-1B holds 1.1 billion masks across 11 million licensed images — roughly 400x more masks than any prior dataset.",
    "SAM powers Instagram's Backdrop and Cutouts editing features, and became one of the most-cited vision papers of the 2020s.",
    "12 authors (Kirillov et al.), submitted April 5, 2023."
   ],
   "latest": "SAM ViT-H (Apr 2023); line continued by SAM 2 (2024) and SAM 3 (2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2304.02643"
    },
    {
     "label": "GitHub · segment-anything",
     "url": "https://github.com/facebookresearch/segment-anything"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/segment-anything"
    },
    {
     "label": "Meta AI",
     "url": "https://ai.meta.com/datasets/segment-anything"
    }
   ],
   "why": "SAM turned segmentation into a commodity: one promptable model that cuts out any object in any image zero-shot, released Apache-2.0 with SA-1B's 1.1 billion masks and a browser demo on day one. It became one of the most-cited vision papers of the 2020s, spawned a family of successors, and powers Instagram's editing tools.",
   "try": [
    {
     "label": "sam-vit-huge on Hugging Face",
     "url": "https://huggingface.co/facebook/sam-vit-huge"
    },
    {
     "label": "Run the official notebook in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/segment-anything/blob/main/notebooks/predictor_example.ipynb"
    },
    {
     "label": "Segment Anything Playground (now running SAM 3)",
     "url": "https://aidemos.meta.com/segment-anything"
    }
   ],
   "params": "ViT-B 91M / ViT-L 308M / ViT-H 636M"
  },
  {
   "id": "sam2",
   "name": "SAM 2",
   "full_name": "SAM 2",
   "tag": "One click on one frame tracks an object through a whole video",
   "year": 2024,
   "year_label": "2024",
   "cat": "vision",
   "size": 2,
   "status": "superseded",
   "open": true,
   "lineage": [
    "sam"
   ],
   "desc": "Extends promptable segmentation to video with a streaming-memory transformer: prompt an object once and SAM 2 tracks its 'masklet' across frames in real time (~44 fps), surviving occlusions and reappearances. Also 6x faster and more accurate than SAM on still images. Apache 2.0 code and weights, July 2024.",
   "facts": [
    "It tracks an object through an entire video from a single click on one frame.",
    "Its SA-V dataset (about 51,000 videos, 643,000 masklets, CC-BY-4.0) is roughly 50x larger than prior video segmentation datasets.",
    "Using SAM 2 in the annotation loop made labeling 8.4x faster than per-frame manual annotation.",
    "SA-V is 53x larger than the previous biggest video object segmentation dataset; masklets are 70% auto-generated (451.7K) plus 190.9K manual, filmed across 47 countries with 14-second average clips.",
    "18 authors (Ravi et al.)."
   ],
   "latest": "SAM 2.1 (Sep 2024), plus the SAM 2.1 Developer Suite with open training code",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2408.00714"
    },
    {
     "label": "GitHub · sam2",
     "url": "https://github.com/facebookresearch/sam2"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/segment-anything-2"
    },
    {
     "label": "Meta AI",
     "url": "https://ai.meta.com/datasets/segment-anything-video"
    }
   ],
   "why": "SAM 2 extended promptable segmentation to video: one click on one frame tracks an object through occlusions at about 44 fps, and it is 6x faster than SAM on stills. Its SA-V dataset (51,000 videos, 643,000 masklets) was roughly 50x larger than prior video segmentation data, making it the standard video-mask tool.",
   "try": [
    {
     "label": "Official SAM 2 demo",
     "url": "https://sam2.metademolab.com/"
    },
    {
     "label": "sam2.1-hiera-large on Hugging Face",
     "url": "https://huggingface.co/facebook/sam2.1-hiera-large"
    },
    {
     "label": "Run the official notebook in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/sam2/blob/main/notebooks/image_predictor_example.ipynb"
    }
   ],
   "params": "Hiera tiny 39M / small 46M / base+ 81M / large 224M"
  },
  {
   "id": "sam3",
   "name": "SAM 3",
   "full_name": "SAM 3",
   "tag": "Type 'yellow school bus' — it masks every single one",
   "year": 2025,
   "year_label": "2025",
   "cat": "vision",
   "size": 3,
   "status": "active",
   "open": true,
   "lineage": [
    "sam2"
   ],
   "desc": "The first SAM that understands language: an 848M-parameter unified detector-plus-tracker that finds, segments, and tracks every instance of a concept from a short text phrase ('red baseball cap') or image exemplar, across images and video. Released November 19, 2025 with open checkpoints, alongside the browser-based Segment Anything Playground.",
   "facts": [
    "You type 'yellow school bus' and it finds and masks every single one — in video, in near real time.",
    "Its SA-Co benchmark covers 270,000 unique concepts — over 50x more than prior benchmarks — and SAM 3 reaches 75-80% of human performance on it.",
    "It ships in real products: object-level video effects in Instagram's Edits app and Meta AI.",
    "Meta also released SA-FARI, a wildlife dataset of 10,000+ camera-trap videos spanning 100+ species.",
    "Its data engine produced 4M unique concept labels with hard negatives."
   ],
   "latest": "SAM 3.1 (Mar 27, 2026) — multiplexing tracks up to 16 objects per forward pass, doubling video throughput from 16 to 32 fps on one H100",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2511.16719"
    },
    {
     "label": "GitHub · sam3",
     "url": "https://github.com/facebookresearch/sam3"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/segment-anything-model-3"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/sam-3-segment-anything-with-concepts"
    },
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2025/11/new-sam-models-detect-objects-create-3d-reconstructions"
    }
   ],
   "why": "SAM 3 gave Segment Anything language: type 'yellow school bus' and it finds, masks and tracks every instance across images and video, reaching 75-80% of human performance on a 270,000-concept benchmark. It ships inside Instagram Edits and Meta AI, and the 3.1 update doubled video throughput to 32 fps on one H100.",
   "try": [
    {
     "label": "Segment Anything Playground",
     "url": "https://aidemos.meta.com/segment-anything"
    },
    {
     "label": "Weights on Hugging Face",
     "url": "https://huggingface.co/facebook/sam3"
    },
    {
     "label": "Run the official notebook in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/sam3/blob/main/examples/sam3_image_batched_inference.ipynb"
    }
   ],
   "params": "848M"
  },
  {
   "id": "sam3d",
   "name": "SAM 3D (Objects + Body)",
   "full_name": "SAM 3D (Objects + Body)",
   "tag": "A single photo becomes a 3D object",
   "year": 2025,
   "year_label": "2025",
   "cat": "vision",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "sam3"
   ],
   "desc": "Two generative models released November 19, 2025 that lift a single 2D photo into 3D: SAM 3D Objects reconstructs full shape, texture, and scene layout even under occlusion and clutter, while SAM 3D Body recovers full-body human mesh, pose, and shape (optionally prompted with keypoints or masks). Checkpoints and inference code are open.",
   "facts": [
    "SAM 3D Body runs on a DINOv3 backbone — two flagship Meta research lines fused in one model.",
    "Together with SAM 3 it powers Facebook Marketplace's 'View in Room' furniture preview.",
    "Meta also released SA-3DAO, a benchmark of artist-made meshes paired with real photos."
   ],
   "latest": "v1 (Nov 19, 2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2511.16624"
    },
    {
     "label": "GitHub · sam-3d-objects",
     "url": "https://github.com/facebookresearch/sam-3d-objects"
    },
    {
     "label": "GitHub · sam-3d-body",
     "url": "https://github.com/facebookresearch/sam-3d-body"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/sam-3d"
    }
   ],
   "why": "SAM 3D lifts a single photo into 3D: Objects reconstructs shape, texture and layout even under occlusion, and Body recovers full human mesh and pose on a DINOv3 backbone. Together with SAM 3 it powers Facebook Marketplace's 'View in Room', and the artist-made SA-3DAO benchmark raises the bar for single-image reconstruction.",
   "try": [
    {
     "label": "Turn an image into 3D in the Playground",
     "url": "https://aidemos.meta.com/segment-anything/editor/convert-image-to-3d"
    },
    {
     "label": "SAM 3D Objects on Hugging Face",
     "url": "https://huggingface.co/facebook/sam-3d-objects"
    },
    {
     "label": "SAM 3D Body on Hugging Face",
     "url": "https://huggingface.co/facebook/sam-3d-body-dinov3"
    }
   ],
   "params": ""
  },
  {
   "id": "dino",
   "name": "DINO",
   "full_name": "DINO",
   "tag": "Attention maps that learned to see objects — with zero labels",
   "year": 2021,
   "year_label": "2021",
   "cat": "vision",
   "size": 2,
   "status": "superseded",
   "open": true,
   "lineage": [],
   "desc": "Self-distillation with no labels: DINO showed that Vision Transformers trained purely self-supervised develop attention maps that segment objects without ever being told what an object is. This 2021 FAIR/Inria result seeded Meta's entire label-free backbone program and popularized emergent properties in ViTs.",
   "facts": [
    "The famous attention-map visualizations — a ViT 'seeing' object outlines with zero labels — became one of the most recognizable images in modern vision research."
   ],
   "latest": "DINO (2021); superseded by DINOv2 (2023) and DINOv3 (2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2104.14294"
    },
    {
     "label": "GitHub · dino",
     "url": "https://github.com/facebookresearch/dino"
    }
   ],
   "why": "DINO showed that a Vision Transformer trained with self-distillation and no labels develops attention maps that segment objects on their own — an emergent property that became one of the most recognizable images in vision research. It seeded Meta's entire label-free backbone program (DINOv2, DINOv3) and mainstreamed emergent behavior in ViTs.",
   "try": [
    {
     "label": "dino-vitb16 on Hugging Face",
     "url": "https://huggingface.co/facebook/dino-vitb16"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/dino"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2104.14294"
    }
   ],
   "params": "ViT-S/16 21M / ViT-B/16 85M"
  },
  {
   "id": "dinov2",
   "name": "DINOv2",
   "full_name": "DINOv2",
   "tag": "Label-free features that beat supervised pipelines",
   "year": 2023,
   "year_label": "2023",
   "cat": "vision",
   "size": 3,
   "status": "superseded",
   "open": true,
   "lineage": [
    "dino"
   ],
   "desc": "A family of self-supervised vision backbones (up to a 1.1B-parameter ViT-g) trained on the curated 142M-image LVD-142M dataset, producing all-purpose features that beat supervised and CLIP-style models on classification, depth, and segmentation without fine-tuning. Initially research-licensed in April 2023, relicensed Apache 2.0 later that year.",
   "facts": [
    "It learned to see — matching parts across totally different objects — without ever reading a single label.",
    "DINOv2 features were adopted for medical imaging (histology, endoscopy) and depth estimation; its RGB-coded PCA-of-features videos went viral among researchers as a way to 'see' what a network sees.",
    "Proved frozen, label-free features could beat supervised and text-supervised pipelines; 26 authors (Oquab et al.).",
    "Used with the World Resources Institute to map forests and estimate tree canopy height from satellite imagery; NASA JPL adopted DINOv2 features for Mars exploration robots."
   ],
   "latest": "DINOv2 with registers (2023-24); superseded by DINOv3",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2304.07193"
    },
    {
     "label": "GitHub · dinov2",
     "url": "https://github.com/facebookresearch/dinov2"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/dino-v2-computer-vision-self-supervised-learning"
    }
   ],
   "why": "DINOv2 proved frozen, label-free features could beat supervised and CLIP-style pipelines on classification, depth and segmentation without fine-tuning. Relicensed Apache 2.0, it spread from medical imaging to forest-canopy mapping with the World Resources Institute and Mars-robotics work at NASA JPL, making it the practical default vision backbone of 2023-25.",
   "try": [
    {
     "label": "Official DINOv2 demo",
     "url": "https://dinov2.metademolab.com/"
    },
    {
     "label": "dinov2-base on Hugging Face",
     "url": "https://huggingface.co/facebook/dinov2-base"
    },
    {
     "label": "Run the segmentation notebook in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/dinov2/blob/main/notebooks/semantic_segmentation.ipynb"
    }
   ],
   "params": "ViT-S 21M / ViT-B 86M / ViT-L 300M / ViT-g 1.1B"
  },
  {
   "id": "dinov3",
   "name": "DINOv3",
   "full_name": "DINOv3",
   "tag": "One frozen backbone, from Instagram photos to Mars",
   "year": 2025,
   "year_label": "2025",
   "cat": "vision",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "dinov2"
   ],
   "desc": "Meta's flagship self-supervised backbone, released August 14, 2025: a 7B-parameter ViT trained on 1.7B images with no labels, plus distilled ViT-B/ViT-L and ConvNeXt variants, all under a commercial-use license. Produces state-of-the-art dense features for detection, segmentation, and depth — including a dedicated satellite-imagery model trained on 493M Maxar tiles.",
   "facts": [
    "A single frozen vision model that works from Instagram photos to satellite passes over the Amazon.",
    "The World Resources Institute uses DINOv3 to monitor deforestation — its satellite variant cut tree-canopy-height error in a Kenyan region from 4.1m to 1.2m vs DINOv2 — and NASA's Jet Propulsion Laboratory uses DINO models for Mars robotics.",
    "SAM 3D Body is built on a DINOv3 backbone.",
    "Frozen backbone, no labels, no fine-tuning — yet it beats specialist models on dense tasks like segmentation and depth.",
    "Trained on roughly 1.7B images with a 67-page report and 26 authors (Siméoni et al.)."
   ],
   "latest": "DINOv3 (Aug 14, 2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2508.10104"
    },
    {
     "label": "GitHub · dinov3",
     "url": "https://github.com/facebookresearch/dinov3"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/dinov3-self-supervised-vision-model"
    }
   ],
   "why": "DINOv3 scaled self-supervised vision to a 7B ViT trained on 1.7B images with no labels, producing dense features that beat specialist models on segmentation and depth while frozen. A satellite variant cut tree-canopy-height error in Kenya from 4.1m to 1.2m, and it now underpins SAM 3D Body and WRI deforestation monitoring.",
   "try": [
    {
     "label": "dinov3-vitb16 on Hugging Face",
     "url": "https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m"
    },
    {
     "label": "Full DINOv3 collection on Hugging Face",
     "url": "https://huggingface.co/collections/facebook/dinov3-68924841bd6b561778e31009"
    },
    {
     "label": "Run the PCA-features notebook in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/dinov3/blob/main/notebooks/pca.ipynb"
    }
   ],
   "params": "ViT-S 21M to ViT-7B 6.7B, plus ConvNeXt distillations"
  },
  {
   "id": "mask-rcnn",
   "name": "Mask R-CNN",
   "full_name": "Mask R-CNN",
   "tag": "The detection workhorse of an era",
   "year": 2017,
   "year_label": "2017",
   "cat": "vision",
   "size": 2,
   "status": "superseded",
   "open": true,
   "lineage": [],
   "desc": "FAIR's instance-segmentation architecture that dominated a research era: extending Faster R-CNN with a mask-prediction branch and RoIAlign, it won the ICCV 2017 Best Paper Award (Marr Prize) and remains a standard baseline and production workhorse for detection and segmentation nearly a decade later.",
   "facts": [
    "One of the most-cited AI papers ever (well over 30,000 citations); authored by Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick at FAIR."
   ],
   "latest": "Maintained today inside Detectron2",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/1703.06870"
    },
    {
     "label": "GitHub · detectron2",
     "url": "https://github.com/facebookresearch/detectron2"
    }
   ],
   "why": "Mask R-CNN added a mask branch and RoIAlign to Faster R-CNN and became the detection and instance-segmentation workhorse of an era — ICCV 2017 Best Paper, well over 30,000 citations, and still a production baseline nearly a decade later. Kaiming He's FAIR team set the template that Detectron2 and SAM later built on.",
   "try": [
    {
     "label": "Use it from torchvision",
     "url": "https://docs.pytorch.org/vision/stable/models/mask_rcnn.html"
    },
    {
     "label": "Train one in the Detectron2 Colab tutorial",
     "url": "https://colab.research.google.com/drive/16jcaJoc6bCFAQ96jDe2HwtXj7BMD_-m5"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/1703.06870"
    }
   ],
   "params": ""
  },
  {
   "id": "detectron2",
   "name": "Detectron2",
   "full_name": "Detectron2",
   "tag": "Detection and segmentation, boxed up as the field's default toolkit",
   "year": 2019,
   "year_label": "2019",
   "cat": "vision",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "mask-rcnn"
   ],
   "desc": "FAIR's modular PyTorch platform for object detection, instance segmentation, keypoints, and panoptic segmentation — the successor to the original Caffe2-based Detectron (2018). For years the default research and industry toolkit for detection, it implements Mask R-CNN, RetinaNet, DensePose, and dozens of successors.",
   "facts": [
    "One of GitHub's most-starred computer-vision libraries with tens of thousands of stars; spawned major derivatives like Detectron2Go for mobile deployment and underpinned countless CVPR papers."
   ],
   "latest": "Detectron2 (actively maintained on GitHub)",
   "links": [
    {
     "label": "GitHub · detectron2",
     "url": "https://github.com/facebookresearch/detectron2"
    }
   ],
   "why": "Detectron2 boxed up detection, instance and panoptic segmentation and keypoints as the field's default PyTorch toolkit: Mask R-CNN, RetinaNet, DensePose and dozens of successors in one modular codebase with about 35,000 GitHub stars. It underpinned countless CVPR papers, spawned Detectron2Go for mobile, and remains actively maintained.",
   "try": [
    {
     "label": "Official Colab tutorial",
     "url": "https://colab.research.google.com/drive/16jcaJoc6bCFAQ96jDe2HwtXj7BMD_-m5"
    },
    {
     "label": "Documentation",
     "url": "https://detectron2.readthedocs.io/en/latest/"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/detectron2"
    }
   ],
   "params": ""
  },
  {
   "id": "pytorch3d",
   "name": "PyTorch3D",
   "full_name": "PyTorch3D",
   "tag": "Fit a 3D mesh from photos by gradient descent",
   "year": 2020,
   "year_label": "2020",
   "cat": "vision",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "pytorch"
   ],
   "desc": "FAIR's library for deep learning with 3D data: batched meshes and point clouds, differentiable rendering, and loss functions that let gradients flow through the graphics pipeline. Released February 2020, it became the standard toolkit for research at the intersection of vision and graphics, including many of Meta's own 3D papers.",
   "facts": [
    "Its differentiable renderer made 'fit a 3D mesh from photos by gradient descent' a homework-sized problem; used across academia and in Meta's Codec Avatars research.",
    "Roughly 9k GitHub stars; its differentiable renderer let networks learn 3D shape from 2D photos years before NeRF tooling was commonplace."
   ],
   "latest": "Actively maintained (PyTorch3D on GitHub)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2007.08501"
    },
    {
     "label": "GitHub · pytorch3d",
     "url": "https://github.com/facebookresearch/pytorch3d"
    },
    {
     "label": "pytorch3d.org",
     "url": "https://pytorch3d.org"
    }
   ],
   "why": "PyTorch3D made 3D differentiable: batched meshes, point clouds and a differentiable renderer that lets gradients flow through the graphics pipeline, so 'fit a mesh from photos by gradient descent' became a homework-sized problem years before NeRF tooling was common. It is the standard toolkit at the vision-graphics intersection and underpins Meta's Codec Avatars research.",
   "try": [
    {
     "label": "Tutorials on pytorch3d.org",
     "url": "https://pytorch3d.org/tutorials/"
    },
    {
     "label": "Fit a textured mesh in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/pytorch3d/blob/main/docs/tutorials/fit_textured_mesh.ipynb"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/pytorch3d"
    }
   ],
   "params": ""
  },
  {
   "id": "imagebind",
   "name": "ImageBind",
   "full_name": "ImageBind",
   "tag": "One embedding space to bind six senses",
   "year": 2023,
   "year_label": "2023",
   "cat": "vision",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "The first model to bind six modalities — images, text, audio, depth, thermal, and IMU motion data — into one joint embedding space, trained only on image-paired data (May 2023). Enables cross-modal retrieval and arithmetic, like finding sounds that match an image, without any dataset pairing all modalities together.",
   "facts": [
    "Embedding arithmetic works across senses: image of a dove + sound of an engine retrieves images of birds near motorbikes.",
    "Code and weights are open under a non-commercial research license."
   ],
   "latest": "ImageBind (May 2023)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2305.05665"
    },
    {
     "label": "GitHub · ImageBind",
     "url": "https://github.com/facebookresearch/ImageBind"
    },
    {
     "label": "Live demo",
     "url": "https://imagebind.metademolab.com"
    }
   ],
   "why": "ImageBind was the first model to bind six modalities — images, text, audio, depth, thermal and IMU — into one embedding space using only image-paired data, so cross-modal retrieval and embedding arithmetic work between senses that were never paired in training. It prefigured the natively multimodal models Meta later shipped.",
   "try": [
    {
     "label": "Official ImageBind demo",
     "url": "https://imagebind.metademolab.com/"
    },
    {
     "label": "Code and weights on GitHub",
     "url": "https://github.com/facebookresearch/ImageBind"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2305.05665"
    }
   ],
   "params": "ImageBind-Huge (ViT-H/14 image encoder, ~630M)"
  },
  {
   "id": "chameleon",
   "name": "Chameleon",
   "full_name": "Chameleon",
   "tag": "Images and text in one token stream, trained from scratch",
   "year": 2024,
   "year_label": "2024",
   "cat": "vision",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR's early-fusion mixed-modal foundation model (May 2024): a single token-based transformer trained from scratch on interleaved images and text, able to understand and generate arbitrary interleavings of both. Meta released 7B and 34B weights under a research license in June 2024 — a precursor to natively multimodal product models.",
   "facts": [
    "The released checkpoints deliberately disabled image generation for safety; the paper reported the 34B model beating much larger models on interleaved-generation human evals."
   ],
   "latest": "Chameleon 7B/34B (Jun 2024 weight release; image-understanding only)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2405.09818"
    },
    {
     "label": "GitHub · chameleon",
     "url": "https://github.com/facebookresearch/chameleon"
    }
   ],
   "why": "Chameleon was Meta's first early-fusion model: one token-based transformer trained from scratch on interleaved images and text, able to understand and generate either, with the 34B version beating much larger models on interleaved-generation human evals. It is the research precursor of the natively multimodal Llama 4 and Muse Spark lines.",
   "try": [
    {
     "label": "chameleon-7b on Hugging Face",
     "url": "https://huggingface.co/facebook/chameleon-7b"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/chameleon"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2405.09818"
    }
   ],
   "params": "7B / 34B (hosted on Hugging Face as chameleon-30b)"
  },
  {
   "id": "perception",
   "name": "Perception Encoder & Perception Language Model",
   "full_name": "Perception Encoder & Perception Language Model",
   "tag": "A vision encoder whose best features hide in its middle layers",
   "year": 2025,
   "year_label": "2025",
   "cat": "vision",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR's April 2025 open perception stack: Perception Encoder (PE) is a large-scale vision encoder whose intermediate layers surprisingly hold the best embeddings, topping CLIP-style models on image and video tasks; the Perception Language Model (PLM, 1B/3B/8B) is a fully open, reproducible VLM released with PLM-VideoBench for fine-grained video understanding.",
   "facts": [
    "PLM was trained without distilling from proprietary models — a deliberate 'fully open and reproducible' stance — and shipped with 2.5M new human-labeled video QA samples, then the largest such release."
   ],
   "latest": "PE + PLM (Apr 2025)",
   "links": [
    {
     "label": "GitHub · perception_models",
     "url": "https://github.com/facebookresearch/perception_models"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/meta-fair-updates-perception-localization-reasoning"
    }
   ],
   "why": "Perception Encoder found that a CLIP-style encoder's best embeddings hide in its intermediate layers, not its output, and used that to top CLIP-class models on image and video tasks. PLM paired it with a fully open, non-distilled VLM and 2.5M new human-labeled video QA samples — a deliberate reproducibility stance against proprietary distillation.",
   "try": [
    {
     "label": "PE-Core-G14-448 on Hugging Face",
     "url": "https://huggingface.co/facebook/PE-Core-G14-448"
    },
    {
     "label": "Perception-LM-8B on Hugging Face",
     "url": "https://huggingface.co/facebook/Perception-LM-8B"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/perception_models"
    }
   ],
   "params": "PE Core: B/16 90M, L/14 320M, G/14 1.88B; PLM: 1B / 3B / 8B"
  },
  {
   "id": "sapiens",
   "name": "Sapiens",
   "full_name": "Sapiens",
   "tag": "308 keypoints per body, learned from 300 million human images",
   "year": 2024,
   "year_label": "2024",
   "cat": "vision",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "Reality Labs' Codec Avatars team's foundation models for seeing people (ECCV 2024): ViTs from 0.3B to 2B parameters pretrained on 300 million human images at native 1024px, delivering state-of-the-art 2D pose estimation, body-part segmentation, depth, and surface-normal prediction — the perceptual groundwork for photoreal telepresence avatars.",
   "facts": [
    "Its 308-keypoint pose vocabulary is far denser than standard 17-keypoint benchmarks — detailed enough to drive lifelike avatar hands and faces.",
    "5.4k GitHub stars.",
    "Beat prior SOTA by 7.6 mAP on pose and 17.1 mIoU on segmentation.",
    "Sapiens2's 5B model is reported as the highest-FLOPs vision transformer to date (~15.7 TFLOPs per inference).",
    "The lineage feeds Codec Avatars: the large-scale avatar pretraining work shares authors (e.g."
   ],
   "latest": "Sapiens (Aug 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2408.12569"
    },
    {
     "label": "GitHub · sapiens",
     "url": "https://github.com/facebookresearch/sapiens"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/facebook/sapiens2"
    },
    {
     "label": "meta.com",
     "url": "https://www.meta.com/emerging-tech/codec-avatars/sapiens"
    },
    {
     "label": "rawalkhirodkar.github.io",
     "url": "https://rawalkhirodkar.github.io/sapiens"
    }
   ],
   "why": "Sapiens pretrained ViTs up to 2B parameters on 300 million human images at native 1024px and set state of the art on pose, body-part segmentation, depth and normals — with a 308-keypoint vocabulary dense enough to drive hands and faces. It is the perception groundwork for Reality Labs' photoreal Codec Avatars.",
   "try": [
    {
     "label": "Sapiens models on Hugging Face",
     "url": "https://huggingface.co/facebook/sapiens"
    },
    {
     "label": "sapiens-pose-1b on Hugging Face",
     "url": "https://huggingface.co/facebook/sapiens-pose-1b"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2408.12569"
    }
   ],
   "params": "0.3B / 0.6B / 1B / 2B (Sapiens2 adds 5B)"
  },
  {
   "id": "cotracker",
   "name": "CoTracker",
   "full_name": "CoTracker",
   "tag": "Follow any pixel through a video — even when it disappears",
   "year": 2023,
   "year_label": "2023",
   "cat": "vision",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "Meta plus Oxford VGG's line of transformer point trackers that follow any pixel through a video jointly with its neighbors, staying locked on through occlusions and out-of-frame excursions. CoTracker3 (October 2024) simplified the architecture and used pseudo-labeled real videos to beat prior trackers with 1,000x less training data.",
   "facts": [
    "Tracking points jointly (co-tracking) rather than independently was the key trick; the mesmerizing rainbow point-trail demo videos made it a research-Twitter favorite."
   ],
   "latest": "CoTracker3 (Oct 15, 2024)",
   "links": [
    {
     "label": "GitHub · co-tracker",
     "url": "https://github.com/facebookresearch/co-tracker"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/facebook/cotracker3"
    }
   ],
   "why": "CoTracker made dense point tracking practical by tracking points jointly rather than independently, staying locked on through occlusions and off-screen excursions. CoTracker3 then matched or beat prior trackers with 1,000x less training data via pseudo-labeled real videos, and its rainbow point-trail demos made it a research favorite.",
   "try": [
    {
     "label": "Run the official demo in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/co-tracker/blob/main/notebooks/demo.ipynb"
    },
    {
     "label": "cotracker3 on Hugging Face",
     "url": "https://huggingface.co/facebook/cotracker3"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/co-tracker"
    }
   ],
   "params": ""
  },
  {
   "id": "vggt",
   "name": "VGGT",
   "full_name": "VGGT (Visual Geometry Grounded Transformer)",
   "tag": "CVPR 2025 Best Paper: geometry without the geometry pipeline",
   "year": 2025,
   "year_label": "2025",
   "cat": "vision",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "CVPR 2025 Best Paper winner from Meta AI and Oxford's Visual Geometry Group: a single feed-forward transformer that infers camera parameters, depth maps, point maps, and 3D point tracks from one to hundreds of images in seconds — replacing whole classical structure-from-motion pipelines with one network pass. Code is open source.",
   "facts": [
    "Picked as the top paper out of more than 13,000 CVPR submissions; reconstructs scenes in seconds where optimization-based SfM takes minutes to hours."
   ],
   "latest": "VGGT (CVPR 2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2503.11651"
    },
    {
     "label": "GitHub · vggt",
     "url": "https://github.com/facebookresearch/vggt"
    },
    {
     "label": "cs.ox.ac.uk",
     "url": "https://www.cs.ox.ac.uk/news/2456-full.html"
    }
   ],
   "why": "VGGT replaced the classical structure-from-motion pipeline with one feed-forward transformer that infers cameras, depth, point maps and tracks from one to hundreds of images in seconds. Chosen Best Paper from over 13,000 CVPR 2025 submissions, it reset expectations for how much 3D geometry a single network pass can recover.",
   "try": [
    {
     "label": "Try it in the Hugging Face Space",
     "url": "https://huggingface.co/spaces/facebook/vggt"
    },
    {
     "label": "VGGT-1B on Hugging Face",
     "url": "https://huggingface.co/facebook/VGGT-1B"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2503.11651"
    }
   ],
   "params": "1B"
  },
  {
   "id": "video-seal",
   "name": "Video Seal",
   "full_name": "Video Seal",
   "tag": "Invisible video watermarks that survive blur, crops and compression",
   "year": 2024,
   "year_label": "2024",
   "cat": "vision",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "An open framework and state-of-the-art model for invisible neural video watermarking (December 2024), aimed at AI-content provenance: temporal watermark propagation converts any image watermarker into an efficient video one without stamping every high-res frame. MIT-licensed with training code, baselines, and a public demo.",
   "facts": [
    "Watermarks survive blurring, cropping, and compression edits; released MIT-licensed as deepfake concerns peaked — Meta's provenance counterpart to its generative video push."
   ],
   "latest": "Video Seal (Dec 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2412.09492"
    },
    {
     "label": "GitHub · videoseal",
     "url": "https://github.com/facebookresearch/videoseal"
    },
    {
     "label": "Live demo",
     "url": "https://aidemos.meta.com/videoseal"
    }
   ],
   "why": "Video Seal made invisible neural video watermarking open and efficient: temporal propagation turns any image watermarker into a video one without stamping every high-res frame, and marks survive blur, crops and compression. Released MIT-licensed as deepfake concerns peaked, it is the provenance counterpart to Meta's generative video push.",
   "try": [
    {
     "label": "Official Video Seal demo",
     "url": "https://aidemos.meta.com/videoseal"
    },
    {
     "label": "pip install videoseal",
     "url": "https://pypi.org/project/videoseal"
    },
    {
     "label": "Run the official notebook in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/videoseal/blob/main/notebooks/colab.ipynb"
    }
   ],
   "params": ""
  },
  {
   "id": "ijepa",
   "name": "I-JEPA",
   "full_name": "I-JEPA",
   "tag": "LeCun's world-model bet, realized: predict representations, not pixels",
   "year": 2023,
   "year_label": "2023",
   "cat": "world",
   "size": 1,
   "status": "superseded",
   "open": true,
   "lineage": [],
   "desc": "The first realized piece of Yann LeCun's world-model vision (June 2023): an Image Joint-Embedding Predictive Architecture that learns by predicting abstract representations of masked image regions rather than pixels, delivering strong semantic features with far less compute than pixel-reconstruction or contrastive methods.",
   "facts": [
    "Trained a 632M-parameter ViT-Huge in under 72 hours on 16 A100s — an order of magnitude cheaper than comparable pixel-space methods at the time."
   ],
   "latest": "I-JEPA (Jun 2023, CVPR 2023)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2301.08243"
    },
    {
     "label": "GitHub · ijepa",
     "url": "https://github.com/facebookresearch/ijepa"
    }
   ],
   "why": "I-JEPA was the first working piece of LeCun's world-model program: learn by predicting abstract representations of masked regions rather than pixels. It reached strong semantic features while training a 632M ViT-Huge in under 72 hours on 16 A100s — an order of magnitude cheaper than pixel-reconstruction methods — and launched the JEPA family.",
   "try": [
    {
     "label": "ijepa_vith14_1k on Hugging Face",
     "url": "https://huggingface.co/facebook/ijepa_vith14_1k"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/ijepa"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2301.08243"
    }
   ],
   "params": "ViT-H/14 632M"
  },
  {
   "id": "vjepa",
   "name": "V-JEPA",
   "full_name": "V-JEPA",
   "tag": "Released the same month as Sora — and betting the opposite way",
   "year": 2024,
   "year_label": "2024",
   "cat": "world",
   "size": 1,
   "status": "superseded",
   "open": true,
   "lineage": [
    "ijepa"
   ],
   "desc": "Extends JEPA from images to video (February 2024): learns physical-world intuitions by predicting masked spatio-temporal regions of video in representation space, entirely self-supervised from unlabeled video. Achieved strong frozen-evaluation results on motion-centric benchmarks and set the stage for action-conditioned world models.",
   "facts": [
    "Released the same month as OpenAI's Sora, it embodied the opposite bet: LeCun argued predicting abstract representations, not generating pixels, is the path to machines that understand the world."
   ],
   "latest": "V-JEPA (Feb 2024); superseded by V-JEPA 2",
   "links": [
    {
     "label": "GitHub · jepa",
     "url": "https://github.com/facebookresearch/jepa"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/vjepa"
    }
   ],
   "why": "V-JEPA extended representation prediction from images to video, learning physical intuitions from unlabeled clips with no pixel generation. Released the same month as Sora, it embodied the opposite bet — abstract prediction over pixel synthesis — set strong frozen-evaluation results on motion benchmarks, and set the stage for action-conditioned world models.",
   "try": [
    {
     "label": "Code and models on GitHub",
     "url": "https://github.com/facebookresearch/jepa"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2404.08471"
    }
   ],
   "params": "ViT-L 300M / ViT-H 632M"
  },
  {
   "id": "vjepa2",
   "name": "V-JEPA 2",
   "full_name": "V-JEPA 2",
   "tag": "It learned physics from video — then drove a robot arm it had never met",
   "year": 2025,
   "year_label": "2025",
   "cat": "world",
   "size": 3,
   "status": "active",
   "open": true,
   "lineage": [
    "vjepa"
   ],
   "desc": "Meta's flagship world model, released June 11, 2025: trained on over 1 million hours of video plus only ~62 hours of robot data, its action-conditioned variant (V-JEPA 2-AC) enabled zero-shot planning on real Franka robot arms in unseen labs. Meta simultaneously released three physical-reasoning benchmarks (IntPhys 2, MVPBench, CausalVQA).",
   "facts": [
    "It learned physics by watching YouTube-scale video — then picked up objects with a robot arm it had never controlled.",
    "Zero-shot robot results: 100% success on reaching, 80% on cup pick-and-place, on robots the model never trained on.",
    "Planning ~15x faster than NVIDIA's Cosmos world model per the paper (Meta's press materials said up to 30x).",
    "V-JEPA 2 also jumped the EK100 action-anticipation record from 27.6 to 39.7 recall@5.",
    "Trained on 1M+ hours of video plus under 62 hours of unlabeled robot footage, then deployed zero-shot on robot arms in two different labs 'without collecting any data from the robots in these environments.' Also set records on Epic-Kitchens-100 action…"
   ],
   "latest": "V-JEPA 2.1 (Mar 16, 2026) — new recipe for high-quality, temporally consistent dense features; family now spans 80M to 2B parameters",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2506.09985"
    },
    {
     "label": "GitHub · vjepa2",
     "url": "https://github.com/facebookresearch/vjepa2"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/v-jepa-2-world-model-benchmarks"
    }
   ],
   "why": "V-JEPA 2 is the strongest evidence yet for JEPA-style world models: trained on over a million hours of video plus under 62 hours of robot data, its action-conditioned variant planned zero-shot on Franka arms in labs it never saw — 100% on reaching, 80% on pick-and-place — and set new action-anticipation records.",
   "try": [
    {
     "label": "vjepa2-vitl on Hugging Face",
     "url": "https://huggingface.co/facebook/vjepa2-vitl-fpc64-256"
    },
    {
     "label": "Run the official demo notebook in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/vjepa2/blob/main/notebooks/vjepa2_demo.ipynb"
    },
    {
     "label": "Full V-JEPA 2 collection on Hugging Face",
     "url": "https://huggingface.co/collections/facebook/v-jepa-2-6841bad8413014e185b497a6"
    }
   ],
   "params": "ViT-L 300M / ViT-H 600M / ViT-g 1B (V-JEPA 2.1 spans 80M-2B)"
  },
  {
   "id": "vljepa",
   "name": "VL-JEPA",
   "full_name": "VL-JEPA",
   "tag": "Predict the embedding, skip the tokens: JEPA outlives its champion",
   "year": 2025,
   "year_label": "2025",
   "cat": "world",
   "size": 1,
   "status": "active",
   "open": false,
   "lineage": [
    "vjepa2"
   ],
   "desc": "A vision-language JEPA (paper December 11, 2025): instead of autoregressively generating tokens like classic VLMs, it predicts continuous embeddings of target text. At only 1.6B parameters it surpasses CLIP, SigLIP2, and Meta's own Perception Encoder across eight video classification and eight retrieval benchmarks while matching classical VLMs on VQA.",
   "facts": [
    "Achieves its results with 50% fewer trainable parameters than equivalent token-space VLM training — published the month after Yann LeCun's departure was announced, proving the JEPA program outlives its champion at Meta."
   ],
   "latest": "VL-JEPA (Dec 2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2512.10942"
    }
   ],
   "why": "VL-JEPA showed JEPA's predict-the-embedding idea works for vision-language: at 1.6B parameters it beats CLIP, SigLIP2 and Meta's own Perception Encoder on video classification and retrieval while matching token-generating VLMs on VQA with 50% fewer trainable parameters. Published the month after LeCun's exit, it proved the JEPA program outlived its champion.",
   "try": [
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2512.10942"
    }
   ],
   "params": "1.6B"
  },
  {
   "id": "habitat",
   "name": "Habitat",
   "full_name": "Habitat",
   "tag": "A simulator where robots learn to walk before they exist",
   "year": 2019,
   "year_label": "2019",
   "cat": "world",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR's photorealistic 3D simulator for training embodied agents at thousands of frames per second. Habitat 1.0 (2019) tackled navigation; Habitat 2.0 (2021) added interactive rearrangement; Habitat 3.0 (2023) added simulated humanoid avatars and human-in-the-loop tools for studying human-robot collaboration in homes, alongside the HSSD and HM3D scene datasets.",
   "facts": [
    "Simulation speed was the founding obsession — agents can rack up years of navigation experience overnight.",
    "The annual Habitat Challenge became embodied AI's benchmark proving ground, and Habitat 3.0 lets a real human 'possess' the simulated human via VR to test robots against live people."
   ],
   "latest": "Habitat 3.0 (Oct 2023)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2310.13724"
    },
    {
     "label": "GitHub · habitat-lab",
     "url": "https://github.com/facebookresearch/habitat-lab"
    },
    {
     "label": "aihabitat.org",
     "url": "https://aihabitat.org"
    }
   ],
   "why": "Habitat made embodied AI trainable at scale: a photorealistic simulator running thousands of frames per second so agents can accumulate years of navigation experience overnight. Its annual Challenge became the field's proving ground, and Habitat 3.0 added simulated humans — even VR-possessed by real people — to study human-robot collaboration.",
   "try": [
    {
     "label": "aihabitat.org",
     "url": "https://aihabitat.org/"
    },
    {
     "label": "pip install habitat-lab",
     "url": "https://pypi.org/project/habitat-lab"
    },
    {
     "label": "conda install habitat-sim",
     "url": "https://anaconda.org/aihabitat/habitat-sim"
    }
   ],
   "params": ""
  },
  {
   "id": "partnr",
   "name": "PARTNR",
   "full_name": "PARTNR",
   "tag": "100,000 household chores that showed robot planners slow humans down",
   "year": 2024,
   "year_label": "2024",
   "cat": "world",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "habitat"
   ],
   "desc": "The largest benchmark for human-robot collaboration: 100,000 natural-language household tasks across 60 simulated houses and 5,800+ objects, built on Habitat 3.0. Released with a dataset of human demonstrations, it evaluates whether LLM-driven robot planners can coordinate with people on chores like tidying and cooking.",
   "facts": [
    "Findings were humbling: state-of-the-art LLM planners coordinated poorly with human partners, often slowing the human down rather than helping — exactly the gap the benchmark was built to expose."
   ],
   "latest": "PARTNR (Nov 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2411.00081"
    },
    {
     "label": "GitHub · partnr-planner",
     "url": "https://github.com/facebookresearch/partnr-planner"
    }
   ],
   "why": "PARTNR is the largest human-robot collaboration benchmark — 100,000 language-specified household tasks across 60 houses and 5,800+ objects — and its finding was humbling: state-of-the-art LLM planners coordinated so poorly that they often slowed the human down. It gave embodied-AI research a concrete target for genuinely helpful robots.",
   "try": [
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/partnr-planner"
    },
    {
     "label": "Episodes dataset on Hugging Face",
     "url": "https://huggingface.co/datasets/ai-habitat/partnr_episodes"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2411.00081"
    }
   ],
   "params": ""
  },
  {
   "id": "openeqa",
   "name": "OpenEQA",
   "full_name": "OpenEQA",
   "tag": "1,600 questions about real spaces — top VLMs were 'nearly blind'",
   "year": 2024,
   "year_label": "2024",
   "cat": "world",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR's open-vocabulary Embodied Question Answering benchmark (CVPR 2024): over 1,600 human-authored questions about 180+ real homes and offices, testing whether an agent that has 'seen' a space can answer questions about it — from episodic memory or by actively exploring. Includes an automatic LLM-based scorer validated against humans.",
   "facts": [
    "Headline result: on spatial questions, top VLMs were 'nearly blind' — access to visual input barely beat language-only guessing, while humans scored far higher.",
    "Framed by Meta as a milestone toward smart-glasses assistants that remember your world."
   ],
   "latest": "OpenEQA (Apr 2024)",
   "links": [
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/openeqa-embodied-question-answering-robotics-ar-glasses"
    },
    {
     "label": "open-eqa.github.io",
     "url": "https://open-eqa.github.io"
    }
   ],
   "why": "OpenEQA exposed how little vision-language models actually see: on questions about real homes and offices, GPT-4V scored 48.5% against 85.9% for humans, and on spatial questions barely beat text-only guessing. Its 1,600+ human-written questions became the reference test for the memory-equipped smart-glasses assistants Meta is building toward.",
   "try": [
    {
     "label": "Project page",
     "url": "https://open-eqa.github.io/"
    },
    {
     "label": "Dataset and code on GitHub",
     "url": "https://github.com/facebookresearch/open-eqa"
    },
    {
     "label": "Read the paper (PDF)",
     "url": "https://open-eqa.github.io/assets/pdfs/paper.pdf"
    }
   ],
   "params": ""
  },
  {
   "id": "motivo",
   "name": "Meta Motivo",
   "full_name": "Meta Motivo",
   "tag": "Puppet a physics-based humanoid with a prompt",
   "year": 2024,
   "year_label": "2024",
   "cat": "world",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "A first-of-its-kind behavioral foundation model (December 2024) for physics-based humanoid control: using the FB-CPR algorithm to embed states, motions, and rewards in one latent space, it performs whole-body tasks zero-shot — motion tracking, pose reaching, reward optimization — with human-like movement, and adapts to gravity or wind changes it never trained on.",
   "facts": [
    "Meta pitched it as a path to lifelike metaverse NPCs and democratized character animation; the interactive demo lets anyone puppet a physics humanoid with prompts."
   ],
   "latest": "Meta Motivo (Dec 2024)",
   "links": [
    {
     "label": "GitHub · metamotivo",
     "url": "https://github.com/facebookresearch/metamotivo"
    },
    {
     "label": "Live demo",
     "url": "https://metamotivo.metademolab.com"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/meta-fair-updates-agents-robustness-safety-architecture"
    }
   ],
   "why": "Meta Motivo is the first behavioral foundation model for physics-based humanoids: one latent space for states, motions and rewards lets it track motion, reach poses and optimize rewards zero-shot with human-like movement, even under gravity or wind it never trained on. It points toward prompt-driven character animation and lifelike NPCs.",
   "try": [
    {
     "label": "Puppet the humanoid in the official demo",
     "url": "https://metamotivo.metademolab.com/"
    },
    {
     "label": "metamotivo-M-1 on Hugging Face",
     "url": "https://huggingface.co/facebook/metamotivo-M-1"
    },
    {
     "label": "Run the tutorial notebook in Colab",
     "url": "https://colab.research.google.com/github/facebookresearch/metamotivo/blob/main/tutorial.ipynb"
    }
   ],
   "params": "M-1 288M / S-1 25M"
  },
  {
   "id": "locate3d",
   "name": "Meta Locate 3D",
   "full_name": "Meta Locate 3D",
   "tag": "'The flower vase near the TV console' — located in 3D",
   "year": 2025,
   "year_label": "2025",
   "cat": "world",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "An end-to-end model (April 2025) that localizes objects in 3D scenes from natural-language queries like 'the flower vase near the TV console' — operating directly on point clouds from RGB-D sensor streams via a 3D-JEPA encoder and language-conditioned decoder that outputs 3D bounding boxes and masks, ready for real robots.",
   "facts": [
    "Shipped with a new 130,000-annotation referring-expression dataset across ARKitScenes, ScanNet, and ScanNet++ (1,346 scenes) — roughly doubling the world's supply of such 3D language annotations."
   ],
   "latest": "Locate 3D (Apr 2025)",
   "links": [
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/meta-fair-updates-perception-localization-reasoning"
    }
   ],
   "why": "Locate 3D turns a sentence — 'the flower vase near the TV console' — into a 3D box and mask directly from RGB-D point clouds, using a self-supervised 3D-JEPA encoder, so robots can ground language in real scenes. Its 130,000-annotation dataset roughly doubled the world's supply of 3D referring expressions.",
   "try": [
    {
     "label": "Official Locate 3D demo",
     "url": "https://locate3d.atmeta.com/"
    },
    {
     "label": "locate-3d on Hugging Face",
     "url": "https://huggingface.co/facebook/locate-3d"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/locate-3d"
    }
   ],
   "params": "576M"
  },
  {
   "id": "ego4d",
   "name": "Ego4D",
   "full_name": "Ego4D",
   "tag": "3,670 hours of life, seen through human eyes",
   "year": 2021,
   "year_label": "2021",
   "cat": "world",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "The dataset that created modern egocentric AI research: 3,670 hours of daily-life first-person video from over 900 camera wearers across 74 locations in 9 countries, built by FAIR with a consortium of 13+ universities. Its benchmark suite (episodic memory, hands-and-objects, forecasting, social) still anchors annual competitions.",
   "facts": [
    "At release it was roughly 20x larger than any prior egocentric video dataset.",
    "It directly feeds Meta's smart-glasses ambitions — models that remember where you left your keys were an explicit motivating demo."
   ],
   "latest": "Ego4D (public release 2022; CVPR 2022 paper)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2110.07058"
    },
    {
     "label": "ego4d-data.org",
     "url": "https://ego4d-data.org"
    }
   ],
   "why": "Ego4D created modern egocentric AI: 3,670 hours of first-person daily life from 900+ wearers in nine countries — about 20x any prior egocentric dataset — with benchmarks for episodic memory, hands-and-objects, forecasting and social interaction. It is the training ground for the 'where did I leave my keys' assistants behind Meta's glasses.",
   "try": [
    {
     "label": "Get the dataset (start here)",
     "url": "https://ego4d-data.org/docs/start-here/"
    },
    {
     "label": "pip install ego4d (dataset CLI)",
     "url": "https://pypi.org/project/ego4d"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2110.07058"
    }
   ],
   "params": ""
  },
  {
   "id": "egoexo4d",
   "name": "Ego-Exo4D",
   "full_name": "Ego-Exo4D",
   "tag": "Cooking, bike repair, bouldering — seen first- and third-person at once",
   "year": 2023,
   "year_label": "2023",
   "cat": "world",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "ego4d"
   ],
   "desc": "A multimodal dataset pairing time-synced first-person (Aria glasses) and third-person video of skilled activities — cooking, bike repair, soccer, bouldering, dance, music — from 800+ participants in 13 cities. Built by FAIR, Project Aria, and 15 university partners; v2 spans about 1,300 hours across 5,035 captures with expert commentary annotations.",
   "facts": [
    "Every capture includes eye gaze, 7-channel audio, and expert 'coach' commentary — the dataset was explicitly designed for AI that could one day teach humans physical skills through AR glasses."
   ],
   "latest": "Ego-Exo4D v2 (2024; ~1,300 hours)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2311.18259"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/ego-exo4d-video-learning-perception"
    },
    {
     "label": "ego-exo4d-data.org",
     "url": "https://ego-exo4d-data.org"
    }
   ],
   "why": "Ego-Exo4D pairs synchronized first-person Aria and third-person video of skilled activities — cooking, bike repair, bouldering, music — from 800+ participants in 13 cities, with gaze, 7-channel audio and expert coach commentary. It is the dataset for AI that could one day teach physical skills through AR glasses.",
   "try": [
    {
     "label": "Dataset site",
     "url": "https://ego-exo4d-data.org/"
    },
    {
     "label": "Documentation and download",
     "url": "https://docs.ego-exo4d-data.org/"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2311.18259"
    }
   ],
   "params": ""
  },
  {
   "id": "aria",
   "name": "Project Aria",
   "full_name": "Project Aria",
   "tag": "Research glasses for teaching machines to see like us",
   "year": 2020,
   "year_label": "2020",
   "cat": "world",
   "size": 2,
   "status": "active",
   "open": false,
   "lineage": [
    "ego4d"
   ],
   "desc": "Reality Labs' sensor-packed research glasses program (no display — pure perception). Aria Gen 2, announced February 2025, packs four HDR global-shutter cameras, eye tracking, a heart-rate PPG sensor, custom low-power co-processor running on-device SLAM and hand tracking, in a 74-76g folding frame with about 8 hours of battery.",
   "facts": [
    "Hardware is application-only, but the tooling (projectaria_tools) and datasets (Aria Everyday Activities, Aria Digital Twin, Aria Gen 2 Pilot Dataset) are openly released.",
    "Gen 2 uses sub-GHz radio for sub-millisecond multi-device time sync and a contact microphone that separates the wearer's voice from bystanders'."
   ],
   "latest": "Aria Gen 2 — applications open; broad rollout to qualified researchers targeted Q2 2026",
   "links": [
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/aria-gen-2-research-glasses-under-the-hood-reality-labs"
    },
    {
     "label": "projectaria.com",
     "url": "https://www.projectaria.com"
    },
    {
     "label": "facebookresearch.github.io",
     "url": "https://facebookresearch.github.io/projectaria_tools/gen2"
    }
   ],
   "why": "Project Aria is how Meta collects the human perspective at scale: display-free research glasses with HDR cameras, eye tracking and on-device SLAM, loaned to researchers who in turn produce the Ego-Exo4D, Digital Twin and Everyday Activities datasets. Gen 2 adds sub-millisecond multi-device sync and a wearer-isolating contact microphone.",
   "try": [
    {
     "label": "Apply for Aria Gen 2 (projectaria.com)",
     "url": "https://www.projectaria.com/"
    },
    {
     "label": "pip install projectaria-tools",
     "url": "https://pypi.org/project/projectaria-tools"
    },
    {
     "label": "Browse the Aria Dataset Explorer",
     "url": "https://explorer.projectaria.com/"
    }
   ],
   "params": ""
  },
  {
   "id": "wav2vec2",
   "name": "wav2vec 2.0",
   "full_name": "wav2vec 2.0",
   "tag": "Speech recognition from raw audio, almost no labels",
   "year": 2020,
   "year_label": "2020",
   "cat": "speech",
   "size": 2,
   "status": "superseded",
   "open": true,
   "lineage": [],
   "desc": "Self-supervised speech representation learning (NeurIPS 2020): the model learns from raw unlabeled audio, then needs astonishingly little labeled data — 10 minutes of transcriptions plus 53K hours of unlabeled speech beat prior systems trained on 100x more labels. The foundation under MMS, Seamless, and 2025's Omnilingual ASR.",
   "facts": [
    "With just 10 minutes of labeled audio it reached 4.8/8.2 WER on LibriSpeech — a result that reshaped speech research economics.",
    "Its 2025 descendant, Omnilingual w2v 2.0, scaled the idea to 7B parameters."
   ],
   "latest": "wav2vec 2.0 (2020); descendants XLS-R, MMS, Omnilingual w2v 7B (2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2006.11477"
    },
    {
     "label": "GitHub · fairseq",
     "url": "https://github.com/facebookresearch/fairseq/tree/main/examples/wav2vec"
    }
   ],
   "why": "Showed speech recognition could learn mostly from unlabeled audio: 53K unlabeled hours plus only 10 minutes of transcripts beat systems trained on 100x more labels. That recipe became the backbone of Meta's MMS, Seamless and Omnilingual ASR, and of most self-supervised speech research since.",
   "try": [
    {
     "label": "Model on Hugging Face (base, 960h)",
     "url": "https://huggingface.co/facebook/wav2vec2-base-960h"
    },
    {
     "label": "Run it with Transformers (docs)",
     "url": "https://huggingface.co/docs/transformers/model_doc/wav2vec2"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2006.11477"
    }
   ],
   "params": "95M (BASE) / 317M (LARGE)"
  },
  {
   "id": "mms",
   "name": "Massively Multilingual Speech",
   "full_name": "Massively Multilingual Speech (MMS)",
   "tag": "Speech tech for 1,100+ languages",
   "year": 2023,
   "year_label": "2023",
   "cat": "speech",
   "size": 2,
   "status": "superseded",
   "open": true,
   "lineage": [
    "wav2vec2"
   ],
   "desc": "Speech-to-text and text-to-speech for 1,107 languages and spoken-language identification for 4,000+ — a 10x leap in coverage, built by pairing wav2vec 2.0 with a surprising data source: recordings of read religious texts (notably the Bible) available in thousands of languages. Halved Whisper's word error rate on its shared languages at the time.",
   "facts": [
    "The New Testament exists in audio in over 1,100 languages (~32 hours each) — this pun-free 'found dataset' unlocked the coverage.",
    "Despite religious-text training, analyses found no measurable theological bias in outputs."
   ],
   "latest": "MMS 1.0 (May 2023)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2305.13516"
    },
    {
     "label": "GitHub · fairseq",
     "url": "https://github.com/facebookresearch/fairseq/tree/main/examples/mms"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/multilingual-model-speech-recognition"
    }
   ],
   "why": "Took open speech recognition from roughly 100 languages to 1,107, with language identification for 4,000+, by pairing wav2vec 2.0 with recordings of read religious texts. It halved Whisper's word error rate on shared languages and made Meta the main supplier of speech tech for low-resource languages.",
   "try": [
    {
     "label": "ASR model on Hugging Face (mms-1b-all)",
     "url": "https://huggingface.co/facebook/mms-1b-all"
    },
    {
     "label": "Run ASR/TTS/LID with Transformers (docs)",
     "url": "https://huggingface.co/docs/transformers/model_doc/mms"
    },
    {
     "label": "TTS model on Hugging Face (English)",
     "url": "https://huggingface.co/facebook/mms-tts-eng"
    }
   ],
   "params": "1B (mms-1b-all)"
  },
  {
   "id": "nllb",
   "name": "No Language Left Behind",
   "full_name": "No Language Left Behind (NLLB-200)",
   "tag": "200 languages, one model, published in Nature",
   "year": 2022,
   "year_label": "2022",
   "cat": "speech",
   "size": 3,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "A single translation model covering 200 languages (202 language–script variants) — including 150+ low-resource ones like Asturian and Luganda — released July 6, 2022 with the FLORES-200 benchmark and training data. It averaged a 44% BLEU improvement over prior state of the art and later powered Wikipedia's Content Translation tool. Peer-reviewed in Nature (2024).",
   "facts": [
    "One model that can translate between Asturian and Assamese — a pair no commercial system had ever connected.",
    "The flagship is a 54.5B sparse mixture-of-experts; ~150 of its languages are low-resource, and it covers 55 African languages where fewer than 25 previously had wide translation support.",
    "Evaluated across more than 40,000 language directions.",
    "Techniques from the project feed the 25 billion+ translations served daily across Meta's apps, and NLLB powers Wikipedia's Content Translation tool for low-resource languages.",
    "Credited to 'NLLB Team' plus 38 named authors."
   ],
   "latest": "NLLB-200 54.5B MoE (Jul 2022); published in Nature Jun 2024",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2207.04672"
    },
    {
     "label": "GitHub · fairseq",
     "url": "https://github.com/facebookresearch/fairseq/tree/nllb"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/nllb-200-high-quality-machine-translation"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/no-language-left-behind"
    }
   ],
   "why": "Proved a single model could translate 200 languages, including 150+ low-resource ones, with a 44% average BLEU gain over the prior state of the art. Its FLORES-200 benchmark became the standard multilingual test set, NLLB powers Wikipedia's Content Translation tool, and the work was peer-reviewed in Nature in 2024.",
   "try": [
    {
     "label": "Model on Hugging Face (distilled 600M)",
     "url": "https://huggingface.co/facebook/nllb-200-distilled-600M"
    },
    {
     "label": "Run it with Transformers (docs)",
     "url": "https://huggingface.co/docs/transformers/model_doc/nllb"
    },
    {
     "label": "Read the Nature paper",
     "url": "https://www.nature.com/articles/s41586-024-07335-x"
    }
   ],
   "params": "600M / 1.3B / 3.3B dense; 54.5B MoE"
  },
  {
   "id": "seamless",
   "name": "SeamlessM4T & the Seamless family",
   "full_name": "SeamlessM4T & the Seamless family",
   "tag": "A real-life Babel Fish that keeps your voice",
   "year": 2023,
   "year_label": "2023",
   "cat": "speech",
   "size": 3,
   "status": "active",
   "open": true,
   "lineage": [
    "nllb",
    "mms"
   ],
   "desc": "One model for speech-to-speech, speech-to-text, text-to-speech, text-to-text translation and ASR across ~100 languages (Aug 2023). The v2 suite (Nov 2023) added SeamlessExpressive — preserving your tone, pauses, and emotion across languages — and SeamlessStreaming, translating with ~2-second latency before the speaker finishes. Published in Nature in January 2025.",
   "facts": [
    "A real-life Babel Fish — speak in one of ~100 languages, hear it in another, published on the pages of Nature.",
    "The seamless_communication repo has ~11,900 stars.",
    "SeamlessExpressive preserves your speech rate, pauses, emotion and vocal style across languages.",
    "SeamlessAlign's 470,000 hours is the largest open multimodal translation corpus ever mined — about 53 years of continuous audio.",
    "68 credited contributors ('Seamless Communication' team)."
   ],
   "latest": "SeamlessM4T v2 + SeamlessExpressive/Streaming (Nov 2023); Nature paper Jan 2025",
   "links": [
    {
     "label": "arXiv · 2308.11596",
     "url": "https://arxiv.org/abs/2308.11596"
    },
    {
     "label": "arXiv · 2312.05187",
     "url": "https://arxiv.org/abs/2312.05187"
    },
    {
     "label": "Nature · d41586 025 00497 2",
     "url": "https://www.nature.com/articles/d41586-025-00497-2"
    },
    {
     "label": "Nature · s41586 024 08359 z",
     "url": "https://www.nature.com/articles/s41586-024-08359-z"
    },
    {
     "label": "GitHub · seamless_communication",
     "url": "https://github.com/facebookresearch/seamless_communication"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/seamless-m4t"
    }
   ],
   "why": "Collapsed speech-to-speech, speech-to-text, text-to-speech and text translation into one model for roughly 100 languages, then added expressive translation that preserves tone and pauses and streaming translation at about two seconds of latency. Published in Nature in January 2025, it is the reference open system for speech translation.",
   "try": [
    {
     "label": "Model on Hugging Face (SeamlessM4T v2 Large)",
     "url": "https://huggingface.co/facebook/seamless-m4t-v2-large"
    },
    {
     "label": "Run it with Transformers (docs)",
     "url": "https://huggingface.co/docs/transformers/model_doc/seamless_m4t_v2"
    },
    {
     "label": "Meta's Seamless demo site",
     "url": "https://seamless.metademolab.com"
    }
   ],
   "params": "2.3B (SeamlessM4T v2 Large)"
  },
  {
   "id": "hokkien",
   "name": "Universal Speech Translator",
   "full_name": "Universal Speech Translator (Hokkien)",
   "tag": "Translating a language with no standard written form",
   "year": 2022,
   "year_label": "2022",
   "cat": "speech",
   "size": 1,
   "status": "superseded",
   "open": true,
   "lineage": [
    "wav2vec2"
   ],
   "desc": "First AI speech-to-speech translation system for a primarily oral language: Hokkien, spoken by millions in the Chinese diaspora but lacking a standard written form. Built with clever pivots through Mandarin text and released open-source (October 2022) as the first milestone of Meta's Universal Speech Translator program — a direct ancestor of Seamless.",
   "facts": [
    "Because Hokkien has no widely used writing system, the team could not use standard text-based pipelines — they trained on speech directly, using Mandarin as a pivot and working with native speakers to evaluate."
   ],
   "latest": "Hokkien speech-to-speech system (Oct 2022)",
   "links": [
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/ai-translation-hokkien"
    },
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2022/10/hokkien-ai-speech-translation"
    }
   ],
   "why": "The first AI speech-to-speech translation for a language with no standard written form, pivoting through Mandarin text and scored with syllable-level Tai-lo BLEU. Nearly half of the world's 7,000+ languages are primarily oral, so it broke the text-first assumption in machine translation and directly seeded the Seamless program.",
   "try": [
    {
     "label": "Model on Hugging Face (English to Hokkien)",
     "url": "https://huggingface.co/facebook/xm_transformer_s2ut_en-hk"
    },
    {
     "label": "Code and recipes (fairseq, ust branch)",
     "url": "https://github.com/facebookresearch/fairseq/tree/ust/examples/hokkien"
    },
    {
     "label": "Meta AI blog post",
     "url": "https://ai.meta.com/blog/ai-translation-hokkien"
    }
   ],
   "params": ""
  },
  {
   "id": "omnilingual-asr",
   "name": "Omnilingual ASR",
   "full_name": "Omnilingual ASR",
   "tag": "Speech recognition for 1,600+ languages — most for the first time",
   "year": 2025,
   "year_label": "2025",
   "cat": "speech",
   "size": 3,
   "status": "active",
   "open": true,
   "lineage": [
    "mms"
   ],
   "desc": "A FAIR-built suite of open-source speech-recognition models released November 10, 2025 that transcribes more than 1,600 languages — including ~500 low-resource languages never before supported by any ASR system — and extends zero-shot to 5,400+ languages from just a few paired examples. Released under Apache 2.0 in sizes from 300M to 7B parameters.",
   "facts": [
    "Covers nearly every spoken language with a known script — 1,600+ natively, 5,400+ via zero-shot in-context learning.",
    "Scales self-supervised speech pre-training to 7B parameters with an LLM-inspired decoder.",
    "A new language can be added with only a handful of paired audio-text examples, no retraining.",
    "Trained on 4.3 million hours of audio; achieves character error rate below 10% for 78% of supported languages.",
    "Effectively covers nearly every spoken language with a known script — and communities can add a new language with just a few examples."
   ],
   "latest": "Omnilingual ASR (Nov 10, 2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2511.09690"
    },
    {
     "label": "GitHub · omnilingual-asr",
     "url": "https://github.com/facebookresearch/omnilingual-asr"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/omnilingual-asr-open-source-multilingual-speech-recognition-for-1600-languages"
    }
   ],
   "why": "Expanded open speech recognition from MMS's 1,107 languages to more than 1,600, about 500 of them never transcribed by any ASR system, with zero-shot extension to 5,400+ from a few examples. Apache 2.0 weights from 300M to 7B and a pip package make it the most language-inclusive ASR anyone can deploy.",
   "try": [
    {
     "label": "Explore languages (Meta demo)",
     "url": "https://aidemos.atmeta.com/omnilingualasr"
    },
    {
     "label": "pip install omnilingual-asr",
     "url": "https://pypi.org/project/omnilingual-asr"
    },
    {
     "label": "Model on Hugging Face (7B LLM-ASR)",
     "url": "https://huggingface.co/facebook/omniASR-LLM-7B"
    }
   ],
   "params": "300M / 1B / 3B / 7B (CTC and LLM variants)"
  },
  {
   "id": "lang-partner",
   "name": "Language Technology Partner Program + BOUQuET",
   "full_name": "Language Technology Partner Program + BOUQuET",
   "tag": "Partners contribute underserved languages; the resulting models go open",
   "year": 2025,
   "year_label": "2025",
   "cat": "speech",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "nllb"
   ],
   "desc": "Program sourcing speech recordings, transcriptions, and translations for underserved languages from partners — governments, communities, researchers — in support of UNESCO's Decade of Indigenous Languages; resulting models are open-sourced. Announced alongside BOUQuET, an open multilingual translation benchmark. The Government of Nunavut joined early, contributing Inuktitut and Inuinnaqtun data.",
   "facts": [
    "Partners commit 10+ hours of transcribed speech and 200+ translated sentences per language.",
    "The Omnilingual ASR Corpus (Nov 2025), spanning 350 underserved languages, was curated with these global partners."
   ],
   "latest": "Announced Feb 7, 2025 (with UNESCO)",
   "links": [
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2025/02/announcing-language-technology-partner-program"
    },
    {
     "label": "TechCrunch",
     "url": "https://techcrunch.com/2025/02/07/meta-launches-new-program-to-improve-speech-and-translation-ai"
    }
   ],
   "why": "Turns Meta's low-resource language work into a data supply chain: partners contribute 10+ hours of transcribed speech and translated sentences, resulting models are open-sourced, and the effort supports UNESCO's Decade of Indigenous Languages. BOUQuET adds a paragraph-level, multi-way translation benchmark spanning 275 language varieties under CC-BY-4.0.",
   "try": [
    {
     "label": "BOUQuET dataset on Hugging Face",
     "url": "https://huggingface.co/datasets/facebook/bouquet"
    },
    {
     "label": "Read the BOUQuET paper (EMNLP 2025)",
     "url": "https://aclanthology.org/2025.emnlp-main.1400/"
    },
    {
     "label": "Program announcement (Meta Newsroom)",
     "url": "https://about.fb.com/news/2025/02/announcing-language-technology-partner-program"
    }
   ],
   "params": ""
  },
  {
   "id": "audiocraft",
   "name": "AudioCraft",
   "full_name": "AudioCraft (MusicGen, AudioGen, EnCodec)",
   "tag": "Type a vibe, get a song",
   "year": 2023,
   "year_label": "2023",
   "cat": "speech",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "Meta's open generative-audio library (August 2023): MusicGen turns text prompts into music, AudioGen generates environmental sounds, and EnCodec is the neural audio codec underneath. Later gained MAGNeT (non-autoregressive, faster generation) and JASCO (chord/beat-conditioned music). Code is MIT; model weights CC-BY-NC.",
   "facts": [
    "~23,600 GitHub stars — Meta's most popular audio repo.",
    "MusicGen was trained on 20,000 hours of licensed music.",
    "EnCodec compresses audio at rates that beat MP3 at 10x smaller bitrates in listening tests."
   ],
   "latest": "AudioCraft 1.x with MAGNeT & JASCO additions (2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2306.05284"
    },
    {
     "label": "GitHub · audiocraft",
     "url": "https://github.com/facebookresearch/audiocraft"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/audiocraft-musicgen-audiogen-encodec-generative-ai-audio"
    }
   ],
   "why": "The first widely used open text-to-music stack: MusicGen, AudioGen and the EnCodec codec shipped with MIT code and downloadable weights, turning music generation into reproducible research rather than a closed product. The official Hugging Face Space has drawn over 5,000 likes and the library seeded a wave of fine-tunes.",
   "try": [
    {
     "label": "Official MusicGen demo (Hugging Face Space)",
     "url": "https://huggingface.co/spaces/facebook/MusicGen"
    },
    {
     "label": "pip install audiocraft",
     "url": "https://pypi.org/project/audiocraft"
    },
    {
     "label": "Model on Hugging Face (musicgen-small)",
     "url": "https://huggingface.co/facebook/musicgen-small"
    }
   ],
   "params": "MusicGen 300M / 1.5B / 3.3B"
  },
  {
   "id": "voicebox",
   "name": "Voicebox",
   "full_name": "Voicebox",
   "tag": "So good at voices Meta wouldn't release it",
   "year": 2023,
   "year_label": "2023",
   "cat": "speech",
   "size": 1,
   "status": "closed",
   "open": false,
   "lineage": [],
   "desc": "First generative model to solve speech tasks it wasn't explicitly trained for (June 2023): text-guided flow matching enables voice editing, noise removal, style transfer across six languages, and voice cloning from a 2-second sample. Meta deliberately withheld the model and weights, citing voice-impersonation risks — publishing the research with an audio-watermark classifier instead.",
   "facts": [
    "A rare and showcase-worthy 'too dangerous to release' decision: 20x faster than prior diffusion-style speech models, could clone a voice from 2 seconds of audio, and Meta shipped a detector-classifier paper instead of the model."
   ],
   "latest": "Voicebox (Jun 2023, paper + demos only)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2306.15687"
    },
    {
     "label": "Live demo",
     "url": "https://voicebox.metademolab.com"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/voicebox-generative-ai-model-speech"
    }
   ],
   "why": "Showed one flow-matching model could do zero-shot TTS, editing, denoising and cross-lingual style transfer from a 2-second sample, beating VALL-E on word error rate (1.9% vs 5.9%) while up to 20x faster. Meta withheld the weights over impersonation risk, an early high-profile case of publishing without releasing.",
   "try": [
    {
     "label": "Official demo page (audio samples)",
     "url": "https://voicebox.metademolab.com"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2306.15687"
    }
   ],
   "params": ""
  },
  {
   "id": "audiobox",
   "name": "Audiobox",
   "full_name": "Audiobox",
   "tag": "Describe the voice and the room — it generates both",
   "year": 2023,
   "year_label": "2023",
   "cat": "speech",
   "size": 1,
   "status": "closed",
   "open": false,
   "lineage": [
    "voicebox"
   ],
   "desc": "Voicebox's successor unifying speech, sound-effect, and soundscape generation with natural-language prompts — 'a young woman speaks with a high pitch, in a large cathedral' — plus voice restyling that keeps the speaker but changes the acoustic scene. Released as interactive research demos with responsible-AI guardrails (audio watermarking, voice authentication) rather than open weights.",
   "facts": [
    "Combined a described voice with a described environment for the first time.",
    "Its demo site included interactive story-soundscape toys; Meta added automatic watermarking to every generated clip."
   ],
   "latest": "Audiobox (Dec 2023, research demos)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2312.15821"
    },
    {
     "label": "Live demo",
     "url": "https://audiobox.metademolab.com"
    }
   ],
   "why": "Extended Voicebox to sound effects and soundscapes and was the first model to combine voice prompts with text descriptions for freeform voice restyling, surpassing AudioLDM2, VoiceLDM and TANGO on quality. It shipped as a watermarked, research-only demo that Meta retired in February 2026, so today only the paper remains.",
   "try": [
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2312.15821"
    },
    {
     "label": "Meta AI blog post",
     "url": "https://ai.meta.com/blog/audiobox-generating-audio-voice-natural-language-prompts"
    }
   ],
   "params": ""
  },
  {
   "id": "spirit-lm",
   "name": "Spirit LM",
   "full_name": "Spirit LM",
   "tag": "Text and speech in one token stream — with the emotion intact",
   "year": 2024,
   "year_label": "2024",
   "cat": "speech",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "Meta's first open multimodal language model that freely mixes text and speech in one token stream (paper Feb 2024, weights Oct 2024). The Expressive variant adds pitch and style tokens so it can carry emotion — anger, surprise, excitement — across modalities, enabling speech that actually sounds like it means it. FAIR Noncommercial Research License.",
   "facts": [
    "Trained by interleaving text and speech tokens word-by-word from aligned corpora.",
    "Two flavors: Base (semantic speech tokens) and Expressive (adds pitch + style tokens) — an ancestor of today's natively speaking assistants."
   ],
   "latest": "Spirit LM Base & Expressive 7B (weights Oct 18, 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2402.05755"
    },
    {
     "label": "GitHub · spiritlm",
     "url": "https://github.com/facebookresearch/spiritlm"
    }
   ],
   "why": "Meta's first open model that interleaves spoken and written tokens in one stream, so it can continue a text prompt in speech or the reverse, and its Expressive variant carries pitch and style across modalities. It is a reference design for speech-native language models, released under a non-commercial research license.",
   "try": [
    {
     "label": "Code and checkpoint instructions (GitHub)",
     "url": "https://github.com/facebookresearch/spiritlm"
    },
    {
     "label": "Listen to generation samples",
     "url": "https://speechbot.github.io/spiritlm"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2402.05755"
    }
   ],
   "params": "7B"
  },
  {
   "id": "sam-audio",
   "name": "SAM Audio",
   "full_name": "SAM Audio (Segment Anything in Audio)",
   "tag": "Segment Anything — for sound",
   "year": 2025,
   "year_label": "2025",
   "cat": "speech",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "sam3"
   ],
   "desc": "Family of open-weight foundation models that isolate any sound from a complex mixture using text prompts, visual prompts (click the object making the sound in a video), or time-span prompts. A flow-matching diffusion transformer built on Meta's Perception Encoder Audiovisual (PE-AV), it achieves state-of-the-art results across speech, music, instrument, and general sound separation.",
   "facts": [
    "Announced Dec 16, 2025 (about.fb.com); arXiv:2512.18099 submitted Dec 19, 2025 (14 authors, first author Bowen Shi).",
    "Model sizes span ~500M to 3B parameters yet run faster than real time (RTF ≈ 0.7).",
    "Six main checkpoints on Hugging Face (facebook/sam-audio-{small,base,large} plus -tv variants tuned for target correctness/visual prompting) — gated access, ~19K downloads/month for sam-audio-large alone.",
    "GitHub repo has ~3.6K stars.",
    "Time-span prompting was billed by Meta as an industry first."
   ],
   "latest": "SAM Audio v1 (Dec 16, 2025) — checkpoints small/base/large plus -tv variants; no 1.1 or newer release as of Sep 2026 (no tagged GitHub releases; arXiv paper still v1)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2512.18099"
    },
    {
     "label": "GitHub · sam-audio",
     "url": "https://github.com/facebookresearch/sam-audio"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/facebook/sam-audio-large"
    },
    {
     "label": "Live demo",
     "url": "https://aidemos.meta.com/segment-anything/editor/segment-audio"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/sam-audio"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/sam-audio-segment-anything-in-audio"
    },
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2025/12/our-new-sam-audio-model-transforms-audio-editing"
    }
   ],
   "why": "Brings the Segment Anything idea to sound: describe a source, click the object making it in a video, or mark a time span, and it isolates that sound from a mixture with state-of-the-art results across speech, music and general audio. Open weights and a browser playground make source separation usable by editors, not only researchers.",
   "try": [
    {
     "label": "Try it in the Segment Anything Playground",
     "url": "https://aidemos.meta.com/segment-anything/editor/segment-audio"
    },
    {
     "label": "Model on Hugging Face (large)",
     "url": "https://huggingface.co/facebook/sam-audio-large"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/sam-audio"
    }
   ],
   "params": ""
  },
  {
   "id": "make-a-video",
   "name": "Make-A-Video",
   "full_name": "Make-A-Video",
   "tag": "Text-to-video before it was normal",
   "year": 2022,
   "year_label": "2022",
   "cat": "genmedia",
   "size": 1,
   "status": "superseded",
   "open": false,
   "lineage": [],
   "desc": "Meta's pioneering text-to-video system (September 2022), among the first to generate video from prompts without any paired text-video training data — it learned appearance from text-image pairs and motion from unlabeled video. A research showcase only; never released as a product or open model, but it kicked off the T2V race.",
   "facts": [
    "Announced days before Google's Imagen Video — September 2022 was the month text-to-video arrived; the flying-superhero-dog clip became its calling card."
   ],
   "latest": "Make-A-Video (Sep 2022); superseded by Emu Video, then Movie Gen",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2209.14792"
    },
    {
     "label": "makeavideo.studio",
     "url": "https://makeavideo.studio"
    }
   ],
   "why": "One of the first text-to-video systems, and notable for needing no paired text-video data: it learned appearance from text-image pairs and motion from unlabeled video. It was never released, but it kicked off the text-to-video race that led to Emu Video, Movie Gen and Meta's Muse Video.",
   "try": [
    {
     "label": "Official showcase site (samples)",
     "url": "https://makeavideo.studio"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2209.14792"
    }
   ],
   "params": ""
  },
  {
   "id": "emu",
   "name": "Emu",
   "full_name": "Emu",
   "tag": "A few thousand perfect photos beat millions of mediocre ones",
   "year": 2023,
   "year_label": "2023",
   "cat": "genmedia",
   "size": 1,
   "status": "superseded",
   "open": false,
   "lineage": [],
   "desc": "Meta's production image-generation model, unveiled at Connect in September 2023: a latent-diffusion model 'quality-tuned' on just a few thousand exceptionally curated images, which dramatically improved aesthetics. Emu powered Meta AI's image generation and the 'Imagine' features across Meta's apps — closed weights, product-first.",
   "facts": [
    "The counterintuitive finding: a few thousand hand-picked photos beat millions of mediocre ones for aesthetic alignment.",
    "Emu also powered Instagram's AI 'Restyle' and Backdrop editing features."
   ],
   "latest": "Emu (2023); superseded in products by Muse Image (2026)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2309.15807"
    }
   ],
   "why": "Demonstrated 'quality tuning': fine-tuning a latent diffusion model on only a few thousand exceptionally curated images dramatically raised aesthetic quality. Emu powered image generation in Meta AI and the Imagine features across Meta's apps for more than two years, until Muse Image replaced it in 2026.",
   "try": [
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2309.15807"
    },
    {
     "label": "Connect 2023 announcement (Meta Newsroom)",
     "url": "https://about.fb.com/news/2023/09/introducing-ai-powered-assistants-characters-and-creative-tools/"
    }
   ],
   "params": ""
  },
  {
   "id": "emu-video-edit",
   "name": "Emu Video & Emu Edit",
   "full_name": "Emu Video & Emu Edit",
   "tag": "Two Emu offshoots: 4-second clips and edit-by-instruction — demos only",
   "year": 2023,
   "year_label": "2023",
   "cat": "genmedia",
   "size": 1,
   "status": "superseded",
   "open": false,
   "lineage": [
    "emu",
    "make-a-video"
   ],
   "desc": "Twin research models announced November 2023. Emu Video generates 4-second 512px videos by factorizing text-to-video into image generation then image-conditioned video generation, winning head-to-head human evals against contemporaries. Emu Edit performs precise instruction-based image editing via multi-task training on 10 million samples. Demos only; weights never released.",
   "facts": [
    "In human evaluations Meta reported Emu Video was preferred over Runway Gen-2 and Pika at the time; Emu Edit's trick was recognizing the edit task type from the instruction itself."
   ],
   "latest": "Emu Video / Emu Edit (Nov 2023)",
   "links": [
    {
     "label": "arXiv · 2311.10709",
     "url": "https://arxiv.org/abs/2311.10709"
    },
    {
     "label": "arXiv · 2311.10089",
     "url": "https://arxiv.org/abs/2311.10089"
    },
    {
     "label": "Live demo",
     "url": "https://emu-video.metademolab.com"
    }
   ],
   "why": "Emu Video's factorized recipe (generate an image, then animate it) won head-to-head human evaluations against contemporaries, and Emu Edit showed instruction-based editing could be precise when trained as multi-task recognition plus generation on 10 million samples. Both fed directly into the production image and video features Meta shipped next.",
   "try": [
    {
     "label": "Emu Video demo site",
     "url": "https://emu-video.metademolab.com"
    },
    {
     "label": "Emu Edit showcase site",
     "url": "https://emu-edit.metademolab.com"
    },
    {
     "label": "Read the Emu Video paper",
     "url": "https://arxiv.org/abs/2311.10709"
    }
   ],
   "params": ""
  },
  {
   "id": "movie-gen",
   "name": "Movie Gen",
   "full_name": "Movie Gen",
   "tag": "A sentence in, a 16-second film with its own score out",
   "year": 2024,
   "year_label": "2024",
   "cat": "genmedia",
   "size": 3,
   "status": "superseded",
   "open": false,
   "lineage": [
    "emu-video-edit"
   ],
   "desc": "Meta's media-foundation research suite (October 2024): a 30B-parameter video model generating up to 16 seconds of 1080p footage at 16 fps, a 13B synchronized-audio model, personalized video from a single reference photo, and precise text-driven video editing. Published with detailed research and the Movie Gen Bench evaluation suite, but weights stayed closed.",
   "facts": [
    "Type a sentence, get a 16-second HD film with its own synchronized orchestral score — from one system.",
    "The personalization feature could cast you as the star of a generated clip from one photo.",
    "Meta said at announcement it was planned for Instagram in 2025 — the consumer path ultimately arrived via the Vibes feed and Edits app instead.",
    "A 92-page report with 88 credited authors; trained on the order of 100 million videos and 1 billion images using up to 6,144 H100 GPUs.",
    "Its personalization mode makes a video starring you from one reference photo."
   ],
   "latest": "Movie Gen (Oct 2024); succeeded by Muse Video (2026 preview)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2410.13720"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/movie-gen-media-foundation-models-generative-ai-video"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/movie-gen"
    },
    {
     "label": "Meta AI",
     "url": "https://ai.meta.com/static-resource/movie-gen-research-paper"
    }
   ],
   "why": "A 30B video model producing 16 seconds of 1080p footage, a 13B synchronized-audio model, personalization from one photo and text-driven editing, published with a detailed paper and the Movie Gen Bench prompts. Weights stayed closed, but it set the bar for Meta's Muse Video and for open evaluation of video generators.",
   "try": [
    {
     "label": "Official Movie Gen research page",
     "url": "https://ai.meta.com/research/movie-gen"
    },
    {
     "label": "Movie Gen Video Bench on Hugging Face",
     "url": "https://huggingface.co/datasets/meta-ai-for-media-research/movie_gen_video_bench"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2410.13720"
    }
   ],
   "params": "30B (video) / 13B (audio)"
  },
  {
   "id": "esm",
   "name": "ESM / ESMFold",
   "full_name": "ESM / ESMFold (Evolutionary Scale Modeling)",
   "tag": "617 million protein structures in two weeks",
   "year": 2019,
   "year_label": "2019-2022",
   "cat": "science",
   "size": 3,
   "status": "archived",
   "open": true,
   "lineage": [],
   "desc": "Meta's protein language models: ESM-2 (up to 15B parameters, 2022) learned protein structure from sequences alone, and ESMFold predicted atomic-level 3D structure ~60x faster than AlphaFold2 without multiple sequence alignments (Science, 2023). The core team left Meta in 2023 to found EvolutionaryScale, whose ESM3 (2024) continues the family — outside Meta.",
   "facts": [
    "A language model that never saw a 3D structure during pretraining learned to fold proteins — then mapped 617 million of them in a fortnight.",
    "ESMFold folded 617 million metagenomic proteins in two weeks on ~2,000 GPUs to launch the ESM Metagenomic Atlas (Nov 2022), later expanded to ~772 million structures.",
    "EvolutionaryScale launched in June 2024 with a reported $142M seed round — one of the biggest ever — and ESM3 generated esmGFP, a fluorescent protein it framed as '500 million years of evolution' away from known ones.",
    "Predicted 617 million protein structures in just two weeks on ~2,000 GPUs; structure emerged spontaneously in the language model's representations as it scaled from 8M to 15B parameters."
   ],
   "latest": "ESM-2/ESMFold (2022); Meta repo archived August 2024; ESM3 is EvolutionaryScale's, not Meta's",
   "links": [
    {
     "label": "Science",
     "url": "https://www.science.org/doi/10.1126/science.ade2574"
    },
    {
     "label": "GitHub · esm",
     "url": "https://github.com/facebookresearch/esm"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/protein-folding-esmfold-metagenomics"
    },
    {
     "label": "esmatlas.com",
     "url": "https://esmatlas.com"
    },
    {
     "label": "evolutionaryscale.ai",
     "url": "https://www.evolutionaryscale.ai/blog/esm3-release"
    }
   ],
   "why": "Showed protein language models trained on sequences alone learn structure: ESMFold predicted 3D structure about 60x faster than AlphaFold2 without multiple sequence alignments, enabling the ESM Metagenomic Atlas. The core team left to found EvolutionaryScale in 2023, and Meta archived the repository in August 2024.",
   "try": [
    {
     "label": "Browse the ESM Metagenomic Atlas",
     "url": "https://esmatlas.com"
    },
    {
     "label": "ESMFold on Hugging Face",
     "url": "https://huggingface.co/facebook/esmfold_v1"
    },
    {
     "label": "Fold a sequence in Colab (ColabFold ESMFold)",
     "url": "https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/ESMFold.ipynb"
    }
   ],
   "params": "ESM-2 8M to 15B; ESMFold 690M + 3B ESM-2 backbone"
  },
  {
   "id": "open-catalyst",
   "name": "Open Catalyst Project",
   "full_name": "Open Catalyst Project",
   "tag": "Hunting clean-energy catalysts with AI",
   "year": 2020,
   "year_label": "2020",
   "cat": "science",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "Collaboration between FAIR and Carnegie Mellon using AI to find catalysts for renewable-energy storage. The OC20 dataset (2020) — over 1.2 million DFT relaxations — plus OC22 and the experimental OCx24 release (November 2024) turned catalysis into a benchmark-driven ML field, with public leaderboards and NeurIPS competitions.",
   "facts": [
    "OC20 remains one of the largest chemistry ML datasets ever released; the project's open demo lets anyone simulate adsorption energies in the browser instead of running days of DFT."
   ],
   "latest": "Open Catalyst Experiments 2024 (OCx24, Nov 2024); models and demo now served via fairchem/UMA",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2010.09990"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/open-catalyst-simulations-experiments"
    },
    {
     "label": "opencatalystproject.org",
     "url": "https://opencatalystproject.org"
    }
   ],
   "why": "Turned catalyst discovery into a benchmark-driven machine-learning field: OC20's 1.2 million DFT relaxations, public leaderboards and NeurIPS competitions gave ML researchers a concrete climate problem, and OCx24 in 2024 closed the loop with real experiments. Its models now live on in fairchem and UMA.",
   "try": [
    {
     "label": "Project site",
     "url": "https://opencatalystproject.org"
    },
    {
     "label": "OC20 leaderboard",
     "url": "https://opencatalystproject.org/leaderboard.html"
    },
    {
     "label": "Interactive Open Catalyst demo",
     "url": "https://open-catalyst.metademolab.com/"
    }
   ],
   "params": ""
  },
  {
   "id": "opendac",
   "name": "OpenDAC / ODAC25",
   "full_name": "OpenDAC / ODAC25",
   "tag": "70 million quantum calculations for materials that pull CO₂ from air",
   "year": 2023,
   "year_label": "2023",
   "cat": "science",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "open-catalyst"
   ],
   "desc": "FAIR Chemistry's direct-air-capture project with Georgia Tech: datasets and models for CO₂-capturing metal-organic frameworks. ODAC23 provided ~38M DFT calculations on ~8,400 MOFs; ODAC25 (2025) scaled to nearly 70 million DFT calculations covering CO₂, H₂O, N₂ and O₂ adsorption in nearly 15,000 MOFs, including functionalized and synthetic frameworks.",
   "facts": [
    "ODAC25's ~70M single-point calculations make it the largest open MOF adsorption dataset — aimed squarely at making carbon-removal sorbent discovery an ML problem."
   ],
   "latest": "ODAC25 (2025)",
   "links": [
    {
     "label": "GitHub · fairchem",
     "url": "https://github.com/facebookresearch/fairchem"
    },
    {
     "label": "fair-chem.github.io",
     "url": "https://fair-chem.github.io/odac25"
    },
    {
     "label": "opencatalystproject.org",
     "url": "https://opencatalystproject.org"
    }
   ],
   "why": "Applies the Open Catalyst playbook to carbon removal: ODAC25's nearly 70 million DFT calculations on about 15,000 metal-organic frameworks, released CC-BY-4.0, let researchers train interatomic potentials to screen sorbents for direct air capture instead of running each framework through expensive quantum chemistry.",
   "try": [
    {
     "label": "Dataset and models on Hugging Face",
     "url": "https://huggingface.co/facebook/ODAC25"
    },
    {
     "label": "ODAC25 documentation",
     "url": "https://fair-chem.github.io/odac25"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2508.03162"
    }
   ],
   "params": ""
  },
  {
   "id": "omol25",
   "name": "Open Molecules 2025",
   "full_name": "Open Molecules 2025 (OMol25)",
   "tag": "100 million quantum-chemistry calculations, given away",
   "year": 2025,
   "year_label": "2025",
   "cat": "science",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "open-catalyst"
   ],
   "desc": "A landmark quantum-chemistry dataset released May 2025: over 100 million DFT calculations spanning biomolecules, electrolytes and metal complexes, with systems up to 350 atoms — an order of magnitude beyond typical 20-30-atom calculations. Built to train neural network potentials that approach DFT accuracy at a fraction of the cost, and released openly on Hugging Face.",
   "facts": [
    "Producing OMol25 consumed billions of CPU core-hours of density functional theory — one of the largest compute donations to open chemistry ever made."
   ],
   "latest": "OMol25 (May 2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2505.08762"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/facebook/OMol25"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/meta-fair-science-new-open-source-releases"
    }
   ],
   "why": "With over 100 million DFT calculations on systems up to 350 atoms, spanning biomolecules, electrolytes and metal complexes, OMol25 is an order of magnitude beyond typical quantum-chemistry datasets. It is the training bed for UMA and for neural network potentials that approach DFT accuracy at a fraction of the cost.",
   "try": [
    {
     "label": "Dataset and models on Hugging Face",
     "url": "https://huggingface.co/facebook/OMol25"
    },
    {
     "label": "Try the UMA demo trained on it",
     "url": "https://huggingface.co/spaces/facebook/fairchem_uma_demo"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2505.08762"
    }
   ],
   "params": ""
  },
  {
   "id": "uma",
   "name": "UMA — Universal Models for Atoms",
   "full_name": "UMA — Universal Models for Atoms",
   "tag": "~10,000x faster than DFT — a universal model for atoms",
   "year": 2025,
   "year_label": "2025",
   "cat": "science",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "open-catalyst",
    "omol25",
    "omat24",
    "opendac"
   ],
   "desc": "A family of machine-learned interatomic potentials trained on about half a billion 3D atomic structures across molecules, materials and catalysts — FAIR Chemistry's 'one model for all of chemistry.' Using a Mixture of Linear Experts architecture, UMA matches DFT-level end results roughly 10,000x faster, and powers downstream tools like FastCSP for crystal-structure prediction (August 2025).",
   "facts": [
    "Trained on ~500 million unique 3D structures — likely the largest atomistic training corpus assembled; October 2025 added multi-node/multi-GPU and LAMMPS interfaces for large-scale molecular dynamics."
   ],
   "latest": "UMA-1.2 (March 2026) — ~50% faster, ~40% more accurate on the OMol25 test set",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2506.23971"
    },
    {
     "label": "GitHub · fairchem",
     "url": "https://github.com/facebookresearch/fairchem"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/uma-a-family-of-universal-models-for-atoms"
    }
   ],
   "why": "One interatomic potential for molecules, materials and catalysts, trained on about half a billion 3D structures with a mixture-of-linear-experts design that adds capacity without slowing inference. It delivers DFT-level results roughly 10,000x faster and powers downstream tools such as FastCSP for crystal-structure prediction.",
   "try": [
    {
     "label": "Official UMA demo (Hugging Face Space)",
     "url": "https://huggingface.co/spaces/facebook/fairchem_uma_demo"
    },
    {
     "label": "Checkpoints on Hugging Face",
     "url": "https://huggingface.co/facebook/UMA"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2506.23971"
    }
   ],
   "params": "UMA-M: 1.4B total, ~50M active per structure"
  },
  {
   "id": "omat24",
   "name": "Open Materials 2024",
   "full_name": "Open Materials 2024 (OMat24)",
   "tag": "Open data that topped Matbench Discovery for predicting stable crystals",
   "year": 2024,
   "year_label": "2024",
   "cat": "science",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "open-catalyst"
   ],
   "desc": "An open inorganic-materials dataset of over 100 million DFT calculations plus pretrained EquiformerV2-based interatomic potentials, released October 2024. The models set state-of-the-art results on the Matbench Discovery leaderboard for predicting stable materials, and the permissively licensed dataset became a foundation for the materials-discovery ML community.",
   "facts": [
    "On release, OMat24-trained models topped Matbench Discovery — open data beating closed efforts at predicting which hypothetical crystals can actually exist."
   ],
   "latest": "OMat24 (October 2024); successor capabilities folded into UMA",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2410.12771"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/facebook/OMat24"
    },
    {
     "label": "rowansci.com",
     "url": "https://rowansci.com/features/omat24"
    }
   ],
   "why": "Over 100 million inorganic DFT calculations plus pretrained potentials that took the top of the Matbench Discovery leaderboard for predicting stable materials. Released under a permissive license, it became a foundation dataset for materials-discovery ML, and its capabilities were folded into UMA.",
   "try": [
    {
     "label": "Dataset and models on Hugging Face",
     "url": "https://huggingface.co/facebook/OMat24"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2410.12771"
    },
    {
     "label": "FAIR Chemistry docs",
     "url": "https://fair-chem.github.io"
    }
   ],
   "params": "EquiformerV2 31M / 86M / 153M; eSEN 30M"
  },
  {
   "id": "fairchem",
   "name": "fairchem",
   "full_name": "fairchem",
   "tag": "One pip install for five generations of FAIR chemistry",
   "year": 2024,
   "year_label": "2024",
   "cat": "science",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "uma"
   ],
   "desc": "FAIR Chemistry's consolidated open-source library — the successor to the ocp codebase — unifying data, training and inference for Open Catalyst, OpenDAC, OMat24, OMol25 and UMA. Provides an ASE calculator interface so any computational chemist can drop UMA into existing workflows, plus LAMMPS and multi-GPU support for production molecular dynamics.",
   "facts": [
    "One pip install now covers five generations of FAIR chemistry datasets and models — the connective tissue of Meta's AI-for-science program."
   ],
   "latest": "fairchem v2.x (2025-2026), serving UMA-1.2 models",
   "links": [
    {
     "label": "GitHub · fairchem",
     "url": "https://github.com/facebookresearch/fairchem"
    },
    {
     "label": "fair-chem.github.io",
     "url": "https://fair-chem.github.io"
    }
   ],
   "why": "The single library behind Open Catalyst, OpenDAC, OMat24, OMol25 and UMA: a pip install and an ASE calculator interface let any computational chemist drop Meta's models into existing workflows, with LAMMPS and multi-GPU support for production molecular dynamics. It is how FAIR Chemistry's datasets become tools people run.",
   "try": [
    {
     "label": "pip install fairchem-core",
     "url": "https://pypi.org/project/fairchem-core"
    },
    {
     "label": "Quickstart (UMA with ASE)",
     "url": "https://fair-chem.github.io/quickstart"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/fairchem"
    }
   ],
   "params": ""
  },
  {
   "id": "brain-decoding",
   "name": "Brain & AI: Decoding Perception from MEG/EEG",
   "full_name": "Brain & AI: Decoding Perception from MEG/EEG",
   "tag": "Reconstructing what you see from brain signals",
   "year": 2022,
   "year_label": "2022-2023",
   "cat": "science",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR's neuroscience program showed that self-supervised AI models align with brain activity — and can decode it. A 2023 Nature Machine Intelligence paper decoded perceived speech from non-invasive MEG/EEG using wav2vec 2.0-style contrastive learning; a companion 2023 system reconstructed seen images from MEG signals in near real time using DINOv2 embeddings.",
   "facts": [
    "From 3 seconds of MEG activity, the speech decoder identified the matching audio segment from over 1,500 candidates with up to 41% top-10 accuracy averaged across participants — using no implants at all."
   ],
   "latest": "Image decoding from MEG (October 2023); speech-perception decoding (2022 preprint, published 2023)",
   "links": [
    {
     "label": "arXiv · 2208.12266",
     "url": "https://arxiv.org/abs/2208.12266"
    },
    {
     "label": "arXiv · 2310.19812",
     "url": "https://arxiv.org/abs/2310.19812"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/meta-fair-science-new-open-source-releases"
    }
   ],
   "why": "Showed self-supervised AI representations align with brain activity well enough to decode it: contrastive learning between wav2vec 2.0 speech embeddings and MEG/EEG identified perceived speech, and DINOv2 embeddings reconstructed seen images from MEG in near real time. Code covering 175 volunteers and 160+ hours of recordings is open.",
   "try": [
    {
     "label": "Code on GitHub (brainmagick)",
     "url": "https://github.com/facebookresearch/brainmagick"
    },
    {
     "label": "Speech decoding paper",
     "url": "https://arxiv.org/abs/2208.12266"
    },
    {
     "label": "Image decoding paper",
     "url": "https://arxiv.org/abs/2310.19812"
    }
   ],
   "params": ""
  },
  {
   "id": "brain2qwerty",
   "name": "Brain2Qwerty",
   "full_name": "Brain2Qwerty",
   "tag": "Typing decoded from non-invasive brain recordings",
   "year": 2025,
   "year_label": "2025",
   "cat": "science",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "brain-decoding"
   ],
   "desc": "A non-invasive brain-to-text system from FAIR and the Basque Center on Cognition, Brain and Language (February 2025): as participants type sentences, a convolution-transformer-language-model stack decodes the text from MEG or EEG alone. MEG reached a 32% character error rate on average — 19% for the best participant — without surgery.",
   "facts": [
    "The open repo has ~900+ GitHub stars; v2 decodes full sentences from continuous MEG in a streaming fashion — an early glimpse of typing-free communication for people who cannot speak or move."
   ],
   "latest": "Brain2Qwerty v2 — trained on 10x more data per participant, up to 78% word accuracy for the best participant",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2502.17480"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/brain-to-text-decoding-a-non-invasive-approach-via-typing"
    },
    {
     "label": "facebookresearch.github.io",
     "url": "https://facebookresearch.github.io/brain2qwerty"
    }
   ],
   "why": "Non-invasive brain-to-text without surgery: decoding typed sentences from MEG reached a 32% character error rate (19% for the best participant) in v1, and v2 lifted the best participant to 78% word accuracy, with accuracy scaling log-linearly with data. Published in Nature Neuroscience with code open.",
   "try": [
    {
     "label": "Project page with results explorer",
     "url": "https://facebookresearch.github.io/brain2qwerty"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/brain2qwerty"
    },
    {
     "label": "Read the Nature Neuroscience paper",
     "url": "https://www.nature.com/articles/s41593-026-02303-2"
    }
   ],
   "params": ""
  },
  {
   "id": "pytorch",
   "name": "PyTorch",
   "full_name": "PyTorch",
   "tag": "The framework the AI world runs on",
   "year": 2016,
   "year_label": "2016",
   "cat": "oss",
   "size": 3,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "The dominant deep-learning framework, created at FAIR and open-sourced in early 2017. Its define-by-run design won over researchers; PyTorch 2.0 (2023) added torch.compile for speed without losing flexibility. Governance moved to the Linux Foundation's PyTorch Foundation in September 2022, which became an umbrella foundation in May 2025 hosting vLLM, DeepSpeed and Ray.",
   "facts": [
    "102.7k GitHub stars and 29.1k forks (Sept 2026).",
    "The large majority of ML research papers that specify a framework use PyTorch, and most of Hugging Face's 3M+ public models ship PyTorch weights.",
    "PyTorch 2.11 (Mar 2026) added a FlashAttention-4 FlexAttention backend; 2.13 (Jul 2026) brought FlexAttention to Apple Silicon with ~12x speedups."
   ],
   "latest": "PyTorch 2.14.0 (Sept 2, 2026); ~2-month release cadence through 2026",
   "links": [
    {
     "label": "GitHub · pytorch",
     "url": "https://github.com/pytorch/pytorch"
    },
    {
     "label": "PyTorch blog · press release pytorch foundati",
     "url": "https://pytorch.org/blog/press-release-pytorch-foundation-expands-welcomes-projects-vllm-deepspeed"
    },
    {
     "label": "PyTorch blog · pytorch 2 11 release blog",
     "url": "https://pytorch.org/blog/pytorch-2-11-release-blog"
    },
    {
     "label": "pytorch.org",
     "url": "https://pytorch.org"
    }
   ],
   "why": "The dominant deep-learning framework: its define-by-run design won researchers over, torch.compile in 2.0 closed the speed gap, and most modern AI papers and open models ship as PyTorch code. Meta handed governance to the Linux Foundation in 2022, and the PyTorch Foundation now also hosts vLLM, DeepSpeed and Ray.",
   "try": [
    {
     "label": "Install PyTorch",
     "url": "https://pytorch.org/get-started/locally"
    },
    {
     "label": "Official tutorials",
     "url": "https://docs.pytorch.org/tutorials"
    },
    {
     "label": "torch on PyPI",
     "url": "https://pypi.org/project/torch"
    }
   ],
   "params": ""
  },
  {
   "id": "torch-domain",
   "name": "PyTorch Domain Libraries",
   "full_name": "PyTorch Domain Libraries (torchvision, torchaudio, torchtune)",
   "tag": "torchvision, torchaudio, torchtune — the batteries that came with PyTorch",
   "year": 2017,
   "year_label": "2017-2026",
   "cat": "oss",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "pytorch"
   ],
   "desc": "The official ecosystem around PyTorch: torchvision (datasets, models, transforms for vision, since 2017), torchaudio (audio I/O and pipelines), and torchtune (2024, PyTorch-native LLM fine-tuning). Torchtune development wound down in 2025 with 150+ contributors credited; torchvision and torchaudio still track every PyTorch release. Several repos now live under the meta-pytorch GitHub org.",
   "facts": [
    "Torchtune reached 5.8k stars before winding down in 2025 (see issue #2883, 'The future of torchtune').",
    "Pretrained torchvision models were the de facto ImageNet backbone zoo for a generation of CV papers."
   ],
   "latest": "torchvision/torchaudio track PyTorch 2.13 (July 2026); torchtune final line v0.6, wound down 2025",
   "links": [
    {
     "label": "GitHub · vision",
     "url": "https://github.com/pytorch/vision"
    },
    {
     "label": "GitHub · torchtune",
     "url": "https://github.com/meta-pytorch/torchtune"
    },
    {
     "label": "PyTorch blog",
     "url": "https://pytorch.org/blog/torchtune-fine-tune-llms"
    }
   ],
   "why": "torchvision and torchaudio are the standard way researchers load datasets, pretrained models and transforms for vision and audio, tracking every PyTorch release since 2017. torchtune brought PyTorch-native LLM fine-tuning recipes before winding down in 2025 with more than 150 contributors credited.",
   "try": [
    {
     "label": "torchvision on PyPI",
     "url": "https://pypi.org/project/torchvision"
    },
    {
     "label": "torchaudio on PyPI",
     "url": "https://pypi.org/project/torchaudio"
    },
    {
     "label": "torchtune on PyPI",
     "url": "https://pypi.org/project/torchtune"
    }
   ],
   "params": ""
  },
  {
   "id": "executorch",
   "name": "ExecuTorch",
   "full_name": "ExecuTorch",
   "tag": "PyTorch in your pocket",
   "year": 2023,
   "year_label": "2023",
   "cat": "oss",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "pytorch"
   ],
   "desc": "PyTorch's on-device inference runtime for phones, wearables and embedded hardware, announced in 2023 and reaching 1.0 general availability in October 2025 with broad CPU/GPU/NPU backend support (including Arm SME2 via KleidiAI). It lets PyTorch models run efficiently at the edge and underpins on-device AI experiences across Meta's apps and devices.",
   "facts": [
    "Went from experimental preview (2023) to 1.0 GA (Oct 2025) with hardware partners including Arm, Apple, Qualcomm and MediaTek shipping dedicated backends."
   ],
   "latest": "ExecuTorch 1.1.0 (January 2026); 1.0 GA shipped October 2025",
   "links": [
    {
     "label": "GitHub · executorch",
     "url": "https://github.com/pytorch/executorch"
    },
    {
     "label": "PyTorch blog",
     "url": "https://pytorch.org/blog/introducing-executorch-1-0"
    },
    {
     "label": "newsroom.arm.com",
     "url": "https://newsroom.arm.com/news/executorch-1-0-ga-release-edge-ai"
    }
   ],
   "why": "PyTorch's answer to on-device inference: export a model once and run it on phone CPUs, GPUs and NPUs from Arm, Qualcomm, MediaTek and Apple without leaving the PyTorch ecosystem. It reached 1.0 in October 2025 and underpins on-device AI across Meta's apps and devices, including the quantized Llama 3.2 models.",
   "try": [
    {
     "label": "Documentation",
     "url": "https://docs.pytorch.org/executorch/stable/index.html"
    },
    {
     "label": "pip install executorch",
     "url": "https://pypi.org/project/executorch"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/pytorch/executorch"
    }
   ],
   "params": ""
  },
  {
   "id": "faiss",
   "name": "faiss",
   "full_name": "faiss",
   "tag": "The quiet backbone of the vector-database era",
   "year": 2017,
   "year_label": "2017",
   "cat": "oss",
   "size": 3,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "The library for efficient similarity search and clustering of dense vectors, with CPU and GPU indexes that scale to billions of vectors. Open-sourced by FAIR in 2017, faiss became foundational plumbing for the vector-database and RAG era — its algorithms are embedded in OpenSearch, Milvus and many commercial vector stores.",
   "facts": [
    "41k GitHub stars (Sept 2026) — the most-starred facebookresearch repo after Segment Anything.",
    "The 2017 launch demonstrated billion-scale GPU k-NN search; the 2024 paper 'The Faiss library' (arXiv 2401.08281) documents a decade of tricks.",
    "Practically every RAG stack touches faiss or an index it inspired."
   ],
   "latest": "v1.15.0 (Aug 2026)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2401.08281"
    },
    {
     "label": "GitHub · faiss · faiss",
     "url": "https://github.com/facebookresearch/faiss"
    },
    {
     "label": "GitHub · faiss · wiki",
     "url": "https://github.com/facebookresearch/faiss/wiki"
    }
   ],
   "why": "The reference library for nearest-neighbor search over billions of dense vectors on CPU or GPU. Open-sourced in 2017, it became the plumbing of the vector-database and RAG era: its indexes are embedded in OpenSearch, Milvus and many commercial vector stores, and a faiss index is still the default first step for embedding search.",
   "try": [
    {
     "label": "pip install faiss-cpu",
     "url": "https://pypi.org/project/faiss-cpu"
    },
    {
     "label": "Getting started (wiki)",
     "url": "https://github.com/facebookresearch/faiss/wiki/Getting-started"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/faiss"
    }
   ],
   "params": ""
  },
  {
   "id": "xformers",
   "name": "xFormers",
   "full_name": "xFormers",
   "tag": "Exact attention without the O(n²) memory bill",
   "year": 2021,
   "year_label": "2021",
   "cat": "oss",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "pytorch"
   ],
   "desc": "Hackable, optimized Transformer building blocks, best known for memory_efficient_attention — exact attention without the O(n²) memory bottleneck, dispatching to the best kernel (including FlashAttention) for the hardware. It became the default speed upgrade for Stable Diffusion and countless training stacks before much of it was absorbed into PyTorch itself.",
   "facts": [
    "11k GitHub stars.",
    "Photoroom measured up to ~100% faster Stable Diffusion image generation just by enabling xFormers' memory-efficient attention."
   ],
   "latest": "v0.0.35 (February 2026), with stable wheels for PyTorch 2.10+",
   "links": [
    {
     "label": "GitHub · xformers",
     "url": "https://github.com/facebookresearch/xformers"
    },
    {
     "label": "facebookresearch.github.io",
     "url": "https://facebookresearch.github.io/xformers"
    }
   ],
   "why": "Made memory-efficient exact attention a one-line drop-in, dispatching to the best kernel for the hardware, including FlashAttention. It was the default speed and memory upgrade for Stable Diffusion and countless training stacks, and much of its work was later absorbed into PyTorch itself.",
   "try": [
    {
     "label": "pip install xformers",
     "url": "https://pypi.org/project/xformers"
    },
    {
     "label": "Documentation",
     "url": "https://facebookresearch.github.io/xformers"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/xformers"
    }
   ],
   "params": ""
  },
  {
   "id": "fasttext",
   "name": "fastText",
   "full_name": "fastText",
   "tag": "Language detection for 176 languages in under a megabyte",
   "year": 2016,
   "year_label": "2016",
   "cat": "oss",
   "size": 1,
   "status": "archived",
   "open": true,
   "lineage": [],
   "desc": "Library for fast word representations and text classification built on subword n-grams, from FAIR in 2016. It shipped pre-trained word vectors for 157 languages and a famously tiny language-identification model, making industrial-strength NLP possible on a laptop CPU. The repo was archived in March 2024 — mission accomplished.",
   "facts": [
    "About 26k GitHub stars.",
    "Its lid.176 language-identification model recognizes 176 languages and compresses to under 1 MB — still deployed all over the industry years after the repo froze."
   ],
   "latest": "v0.9.2 final line; repository archived March 19, 2024",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/1607.04606"
    },
    {
     "label": "GitHub · fastText",
     "url": "https://github.com/facebookresearch/fastText"
    },
    {
     "label": "fasttext.cc",
     "url": "https://fasttext.cc"
    }
   ],
   "why": "Made industrial-strength text classification and word embeddings run on a laptop CPU using subword n-grams, and shipped pretrained vectors for 157 languages plus a famously tiny language-identification model that still runs inside data pipelines everywhere. Archived in March 2024 after its ideas became standard practice.",
   "try": [
    {
     "label": "pip install fasttext",
     "url": "https://pypi.org/project/fasttext"
    },
    {
     "label": "Pretrained vectors for 157 languages",
     "url": "https://fasttext.cc/docs/en/crawl-vectors.html"
    },
    {
     "label": "Language-ID model on Hugging Face",
     "url": "https://huggingface.co/facebook/fasttext-language-identification"
    }
   ],
   "params": ""
  },
  {
   "id": "prophet",
   "name": "Prophet",
   "full_name": "Prophet",
   "tag": "The forecasting tool every analyst has met",
   "year": 2017,
   "year_label": "2017",
   "cat": "oss",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "Time-series forecasting library from Meta's Core Data Science team (2017), fitting additive models with trend, seasonality and holiday effects so analysts can produce credible forecasts without stats PhDs. Available in Python and R, it remains one of the most widely used forecasting tools in industry and is still actively released.",
   "facts": [
    "Roughly 19k GitHub stars and on the order of a million PyPI downloads a month nearly a decade after release; the 'Forecasting at Scale' paper is a citation staple in business analytics."
   ],
   "latest": "v1.4.0 (August 2026) — dropped Python <3.10, added full static typing and uncertainty propagation for extra regressors",
   "links": [
    {
     "label": "GitHub · prophet",
     "url": "https://github.com/facebook/prophet"
    },
    {
     "label": "facebook.github.io",
     "url": "https://facebook.github.io/prophet"
    },
    {
     "label": "peerj.com",
     "url": "https://peerj.com/preprints/3190"
    }
   ],
   "why": "Let analysts produce credible time-series forecasts with trend, seasonality and holiday effects without a statistics PhD. Released in Python and R in 2017, it remains one of industry's most widely used forecasting tools and is still actively maintained, with v1.4.0 in August 2026 adding full static typing.",
   "try": [
    {
     "label": "Quick start (Python and R)",
     "url": "https://facebook.github.io/prophet/docs/quick_start.html"
    },
    {
     "label": "prophet on PyPI",
     "url": "https://pypi.org/project/prophet"
    },
    {
     "label": "prophet on CRAN",
     "url": "https://cran.r-project.org/package=prophet"
    }
   ],
   "params": ""
  },
  {
   "id": "hydra",
   "name": "Hydra",
   "full_name": "Hydra",
   "tag": "Compose YAML configs, override from the shell, sweep in one flag",
   "year": 2019,
   "year_label": "2019",
   "cat": "oss",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "Framework for elegantly configuring complex applications: hierarchical YAML config composition, command-line overrides, and multirun sweeps with pluggable optimizers (Optuna, Nevergrad, Ax). Released by FAIR in 2019, Hydra became the de facto standard for ML experiment configuration and anchors popular templates like Lightning-Hydra.",
   "facts": [
    "About 9k GitHub stars; the hydra/sweeper=nevergrad one-liner turns any script into a hyperparameter search, and 'Lightning + Hydra' is arguably the most cloned research-project template on GitHub."
   ],
   "latest": "v1.3.6 (August 2024)",
   "links": [
    {
     "label": "GitHub · hydra",
     "url": "https://github.com/facebookresearch/hydra"
    },
    {
     "label": "hydra.cc",
     "url": "https://hydra.cc"
    }
   ],
   "why": "Made hierarchical YAML config composition, command-line overrides and multirun sweeps the standard way to organize ML experiments. Hydra anchors popular templates such as Lightning-Hydra and appears in a large share of research codebases, so reading a Meta or academic training repo usually means reading a Hydra config first.",
   "try": [
    {
     "label": "Tutorial and docs",
     "url": "https://hydra.cc/docs/intro"
    },
    {
     "label": "pip install hydra-core",
     "url": "https://pypi.org/project/hydra-core"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/hydra"
    }
   ],
   "params": ""
  },
  {
   "id": "nevergrad",
   "name": "Nevergrad",
   "full_name": "Nevergrad",
   "tag": "Dozens of black-box optimizers behind one ask/tell API",
   "year": 2018,
   "year_label": "2018",
   "cat": "oss",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "Gradient-free (derivative-free) optimization platform from FAIR (2018), bundling dozens of evolutionary and bandit algorithms — CMA-ES, differential evolution, particle swarm and more — behind one ask/tell API, with benchmarks for comparing them. Widely used for hyperparameter tuning and black-box problems where gradients don't exist.",
   "facts": [
    "~4k GitHub stars; plugs directly into Hydra as a sweeper, so many researchers use Nevergrad daily without realizing it."
   ],
   "latest": "v1.0.12 (April 2024)",
   "links": [
    {
     "label": "GitHub · nevergrad",
     "url": "https://github.com/facebookresearch/nevergrad"
    },
    {
     "label": "facebookresearch.github.io",
     "url": "https://facebookresearch.github.io/nevergrad"
    }
   ],
   "why": "Puts dozens of derivative-free optimizers, including CMA-ES, differential evolution, particle swarm and bandit methods, behind one ask/tell API with benchmarks for comparing them. It is a standard choice for hyperparameter tuning and black-box problems where gradients don't exist, and it plugs into Hydra's multirun sweeper.",
   "try": [
    {
     "label": "Documentation",
     "url": "https://facebookresearch.github.io/nevergrad"
    },
    {
     "label": "pip install nevergrad",
     "url": "https://pypi.org/project/nevergrad"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/nevergrad"
    }
   ],
   "params": ""
  },
  {
   "id": "parlai",
   "name": "ParlAI",
   "full_name": "ParlAI",
   "tag": "The dialogue framework that raised BlenderBot, archived in the LLM era",
   "year": 2017,
   "year_label": "2017",
   "cat": "oss",
   "size": 1,
   "status": "archived",
   "open": true,
   "lineage": [],
   "desc": "FAIR's unified framework for training and evaluating dialogue models across 100+ datasets, from chit-chat to VQA. It was the birthplace of the BlenderBot line and years of conversational-AI research. After the LLM era made bespoke dialogue frameworks obsolete, the repository was archived on July 30, 2026.",
   "facts": [
    "10.6k GitHub stars and 4,359 commits at archive time; nearly a decade of dialogue research — including BlenderBot 1/2/3 — ran through ParlAI tasks and teachers."
   ],
   "latest": "v1.7.x final line; repository archived July 30, 2026",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/1705.06476"
    },
    {
     "label": "GitHub · ParlAI",
     "url": "https://github.com/facebookresearch/ParlAI"
    },
    {
     "label": "parl.ai",
     "url": "https://parl.ai"
    }
   ],
   "why": "For nearly a decade the unified framework for training and evaluating dialogue models across 100+ datasets, and the birthplace of the BlenderBot line. General-purpose LLMs made bespoke dialogue frameworks obsolete and the repository was archived on July 30, 2026, but its tasks remain a reference for conversational AI research.",
   "try": [
    {
     "label": "Documentation",
     "url": "https://parl.ai/docs"
    },
    {
     "label": "parlai on PyPI",
     "url": "https://pypi.org/project/parlai"
    },
    {
     "label": "Archived code on GitHub",
     "url": "https://github.com/facebookresearch/ParlAI"
    }
   ],
   "params": ""
  },
  {
   "id": "animated-drawings",
   "name": "Animated Drawings",
   "full_name": "Animated Drawings",
   "tag": "Your kid's doodle, walking",
   "year": 2021,
   "year_label": "2021",
   "cat": "oss",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR's system that automatically rigs and animates children's drawings of human figures. The public demo (late 2021) went viral; in April 2023 Meta open-sourced the code under MIT plus the Amateur Drawings dataset of 178,000+ annotated sketches — collected, with consent, from demo uploads. Published in ACM Transactions on Graphics.",
   "facts": [
    "About 13k GitHub stars.",
    "Some 3.2 million people uploaded ~6.7 million images to the demo; the team wanted 10,000 drawings and got over 3 million, distilling 178,166 into a first-of-its-kind dataset."
   ],
   "latest": "Open-source release April 2023; demo live at sketch.metademolab.com",
   "links": [
    {
     "label": "GitHub · AnimatedDrawings",
     "url": "https://github.com/facebookresearch/AnimatedDrawings"
    },
    {
     "label": "Live demo",
     "url": "https://sketch.metademolab.com"
    },
    {
     "label": "dl.acm.org",
     "url": "https://dl.acm.org/doi/10.1145/3592788"
    }
   ],
   "why": "A rare FAIR project that went viral with the public: upload a child's drawing and it rigs and animates the figure. Consented demo uploads became the 178,000-sketch Amateur Drawings dataset, and the MIT-licensed code, published in ACM Transactions on Graphics, shows how a research demo can double as a data engine.",
   "try": [
    {
     "label": "Animate a drawing (official demo)",
     "url": "https://sketch.metademolab.com"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/AnimatedDrawings"
    },
    {
     "label": "Amateur Drawings annotations (JSON download)",
     "url": "https://dl.fbaipublicfiles.com/amateur_drawings/amateur_drawings_annotations.json"
    }
   ],
   "params": ""
  },
  {
   "id": "cicero",
   "name": "CICERO",
   "full_name": "CICERO",
   "tag": "It won Diplomacy by out-talking humans — and no one called it out",
   "year": 2022,
   "year_label": "2022",
   "cat": "games",
   "size": 3,
   "status": "active",
   "open": true,
   "lineage": [
    "rebel"
   ],
   "desc": "The first AI to achieve human-level play in Diplomacy — a game requiring natural-language negotiation, persuasion and cooperation with people. CICERO combined a dialogue language model with a strategic planning engine that inferred other players' intentions. Published in Science (November 2022); code and models were open-sourced for research.",
   "facts": [
    "An AI that won a game of lying and alliances not by out-calculating humans, but by out-talking them.",
    "Across 40 anonymous online games against 82 humans, CICERO scored more than double the average of its opponents and ranked in the top 10% of repeat players — sending about 130 negotiation messages per game without being unmasked as an AI.",
    "Played 40 webDiplomacy.net blitz games between Aug 19 and Oct 13, 2022, negotiating with humans in free-form English — and players didn't flag it as an AI during league play."
   ],
   "latest": "Science publication + code release (November 2022)",
   "links": [
    {
     "label": "Science",
     "url": "https://www.science.org/doi/10.1126/science.ade9097"
    },
    {
     "label": "GitHub · diplomacy_cicero",
     "url": "https://github.com/facebookresearch/diplomacy_cicero"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/cicero"
    }
   ],
   "why": "The first AI to reach human-level play in Diplomacy, a game won through natural-language negotiation and cooperation. A 2.7B-parameter dialogue model fine-tuned on 40,000+ webDiplomacy games, coupled with a strategic planner, scored more than double the human average and ranked in the top 10% of players. Code and weights are open.",
   "try": [
    {
     "label": "Code and model weights on GitHub",
     "url": "https://github.com/facebookresearch/diplomacy_cicero"
    },
    {
     "label": "Official CICERO research page",
     "url": "https://ai.meta.com/research/cicero"
    },
    {
     "label": "Meta AI blog post",
     "url": "https://ai.meta.com/blog/cicero-ai-negotiates-persuades-and-cooperates-with-people"
    }
   ],
   "params": "2.7B (dialogue language model)"
  },
  {
   "id": "pluribus",
   "name": "Pluribus",
   "full_name": "Pluribus",
   "tag": "Superhuman six-player poker for $150 of compute",
   "year": 2019,
   "year_label": "2019",
   "cat": "games",
   "size": 2,
   "status": "active",
   "open": false,
   "lineage": [],
   "desc": "The first AI to beat elite professionals at six-player no-limit Texas Hold'em — the first superhuman result in any major game with more than two players. Built by Noam Brown (Facebook AI Research) and Tuomas Sandholm (CMU), it won across 10,000 hands against pros including Darren Elias and Chris Ferguson. Published in Science, July 2019; code was not released.",
   "facts": [
    "Famously frugal: Pluribus trained its blueprint strategy in about 8 days on a 64-core server — roughly $150 of cloud compute — while beating 13 pros who had each won over $1M playing poker.",
    "The code stayed closed partly over concerns about online poker integrity."
   ],
   "latest": "Science publication, July 2019 ('Superhuman AI for multiplayer poker')",
   "links": [
    {
     "label": "Science",
     "url": "https://www.science.org/doi/10.1126/science.aay2400"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/pluribus-first-ai-to-beat-pros-in-6-player-poker"
    },
    {
     "label": "cs.cmu.edu",
     "url": "https://www.cs.cmu.edu/news/2019/carnegie-mellon-and-facebook-ai-beats-professionals-six-player-poker"
    }
   ],
   "why": "The first superhuman result in any major game with more than two players: six-player no-limit Texas Hold'em against professionals including Darren Elias and Chris Ferguson over 10,000 hands. Its blueprint strategy took eight days and 12,400 core-hours to compute, and it played live on 28 cores. Code was never released.",
   "try": [
    {
     "label": "CMU announcement with details",
     "url": "https://www.cs.cmu.edu/news/2019/carnegie-mellon-and-facebook-ai-beats-professionals-six-player-poker"
    }
   ],
   "params": ""
  },
  {
   "id": "rebel",
   "name": "ReBeL",
   "full_name": "ReBeL",
   "tag": "AlphaZero for hidden-information games — the poker code stayed locked",
   "year": 2020,
   "year_label": "2020",
   "cat": "games",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "pluribus"
   ],
   "desc": "Recursive Belief-based Learning: a general RL-plus-search algorithm extending AlphaZero-style self-play to imperfect-information games by operating on public belief states. ReBeL reached superhuman heads-up no-limit hold'em with far less poker-specific knowledge than prior bots. Meta open-sourced a Liar's Dice implementation alongside the NeurIPS 2020 paper.",
   "facts": [
    "By converting hidden-information games into continuous-state perfect-information games over beliefs, ReBeL unified the AlphaZero and Pluribus research lines — the poker code itself was withheld to protect online games, so Liar's Dice became the open testbed."
   ],
   "latest": "NeurIPS 2020 paper + open-source Liar's Dice code (December 2020)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2007.13544"
    },
    {
     "label": "GitHub · rebel",
     "url": "https://github.com/facebookresearch/rebel"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/rebel-a-general-game-playing-ai-bot-that-excels-at-poker-and-more"
    }
   ],
   "why": "Extended AlphaZero-style self-play plus search to imperfect-information games by operating on public belief states, reaching superhuman heads-up no-limit hold'em with far less poker-specific knowledge than earlier bots. The Apache-licensed Liar's Dice implementation, with released value-function checkpoints, is the accessible entry point.",
   "try": [
    {
     "label": "Code and checkpoints on GitHub",
     "url": "https://github.com/facebookresearch/rebel"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2007.13544"
    },
    {
     "label": "NeurIPS 2020 proceedings page",
     "url": "https://proceedings.neurips.cc/paper/2020/hash/c61f571dbd2fb949d3fe5ae1608dd48b-Abstract.html"
    }
   ],
   "params": ""
  },
  {
   "id": "elf-opengo",
   "name": "ELF OpenGo",
   "full_name": "ELF OpenGo",
   "tag": "20-0 against Go professionals, on a single GPU — then open-sourced",
   "year": 2018,
   "year_label": "2018",
   "cat": "games",
   "size": 1,
   "status": "archived",
   "open": true,
   "lineage": [],
   "desc": "FAIR's open reimplementation of AlphaZero for Go, built on the ELF reinforcement-learning platform. Trained on 2,000 GPUs over about two weeks, it went 20-0 against four top-30 professional players — while running on a single GPU — and released its code, trained models and analysis so anyone could reproduce a superhuman Go engine.",
   "facts": [
    "Beat pros 20-0 with 50 seconds per move while the humans had unlimited time; the team also released win-rate analyses of 87,000 professional human games — a gift to the Go community."
   ],
   "latest": "Final model + ICML 2019 analysis paper (v2, 2019)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/1902.04522"
    },
    {
     "label": "GitHub · ELF",
     "url": "https://github.com/pytorch/ELF"
    },
    {
     "label": "Meta AI",
     "url": "https://ai.meta.com/tools/elf-opengo"
    }
   ],
   "why": "An open reimplementation of AlphaZero for Go that went 20-0 against four top-30 professionals while running on a single GPU, then released code, models, 20 million self-play games and a playable Windows binary so anyone could reproduce a superhuman engine. The repository was archived in December 2020.",
   "try": [
    {
     "label": "Models, game analysis tool and Windows binary",
     "url": "https://ai.meta.com/tools/elf-opengo"
    },
    {
     "label": "Archived code on GitHub",
     "url": "https://github.com/pytorch/ELF"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/1902.04522"
    }
   ],
   "params": ""
  },
  {
   "id": "purple-llama",
   "name": "Purple Llama",
   "full_name": "Purple Llama (Llama Guard, LlamaFirewall, Prompt Guard, CyberSecEval)",
   "tag": "Open-source safety tooling for open models",
   "year": 2023,
   "year_label": "2023",
   "cat": "safety",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [
    "llama2"
   ],
   "desc": "Meta's open trust-and-safety line, launched December 2023. Llama Guard classifies unsafe prompts/responses; Code Shield and CyberSecEval target insecure code and cyber risk. LlamaCon 2025 added Llama Guard 4 (a single 12B natively multimodal safeguard), LlamaFirewall (guardrail orchestration against prompt injection and risky tool use), and Prompt Guard 2.",
   "facts": [
    "The name is a security pun: 'purple teaming' = red (attack) + blue (defend).",
    "Llama Guard 4 unified separate text and vision guard models into one 12B model.",
    "The PurpleLlama repo has ~4,400 GitHub stars; an AI Defenders Program gives partners early security tooling.",
    "Llama Guard 4 runs on a single GPU and handles up to five images per prompt; the Llama Guard lineage (v1→v4 in 17 months) became the default open moderation layer across the industry."
   ],
   "latest": "Llama Guard 4 12B + LlamaFirewall + Prompt Guard 2 (Apr 2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2312.06674"
    },
    {
     "label": "GitHub · PurpleLlama",
     "url": "https://github.com/meta-llama/PurpleLlama"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-Guard-4-12B"
    },
    {
     "label": "Meta AI blog · ai defenders program llama pro",
     "url": "https://ai.meta.com/blog/ai-defenders-program-llama-protection-tools"
    },
    {
     "label": "Meta AI blog · purple llama open trust safety",
     "url": "https://ai.meta.com/blog/purple-llama-open-trust-safety-generative-ai"
    },
    {
     "label": "llama.com",
     "url": "https://www.llama.com/llama-protections"
    }
   ],
   "why": "Meta's open trust-and-safety toolkit: Llama Guard classifies unsafe prompts and responses, Prompt Guard catches jailbreaks and injections, CyberSecEval measures cyber risk. Llama Guard 4 is a single 12B natively multimodal safeguard pruned from Llama 4 Scout, making open guardrails a standard layer in Llama and third-party deployments.",
   "try": [
    {
     "label": "Llama Guard 4 on Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-Guard-4-12B"
    },
    {
     "label": "Prompt Guard 2 on Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-86M"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/meta-llama/PurpleLlama"
    }
   ],
   "params": "Llama Guard 4: 12B; Prompt Guard 2: 22M / 86M"
  },
  {
   "id": "cyberseceval",
   "name": "CyberSecEval",
   "full_name": "CyberSecEval",
   "tag": "Can your model break code — and can it patch it?",
   "year": 2023,
   "year_label": "2023",
   "cat": "safety",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "purple-llama"
   ],
   "desc": "The most comprehensive open benchmark suite for measuring LLM cybersecurity risks and capabilities, spanning insecure code generation, cyberattack helpfulness, prompt-injection resistance and vulnerability exploitation. CyberSecEval 4 (April 2025) added defensive evaluations: AutoPatchBench (automatic vulnerability patching) and CyberSOCEval, built with CrowdStrike for SOC-style malware and threat-intel analysis.",
   "facts": [
    "Four major versions in under 18 months (Dec 2023 → Apr 2025).",
    "V4's pivot from measuring offense to measuring defense — can your model patch code, not just break it — set the template other labs followed."
   ],
   "latest": "CyberSecEval 4 (April 2025)",
   "links": [
    {
     "label": "GitHub · PurpleLlama",
     "url": "https://github.com/meta-llama/PurpleLlama/tree/main/CybersecurityBenchmarks"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/ai-defenders-program-llama-protection-tools"
    },
    {
     "label": "meta-llama.github.io",
     "url": "https://meta-llama.github.io/PurpleLlama/CyberSecEval"
    }
   ],
   "why": "The most comprehensive open benchmark for LLM cybersecurity risk: insecure code generation, cyberattack helpfulness, prompt-injection resistance and vulnerability exploitation. CyberSecEval 4 added defensive evaluations, AutoPatchBench and CyberSOCEval with CrowdStrike, and it has become a standard reference for cyber-safety evaluation of frontier models.",
   "try": [
    {
     "label": "Benchmark code on GitHub",
     "url": "https://github.com/meta-llama/PurpleLlama/tree/main/CybersecurityBenchmarks"
    },
    {
     "label": "Documentation",
     "url": "https://meta-llama.github.io/PurpleLlama/CyberSecEval"
    },
    {
     "label": "Read the original paper",
     "url": "https://arxiv.org/abs/2312.04724"
    }
   ],
   "params": ""
  },
  {
   "id": "llamafirewall",
   "name": "LlamaFirewall",
   "full_name": "LlamaFirewall",
   "tag": "Prompt-injection success rate: 17.6% without, 1.7% with",
   "year": 2025,
   "year_label": "2025",
   "cat": "safety",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "purple-llama"
   ],
   "desc": "An open-source, system-level guardrail framework for securing AI agents, released April-May 2025 and used in production at Meta. Its three layers — PromptGuard 2 (jailbreak detection), Agent Alignment Checks (a chain-of-thought auditor catching goal hijacking), and CodeShield (static analysis of generated code) — defend against prompt injection and insecure agent behavior.",
   "facts": [
    "In Meta's AgentDojo evaluation, prompt-injection attacks succeeded 17.6% of the time without LlamaFirewall — and 1.7% with it, a >90% reduction.",
    "Any developer who can write a regex can add a custom scanner."
   ],
   "latest": "LlamaFirewall (2025), part of the PurpleLlama repository",
   "links": [
    {
     "label": "GitHub · PurpleLlama",
     "url": "https://github.com/meta-llama/PurpleLlama"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/llamafirewall-an-open-source-guardrail-system-for-building-secure-ai-agents"
    },
    {
     "label": "meta-llama.github.io",
     "url": "https://meta-llama.github.io/PurpleLlama/LlamaFirewall"
    }
   ],
   "why": "A system-level guardrail framework for AI agents, used in production at Meta and released open source: PromptGuard 2 for jailbreak detection, Agent Alignment Checks that audit chain-of-thought for goal hijacking, and CodeShield static analysis of generated code. One of the first open, layered defenses built specifically for agent workflows.",
   "try": [
    {
     "label": "pip install llamafirewall",
     "url": "https://pypi.org/project/llamafirewall"
    },
    {
     "label": "Documentation and tutorials",
     "url": "https://meta-llama.github.io/PurpleLlama/LlamaFirewall"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2505.03574"
    }
   ],
   "params": ""
  },
  {
   "id": "gaia",
   "name": "GAIA Benchmark",
   "full_name": "GAIA Benchmark",
   "tag": "466 questions: humans scored 92%, GPT-4 with plugins scored 15%",
   "year": 2023,
   "year_label": "2023",
   "cat": "safety",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "A benchmark for General AI Assistants, co-created by Meta FAIR (Grégoire Mialon, Yann LeCun, Thomas Scialom) with Hugging Face. Its 466 real-world questions require reasoning, web browsing, multi-modality and tool use — conceptually simple for humans, brutal for AIs — and it became the standard yardstick for agentic systems.",
   "facts": [
    "At release, human respondents scored 92% while GPT-4 with plugins managed 15% — the gap that launched a thousand agent frameworks.",
    "Climbing the GAIA leaderboard became a rite of passage for 2024-2026 agent startups."
   ],
   "latest": "GAIA (November 2023); public leaderboard maintained on Hugging Face",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2311.12983"
    },
    {
     "label": "Hugging Face Space",
     "url": "https://huggingface.co/spaces/gaia-benchmark/leaderboard"
    }
   ],
   "why": "The standard yardstick for agentic AI: 466 real-world questions requiring reasoning, browsing, multi-modality and tool use that are simple for humans and hard for models. Co-created by FAIR and Hugging Face, its public leaderboard became the benchmark agent frameworks report, from early AutoGPT-style systems to today's frontier agents.",
   "try": [
    {
     "label": "Public leaderboard (Hugging Face Space)",
     "url": "https://huggingface.co/spaces/gaia-benchmark/leaderboard"
    },
    {
     "label": "Dataset on Hugging Face",
     "url": "https://huggingface.co/datasets/gaia-benchmark/GAIA"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2311.12983"
    }
   ],
   "params": ""
  },
  {
   "id": "mobilellm",
   "name": "MobileLLM",
   "full_name": "MobileLLM",
   "tag": "Proof that sub-billion models could think",
   "year": 2024,
   "year_label": "2024",
   "cat": "edge",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "FAIR's sub-billion-parameter architecture study proving that deep-and-thin transformers with embedding sharing, grouped-query attention, SwiGLU, and block-wise layer sharing beat wide-shallow designs on-device. Published at ICML 2024. Checkpoints from 125M to 1.5B released with full weights and training code — all non-commercial (CC-BY-NC 4.0 weights, FAIR NC code).",
   "facts": [
    "MobileLLM-125M beat prior 125M SoTA by 2.7% and the 350M beat SoTA by 4.3% on zero-shot commonsense reasoning (46.3% vs GPT-neo-125M's 42.9%; 51.3% vs Pythia-410M's 46.6%).",
    "Training the 125M takes ~3 days on 32 A100s over 1T tokens; the 1.5B takes ~18 days.",
    "GitHub repo has ~1.5k stars.",
    "The HF collection also holds ParetoQ extreme-quantization variants down to 1-bit."
   ],
   "latest": "MobileLLM-1.5B + ParetoQ quantized variants (1/1.58/2/3/4-bit); 125M/350M/600M/1B/1.5B line completed Nov 2024",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2402.14905"
    },
    {
     "label": "GitHub · MobileLLM",
     "url": "https://github.com/facebookresearch/MobileLLM"
    },
    {
     "label": "Hugging Face · MobileLLM 125M",
     "url": "https://huggingface.co/facebook/MobileLLM-125M"
    },
    {
     "label": "Hugging Face · mobilellm",
     "url": "https://huggingface.co/collections/facebook/mobilellm"
    }
   ],
   "why": "Reset how sub-billion language models are designed: deep-and-thin transformers with embedding sharing, grouped-query attention and block-wise layer sharing beat wide-shallow designs at equal size. Published at ICML 2024 with full weights and training code, it became the reference architecture for on-device models, including Meta's R1, Pro and Flash lines.",
   "try": [
    {
     "label": "Model collection on Hugging Face",
     "url": "https://huggingface.co/collections/facebook/mobilellm"
    },
    {
     "label": "Code on GitHub",
     "url": "https://github.com/facebookresearch/MobileLLM"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2402.14905"
    }
   ],
   "params": "125M / 350M / 600M / 1B / 1.5B"
  },
  {
   "id": "mobilellm-r1",
   "name": "MobileLLM-R1 / R1.5",
   "full_name": "MobileLLM-R1 / R1.5",
   "tag": "Sub-billion reasoners matching Qwen3-0.6B on a seventh of the tokens",
   "year": 2025,
   "year_label": "2025",
   "cat": "edge",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "mobilellm"
   ],
   "desc": "Open-recipe sub-billion reasoning models (140M/360M/950M, base + final) for math, code, and science. R1-950M matches or beats Qwen3-0.6B on MATH, MMLU, and LiveCodeBench despite fewer than 5T total training tokens vs Qwen3's 36T. R1.5 (Nov 2025) adds on-policy knowledge distillation. FAIR Noncommercial Research License — must be labeled non-commercial.",
   "facts": [
    "R1.5's on-policy KD added 10-35 points on hard reasoning benchmarks: R1.5-950M scores 39.9 on AIME'24 vs Qwen3-0.6B's 11.3, and the 360M jumped from 28.4 to 63.4 on MATH.",
    "Fully open training recipe: data sources, code, and all checkpoints published."
   ],
   "latest": "MobileLLM-R1.5 140M/360M/950M (Nov 24, 2025); R1 released Sept 12, 2025",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2509.24945"
    },
    {
     "label": "GitHub · MobileLLM-R1",
     "url": "https://github.com/facebookresearch/MobileLLM-R1"
    },
    {
     "label": "Hugging Face · MobileLLM R1 950M",
     "url": "https://huggingface.co/facebook/MobileLLM-R1-950M"
    },
    {
     "label": "Hugging Face · MobileLLM R1.5 950M",
     "url": "https://huggingface.co/facebook/MobileLLM-R1.5-950M"
    },
    {
     "label": "Hugging Face · mobilellm r1 68c4597b104fac45f",
     "url": "https://huggingface.co/collections/facebook/mobilellm-r1-68c4597b104fac45f28f448e"
    }
   ],
   "why": "Showed reasoning ability need not wait for billion-parameter scale: the 950M model matches or beats Qwen3-0.6B on MATH, MMLU and LiveCodeBench despite fewer than 5T training tokens versus Qwen3's 36T. R1.5 in November 2025 added on-policy knowledge distillation, and the fully open recipe lets others reproduce it.",
   "try": [
    {
     "label": "MobileLLM-R1.5-950M on Hugging Face",
     "url": "https://huggingface.co/facebook/MobileLLM-R1.5-950M"
    },
    {
     "label": "R1 / R1.5 collection on Hugging Face",
     "url": "https://huggingface.co/collections/facebook/mobilellm-r1-68c4597b104fac45f28f448e"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2509.24945"
    }
   ],
   "params": "140M / 360M / 950M"
  },
  {
   "id": "mobilellm-pro",
   "name": "MobileLLM-Pro",
   "full_name": "MobileLLM-Pro",
   "tag": "1B parameters, 128k context — built by the smart-glasses org",
   "year": 2025,
   "year_label": "2025",
   "cat": "edge",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "mobilellm"
   ],
   "desc": "Meta Reality Labs' 1.08B-parameter on-device foundation model (base + instruct) with 128k context, released October 2025. Interleaves local and global attention at a 3:1 ratio (512-token local windows), cutting prefill latency 1.8x and shrinking KV cache from 117MB to 40MB at 8k context. Ships int4 variants for CPU, Apple Neural Engine, and Qualcomm HTP. FAIR Noncommercial Research License.",
   "facts": [
    "Beats Gemma 3 1B on MMLU (44.8% vs 29.9%) and crushes both Gemma 3 1B and Llama 3.2 1B on HumanEval coding (59.8% vs 41.5% and 37.8%).",
    "Notably developed by Reality Labs — the smart-glasses org — not FAIR, hinting at its product destination."
   ],
   "latest": "MobileLLM-Pro 1B base/instruct + int4-cpu and int4-accelerator variants (Oct 2025); technical report arXiv:2511.06719 (Nov 2025)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2511.06719"
    },
    {
     "label": "Hugging Face · MobileLLM Pro",
     "url": "https://huggingface.co/facebook/MobileLLM-Pro"
    },
    {
     "label": "Hugging Face · MobileLLM Pro base int4 cpu",
     "url": "https://huggingface.co/facebook/MobileLLM-Pro-base-int4-cpu"
    },
    {
     "label": "Hugging Face · MobileLLM Pro base int4 accele",
     "url": "https://huggingface.co/facebook/MobileLLM-Pro-base-int4-accelerator"
    }
   ],
   "why": "Reality Labs' production-oriented on-device foundation model: 128k context with interleaved local and global attention (3:1) that cuts prefill latency 1.8x and shrinks the KV cache from 117MB to 40MB at 8k context, shipped with int4 variants for CPU, Apple Neural Engine and Qualcomm HTP. A concrete blueprint for LLMs on glasses and phones.",
   "try": [
    {
     "label": "Model on Hugging Face",
     "url": "https://huggingface.co/facebook/MobileLLM-Pro"
    },
    {
     "label": "Chat demo (Hugging Face Space)",
     "url": "https://huggingface.co/spaces/akhaliq/MobileLLM-Pro"
    },
    {
     "label": "Read the technical report",
     "url": "https://arxiv.org/abs/2511.06719"
    }
   ],
   "params": "1.08B"
  },
  {
   "id": "mobilellm-flash",
   "name": "MobileLLM-Flash",
   "full_name": "MobileLLM-Flash",
   "tag": "Architecture search with real phone latency in the loop",
   "year": 2026,
   "year_label": "2026",
   "cat": "edge",
   "size": 1,
   "status": "active",
   "open": false,
   "lineage": [
    "mobilellm"
   ],
   "desc": "Latest MobileLLM generation (350M/650M/1.4B) designed via hardware-in-the-loop architecture search under real mobile latency constraints, with attention skipping for long-context acceleration up to 8k. Paper (Mar 2026) accepted to ACL 2026 Industry Track. Weights had not been publicly released as of Sept 2026.",
   "facts": [
    "Up to 1.8x faster prefill and 1.6x faster decode on mobile CPUs vs comparable-quality models.",
    "Represents a methodology shift for the family: instead of hand-designed architecture rules (MobileLLM's deep-and-thin), the architecture itself is searched with actual phone latency in the loop.",
    "Only community trackers list a weights release; Meta's own HF collection did not include it."
   ],
   "latest": "MobileLLM-Flash 350M/650M/1.4B (paper Mar 16, 2026; ACL Industry Track 2026)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2603.15954"
    },
    {
     "label": "Community tracker",
     "url": "https://github.com/stevelaskaridis/awesome-mobile-llm"
    }
   ],
   "why": "Designs the model around the phone rather than the benchmark: hardware-in-the-loop architecture search under real mobile latency constraints, plus attention skipping to accelerate contexts up to 8k. Accepted to the ACL 2026 Industry Track; as of September 2026 the weights had not been released, so only the paper is available.",
   "try": [
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2603.15954"
    }
   ],
   "params": "350M / 650M / 1.4B"
  },
  {
   "id": "quantized-llama",
   "name": "Quantized Llama 3.2",
   "full_name": "Quantized Llama 3.2 (1B/3B) + ExecuTorch",
   "tag": "Llama 3.2 on a phone — 56% smaller, up to 4x faster",
   "year": 2024,
   "year_label": "2024",
   "cat": "edge",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "llama3",
    "executorch"
   ],
   "desc": "Meta's first official quantized Llama releases (Oct 24, 2024): Llama 3.2 1B/3B Instruct in two flavors — Quantization-Aware Training with LoRA adaptors for accuracy, and SpinQuant post-training quantization for portability. 2-4x faster with 56% smaller size and 41% less memory, running on phones via PyTorch's ExecuTorch on Qualcomm, MediaTek, and Arm. Llama 3.2 Community License — commercial use permitted.",
   "facts": [
    "Measured on a OnePlus 12: 2.5x faster decode, 4.2x faster prefill, 56% average size reduction, 41% less memory vs BF16 — verified also on Samsung S24+/S22.",
    "QLoRA keeps accuracy within 1.95% of BF16 on the 3B.",
    "SpinQuant (learned rotation matrices + GPTQ, INT4 groupwise g32 weights with 8-bit dynamic activations) is itself a FAIR research contribution with its own open GitHub repo.",
    "This is the family's commercially-licensed on-device option."
   ],
   "latest": "Llama 3.2 1B/3B Instruct QLoRA_INT4_EO8 and SpinQuant_INT4_EO8 (Oct 24, 2024)",
   "links": [
    {
     "label": "GitHub · executorch",
     "url": "https://github.com/pytorch/executorch"
    },
    {
     "label": "GitHub · SpinQuant",
     "url": "https://github.com/facebookresearch/SpinQuant"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct-SpinQuant_INT4_EO8"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/meta-llama-quantized-lightweight-models"
    },
    {
     "label": "PyTorch blog",
     "url": "https://pytorch.org/blog/unleashing-ai-mobile"
    }
   ],
   "why": "Meta's first official quantized Llama releases made small-model on-device deployment a supported path rather than a community hack: QAT with LoRA adaptors for accuracy and SpinQuant for portability, 2-4x faster with 56% smaller size and 41% less memory, running via ExecuTorch on Qualcomm, MediaTek and Arm under the commercial Llama 3.2 license.",
   "try": [
    {
     "label": "Llama 3.2 1B SpinQuant on Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct-SpinQuant_INT4_EO8"
    },
    {
     "label": "Llama 3.2 3B QLoRA on Hugging Face",
     "url": "https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct-QLORA_INT4_EO8"
    },
    {
     "label": "Run Llama on-device (ExecuTorch example)",
     "url": "https://github.com/pytorch/executorch/tree/main/examples/models/llama"
    }
   ],
   "params": "1B / 3B"
  },
  {
   "id": "llm-compiler",
   "name": "Meta LLM Compiler",
   "full_name": "Meta LLM Compiler",
   "tag": "77% of a full autotuning search's gains, without compiling once",
   "year": 2024,
   "year_label": "2024",
   "cat": "edge",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "code-llama"
   ],
   "desc": "Foundation models (7B/13B, plus fine-tuned FTD variants) built on Code Llama for compiler optimization: trained on 546B tokens of LLVM-IR and x86/ARM/CUDA assembly to emulate the compiler, tune optimization flags, and disassemble binaries back to IR. Released June 27, 2024 under the bespoke Meta LLM Compiler License permitting both research and commercial use.",
   "facts": [
    "Achieves 77% of the code-size-optimizing potential of a full autotuning search — without running a single compilation.",
    "Disassembly: 45% round-trip success, 14% exact match, 0.96 round-trip BLEU converting x86_64/ARM assembly back to LLVM-IR.",
    "FTD-13B beats the compiler's own -Oz flag by 4.88% on code size.",
    "16k-token context window; trained on over 500B tokens of compiler IR and assembly."
   ],
   "latest": "LLM Compiler 7B/13B + 7B-ftd/13B-ftd (Jun 27, 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2407.02524"
    },
    {
     "label": "Hugging Face",
     "url": "https://huggingface.co/facebook/llm-compiler-13b"
    },
    {
     "label": "Meta AI research",
     "url": "https://ai.meta.com/research/publications/meta-large-language-model-compiler-foundation-models-of-compiler-optimization"
    }
   ],
   "why": "The first foundation models for compiler optimization: trained on 546B tokens of LLVM-IR and assembly, the 13B model emulates compiler optimizations 20% of the time versus Code Llama's 0.8%, and the FTD variants improve code size by 4.88% and disassemble with 0.96 round-trip BLEU. Released under a license permitting commercial use.",
   "try": [
    {
     "label": "LLM Compiler 13B on Hugging Face",
     "url": "https://huggingface.co/facebook/llm-compiler-13b"
    },
    {
     "label": "LLM Compiler 7B-FTD on Hugging Face",
     "url": "https://huggingface.co/facebook/llm-compiler-7b-ftd"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2407.02524"
    }
   ],
   "params": "7B / 13B (plus FTD variants)"
  },
  {
   "id": "mtia",
   "name": "MTIA",
   "full_name": "MTIA (Meta Training & Inference Accelerator)",
   "tag": "Meta's own AI silicon",
   "year": 2023,
   "year_label": "2023",
   "cat": "edge",
   "size": 2,
   "status": "closed",
   "open": false,
   "lineage": [],
   "desc": "Meta's custom AI silicon line, co-developed with Broadcom: MTIA v1 (2023) and v2 (April 2024) served ranking and recommendations; in March 2026 Meta unveiled four inference chips — MTIA 300 (already live in data centers), 400, 450, and 500 — on a six-month release cadence, targeting generative inference like image and video creation, and reducing reliance on Nvidia.",
   "facts": [
    "MTIA 300 was serving production traffic before the March 2026 announcement.",
    "A deeper Broadcom co-development partnership was formalized in April 2026.",
    "The chips deliberately avoid frontier-LLM training — that stays on GPUs."
   ],
   "latest": "MTIA 300 deployed; 400/450/500 lineup announced (Mar 11, 2026)",
   "links": [
    {
     "label": "Meta AI blog · meta mtia scale ai chips for b",
     "url": "https://ai.meta.com/blog/meta-mtia-scale-ai-chips-for-billions"
    },
    {
     "label": "Meta AI blog · next generation meta training ",
     "url": "https://ai.meta.com/blog/next-generation-meta-training-inference-accelerator-AI-MTIA"
    },
    {
     "label": "Meta Newsroom · expanding metas custom silicon",
     "url": "https://about.fb.com/news/2026/03/expanding-metas-custom-silicon-to-power-our-ai-workloads"
    },
    {
     "label": "Meta Newsroom · meta partners with broadcom to",
     "url": "https://about.fb.com/news/2026/04/meta-partners-with-broadcom-to-co-develop-custom-ai-silicon"
    }
   ],
   "why": "Meta's path off Nvidia dependence: MTIA v1 and v2 (TSMC 5nm, 708 INT8 TOPS, 90W) served ranking and recommendation models; MTIA 300 is now in production and the 400, 450 and 500 chips target generative inference on a six-month cadence, versus the industry's one-to-two-year cycle, co-developed with Broadcom.",
   "try": [
    {
     "label": "March 2026 silicon roadmap (Meta Newsroom)",
     "url": "https://about.fb.com/news/2026/03/expanding-metas-custom-silicon-to-power-our-ai-workloads"
    },
    {
     "label": "MTIA v2 deep dive (Meta AI blog)",
     "url": "https://ai.meta.com/blog/next-generation-meta-training-inference-accelerator-AI-MTIA"
    }
   ],
   "params": ""
  },
  {
   "id": "superclusters",
   "name": "GPU superclusters: 350K H100s → Prometheus & Hyperion",
   "full_name": "GPU superclusters: 350K H100s → Prometheus & Hyperion",
   "tag": "Gigawatt-scale factories for intelligence",
   "year": 2024,
   "year_label": "2024",
   "cat": "edge",
   "size": 2,
   "status": "closed",
   "open": false,
   "lineage": [],
   "desc": "The compute behind everything: Zuckerberg pledged 350,000 H100s (~600K H100-equivalents) by end of 2024; Llama 3 trained on twin 24,576-GPU clusters. The next act is gigawatt-scale 'titan clusters' — Prometheus (Ohio, coming online 2026) and Hyperion (Louisiana), expanded in July 2026 from 2GW to 5GW at a cost topping $50 billion. 2026 AI capex: $115-135B.",
   "facts": [
    "Zuckerberg said Hyperion's footprint would cover a significant part of Manhattan.",
    "Hyperion's price tag nearly doubled from $27B to $50B+ in one 2026 announcement."
   ],
   "latest": "Prometheus online 2026 (~1GW); Hyperion expanded to 5GW, $50B+ (Jul 2026)",
   "links": [
    {
     "label": "engineering.fb.com",
     "url": "https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure"
    },
    {
     "label": "qz.com",
     "url": "https://qz.com/meta-louisiana-hyperion-data-center-expansion-5-gigawatts-071326"
    },
    {
     "label": "datacenterfrontier.com",
     "url": "https://www.datacenterfrontier.com/hyperscale/article/55310441/ownership-and-power-challenges-in-metas-hyperion-and-prometheus-data-centers"
    }
   ],
   "why": "The compute that made Llama and Muse possible: two 24,576-GPU H100 clusters trained Llama 3, on the way to 350,000 H100s (about 600,000 H100-equivalents) by end of 2024. The next step is gigawatt-scale Prometheus and Hyperion, the latter expanded to 5GW in July 2026, backed by $115-135B of 2026 AI capex.",
   "try": [
    {
     "label": "Meta Engineering: building the GenAI clusters",
     "url": "https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure"
    },
    {
     "label": "Prometheus and Hyperion analysis (Data Center Frontier)",
     "url": "https://www.datacenterfrontier.com/hyperscale/article/55310441/ownership-and-power-challenges-in-metas-hyperion-and-prometheus-data-centers"
    }
   ],
   "params": ""
  },
  {
   "id": "codec-avatars",
   "name": "Codec Avatars",
   "full_name": "Codec Avatars",
   "tag": "Telepresence that looks exactly like you",
   "year": 2019,
   "year_label": "2019",
   "cat": "reality",
   "size": 2,
   "status": "active",
   "open": false,
   "lineage": [],
   "desc": "Meta's decade-long program to make VR telepresence indistinguishable from reality: neural-network 'codec' avatars captured in a multi-hundred-camera dome and driven live by headset sensors. Publicly revealed March 2019 from the Pittsburgh lab (founded 2015, led by Yaser Sheikh), it remains Reality Labs Research's flagship moonshot, now testing internally at Meta.",
   "facts": [
    "Original 2019 capture rig used 171 cameras.",
    "Codec Avatars 2.0 (shown at MIT, 2022) rendered 5 avatars at 50fps on a Quest 2 and was claimed to cross the uncanny valley; Sheikh said the project went from 'ten miracles away' to 'five'.",
    "The Sept 28, 2023 Lex Fridman–Zuckerberg podcast — the first photoreal-avatar interview — needed a workstation with 4x RTX 4090s per side.",
    "The CVPR 2026 paper scales pretraining to 1M videos with zero-shot robustness even to stylized imagery.",
    "Program is closed, but Meta open-sourced companion datasets (Ava-256, Goliath) and code (audio2photoreal)."
   ],
   "latest": "Large-scale Codec Avatars, CVPR 2026 (pretrained on 1M in-the-wild videos); Gaussian/relightable codec avatar papers through Dec 2025",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2604.02320"
    },
    {
     "label": "meta.com",
     "url": "https://www.meta.com/emerging-tech/codec-avatars"
    },
    {
     "label": "tech.fb.com",
     "url": "https://tech.fb.com/ar-vr/2019/03/codec-avatars-facebook-reality-labs"
    },
    {
     "label": "uploadvr.com",
     "url": "https://www.uploadvr.com/mark-zuckerberg-lex-fridman-interview-photorealistic-codec-avatars"
    }
   ],
   "why": "Reality Labs' flagship moonshot: neural avatars captured in a multi-camera dome and driven live by headset sensors, aiming at telepresence indistinguishable from reality. The 2026 Large-scale Codec Avatars work pretrains on 1 million in-the-wild videos, showing the approach generalizes beyond lab captures, and it anchors Meta's long-term bet on presence.",
   "try": [
    {
     "label": "Official Codec Avatars page",
     "url": "https://www.meta.com/emerging-tech/codec-avatars"
    },
    {
     "label": "Read the CVPR 2026 paper",
     "url": "https://arxiv.org/abs/2604.02320"
    }
   ],
   "params": ""
  },
  {
   "id": "codec-avatar-studio",
   "name": "Codec Avatar Studio",
   "full_name": "Codec Avatar Studio (Ava-256 + Goliath datasets)",
   "tag": "Paired dome scans and headset footage of 256 people, released open",
   "year": 2024,
   "year_label": "2024",
   "cat": "reality",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "codec-avatars"
   ],
   "desc": "Reality Labs' open release of the paired captures behind Codec Avatars, published at NeurIPS 2024 (Datasets & Benchmarks). Ava-256 pairs high-resolution dome scans of 256 subjects with headset-camera footage for universal avatar encoding/decoding; Goliath-4 captures 4 subjects across 8 modalities including relightable heads, hands, and full-body scans. CC-BY-NC 4.0.",
   "facts": [
    "Ava-256: 256 subjects each captured twice — in a multi-dozen-view dome AND wearing a headset with infrared cameras — the paired data 'rarely available outside dedicated industrial labs.' Goliath ships official PyTorch implementations of RelightableHands,…"
   ],
   "latest": "NeurIPS 2024 release (code + data live on GitHub)",
   "links": [
    {
     "label": "OpenReview",
     "url": "https://openreview.net/forum?id=a6DteCxiw6"
    },
    {
     "label": "GitHub · ava-256",
     "url": "https://github.com/facebookresearch/ava-256"
    },
    {
     "label": "GitHub · goliath",
     "url": "https://github.com/facebookresearch/goliath"
    },
    {
     "label": "meta.com",
     "url": "https://www.meta.com/emerging-tech/codec-avatars/ava256"
    }
   ],
   "why": "Opens the data behind Codec Avatars to outside researchers: Ava-256 pairs 80-camera dome scans of 256 subjects with Quest Pro headset footage for universal face encoders and decoders, and Goliath-4 captures four subjects across eight modalities. Published at NeurIPS 2024 (Datasets and Benchmarks) under CC-BY-NC 4.0 with download scripts and training code.",
   "try": [
    {
     "label": "Ava-256 code and download script",
     "url": "https://github.com/facebookresearch/ava-256"
    },
    {
     "label": "Goliath code and data",
     "url": "https://github.com/facebookresearch/goliath"
    },
    {
     "label": "Read the NeurIPS 2024 paper",
     "url": "https://openreview.net/forum?id=a6DteCxiw6"
    }
   ],
   "params": ""
  },
  {
   "id": "audio2photoreal",
   "name": "audio2photoreal",
   "full_name": "audio2photoreal",
   "tag": "Record your voice, watch a photoreal avatar gesture along",
   "year": 2024,
   "year_label": "2024",
   "cat": "reality",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "codec-avatars"
   ],
   "desc": "Open code and dataset for synthesizing full-bodied photorealistic Codec Avatars — face, body, and hands — that gesture naturally from conversational audio alone. Combines vector-quantization sample diversity with diffusion for high-frequency motion detail. Released January 2024 by Meta Reality Labs researchers with UC Berkeley; published at CVPR 2024.",
   "facts": [
    "Ships a Gradio demo where you record your own voice and watch a photoreal avatar gesture along.",
    "Built on dyadic (two-person) conversation data so avatars react to interpersonal dynamics, not just their own speech."
   ],
   "latest": "CVPR 2024 (arXiv Jan 2024); code, dataset and Gradio demo live",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2401.01885"
    },
    {
     "label": "GitHub · audio2photoreal",
     "url": "https://github.com/facebookresearch/audio2photoreal"
    }
   ],
   "why": "Synthesizes full-body photorealistic avatars, including face and hands, that gesture naturally from conversational audio alone, combining vector-quantized sample diversity with diffusion for high-frequency motion. Published at CVPR 2024 with code, a four-participant conversational dataset and a Colab demo under CC-BY-NC 4.0, it is a public window into Codec Avatar animation.",
   "try": [
    {
     "label": "Run the demo in Colab",
     "url": "https://colab.research.google.com/drive/1A6EwKM3PeX7dcKV66zxQWuP-v_dKlX_0"
    },
    {
     "label": "Code and dataset on GitHub",
     "url": "https://github.com/facebookresearch/audio2photoreal"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2401.01885"
    }
   ],
   "params": ""
  },
  {
   "id": "semg",
   "name": "Generic Neuromotor Interface",
   "full_name": "Generic Neuromotor Interface (sEMG Nature paper)",
   "tag": "A wristband that reads intention from your nerves",
   "year": 2025,
   "year_label": "2025",
   "cat": "reality",
   "size": 2,
   "status": "active",
   "open": true,
   "lineage": [],
   "desc": "The landmark Nature paper (July 23, 2025; Nature 645, 702–711) from Reality Labs — the culmination of the 2019 CTRL-labs acquisition. A dry-electrode sEMG wristband plus deep networks trained on thousands of participants decodes gestures, wrist movement, and handwriting out-of-the-box for new users, with data and code released openly (CC-BY-NC 4.0).",
   "facts": [
    "Handwriting decoded at 20.9 words/minute; 0.88 gesture detections/sec; wrist-angle velocity error under 13°/sec; >90% offline gesture accuracy on completely unseen users.",
    "Training corpora spanned up to 6,627 participants (handwriting); the open release covers 300 participants (100 per task, ~280 hours) — the largest public sEMG collection.",
    "The 2kHz wristband streams over Bluetooth at 2.46 μVrms noise.",
    "Personalizing with 20 minutes of a user's data cut handwriting errors 16%."
   ],
   "latest": "Nature 645, 702–711 (published 23 Jul 2025); open corpus + training code on GitHub",
   "links": [
    {
     "label": "Nature",
     "url": "https://www.nature.com/articles/s41586-025-09255-w"
    },
    {
     "label": "GitHub · generic-neuromotor-interface",
     "url": "https://github.com/facebookresearch/generic-neuromotor-interface"
    },
    {
     "label": "meta.com",
     "url": "https://www.meta.com/blog/reality-labs-surface-emg-research-nature-publication-ar-glasses-orion"
    }
   ],
   "why": "The Nature paper behind Meta's Neural Band: a dry-electrode wristband and deep networks trained across thousands of participants decode gestures, wrist movement and handwriting for new users with no calibration. It is the culmination of the 2019 CTRL-labs acquisition, and the open corpus (100 participants per task) plus training code make neuromotor interfaces reproducible.",
   "try": [
    {
     "label": "Read the Nature paper",
     "url": "https://www.nature.com/articles/s41586-025-09255-w"
    },
    {
     "label": "Data, models and code on GitHub",
     "url": "https://github.com/facebookresearch/generic-neuromotor-interface"
    }
   ],
   "params": ""
  },
  {
   "id": "emg2qwerty",
   "name": "emg2qwerty",
   "full_name": "emg2qwerty",
   "tag": "Every keystroke of 108 typists, paired with their wrist signals",
   "year": 2024,
   "year_label": "2024",
   "cat": "reality",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "semg"
   ],
   "desc": "The largest public surface-EMG dataset for touch typing: 346 hours across 1,135 sessions from 108 users wearing wrist sEMG bands while typing on QWERTY keyboards, with ground-truth keystrokes and reproducible baselines. Released by Reality Labs at NeurIPS 2024 (Datasets & Benchmarks track) to spark external neuromotor-interface research.",
   "facts": [
    "346 hours / 108 users / 1,135 sessions — the largest sEMG typing corpus ever released; a direct precursor to the typing and handwriting decoders in the 2025 Nature paper and the shipped Neural Band."
   ],
   "latest": "NeurIPS 2024 release (arXiv Oct 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2410.20081"
    },
    {
     "label": "GitHub · emg2qwerty",
     "url": "https://github.com/facebookresearch/emg2qwerty"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/open-sourcing-surface-electromyography-datasets-neurips-2024"
    }
   ],
   "why": "The largest public surface-EMG typing dataset: 346 hours across 1,135 sessions from 108 users, with ground-truth keystrokes and reproducible baselines. Released at NeurIPS 2024 under CC-BY-NC 4.0 so outside labs can work on wrist-based text entry without building a capture rig, the same problem Meta's Neural Band tackles in products.",
   "try": [
    {
     "label": "Dataset and baselines on GitHub",
     "url": "https://github.com/facebookresearch/emg2qwerty"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2410.20081"
    }
   ],
   "params": ""
  },
  {
   "id": "emg2pose",
   "name": "emg2pose",
   "full_name": "emg2pose",
   "tag": "Hand tracking from the wrist — no camera required",
   "year": 2024,
   "year_label": "2024",
   "cat": "reality",
   "size": 1,
   "status": "active",
   "open": true,
   "lineage": [
    "semg"
   ],
   "desc": "A benchmark for reconstructing full hand pose from wrist sEMG alone: 370 hours of 2kHz, 16-channel EMG from 193 users, paired with ground-truth hand pose from a 26-camera motion-capture rig across 29 gesture stages. Released by Reality Labs at NeurIPS 2024 alongside emg2qwerty — scale comparable to vision-based hand-pose datasets.",
   "facts": [
    "370 hours from 193 users captured in a 26-camera mocap dome — hand tracking with no camera at all, the exact capability that lets the Neural Band track fingers 'without the need to be in view of a camera.'"
   ],
   "latest": "NeurIPS 2024 release (arXiv Dec 2024)",
   "links": [
    {
     "label": "arXiv",
     "url": "https://arxiv.org/abs/2412.02725"
    },
    {
     "label": "GitHub · emg2pose",
     "url": "https://github.com/facebookresearch/emg2pose"
    },
    {
     "label": "Meta AI blog",
     "url": "https://ai.meta.com/blog/open-sourcing-surface-electromyography-datasets-neurips-2024"
    }
   ],
   "why": "A benchmark for recovering full hand pose from wrist EMG alone: 370 hours of 2kHz, 16-channel recordings from 193 users, paired with ground truth from a 26-camera motion-capture rig across 29 gesture stages. Its scale rivals vision-based hand-pose datasets and makes camera-free hand tracking a tractable public research problem.",
   "try": [
    {
     "label": "Dataset and baselines on GitHub",
     "url": "https://github.com/facebookresearch/emg2pose"
    },
    {
     "label": "Read the paper",
     "url": "https://arxiv.org/abs/2412.02725"
    }
   ],
   "params": ""
  },
  {
   "id": "orion",
   "name": "Orion",
   "full_name": "Orion",
   "tag": "The AR glasses prototype a decade in the making",
   "year": 2024,
   "year_label": "2024",
   "cat": "reality",
   "size": 2,
   "status": "closed",
   "open": false,
   "lineage": [],
   "desc": "Meta's first true AR glasses, unveiled September 25, 2024 at Connect after roughly a decade of work: see-through holographic displays with an industry-leading ~70-degree field of view, a wireless compute puck, and an sEMG wristband for control — the research vehicle proving the glasses-plus-neural-interface computing platform. Never for sale; Meta calls it 'one of the most polished product prototypes we've ever developed.'",
   "facts": [
    "~70° diagonal FOV vs HoloLens 2's 52° and Snap Spectacles' 46° — in a glasses form factor.",
    "Each unit costs about $10,000 to build (silicon-carbide lenses are among the components not yet manufacturable affordably at scale, per press hands-ons).",
    "Previously codenamed Project Nazare.",
    "In 2026 Meta ships only a few hundred dev units while thousands of developers apply."
   ],
   "latest": "Developer-kit access expanding in 2026 (~$10,000/unit); consumer successor codenamed Artemis targeted for 2027",
   "links": [
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2024/09/introducing-orion-our-first-true-augmented-reality-glasses"
    },
    {
     "label": "CNBC",
     "url": "https://www.cnbc.com/2024/09/27/hands-on-with-metas-orion-augmented-reality-smart-glasses-prototype.html"
    },
    {
     "label": "uploadvr.com",
     "url": "https://www.uploadvr.com/meta-connect-2024-orion-prototype-ar-glasses"
    }
   ],
   "why": "Meta's first true AR glasses after roughly a decade of work: see-through holographic displays with the largest field of view in the smallest AR form factor to date, a wireless compute puck and an sEMG wristband. Never sold, it proved the glasses-plus-neural-interface platform that Meta Ray-Ban Display and the planned consumer successor build on.",
   "try": [
    {
     "label": "Announcement (Meta Newsroom)",
     "url": "https://about.fb.com/news/2024/09/introducing-orion-our-first-true-augmented-reality-glasses"
    },
    {
     "label": "CNBC hands-on",
     "url": "https://www.cnbc.com/2024/09/27/hands-on-with-metas-orion-augmented-reality-smart-glasses-prototype.html"
    }
   ],
   "params": ""
  },
  {
   "id": "rayban-display",
   "name": "Meta Ray-Ban Display + Neural Band",
   "full_name": "Meta Ray-Ban Display + Neural Band",
   "tag": "The neural wristband ships",
   "year": 2025,
   "year_label": "2025",
   "cat": "reality",
   "size": 2,
   "status": "closed",
   "open": false,
   "lineage": [
    "orion",
    "semg"
   ],
   "desc": "Meta's first consumer smart glasses with a built-in heads-up display, launched September 30, 2025 at $799, bundled with the Meta Neural Band — an EMG wristband that reads muscle signals so you control the glasses without touching anything. In 2026 they gained the Muse Spark assistant, Threads and Instagram integration, live data widgets, and Neural Handwriting.",
   "facts": [
    "The Neural Band descends from Meta's CTRL-labs acquisition — it decodes wrist EMG signals.",
    "2026's Neural Handwriting lets early-access users silently trace prompts on any flat surface instead of saying 'Hey Meta'.",
    "Global rollout began early 2026 with Canada, France, Italy and the UK.",
    "Sold through Best Buy, LensCrafters, Sunglass Hut and Ray-Ban stores.",
    "Six years from CTRL-labs acquisition (2019, reported $500M–$1B) to shipping product."
   ],
   "latest": "Muse Spark AI update + Neural Handwriting early access (summer 2026)",
   "links": [
    {
     "label": "Meta Newsroom",
     "url": "https://about.fb.com/news/2025/09/meta-ray-ban-display-ai-glasses-emg-wristband"
    },
    {
     "label": "CNBC",
     "url": "https://www.cnbc.com/2025/09/17/zuckerberg-799-meta-ray-ban-display-glasses.html"
    },
    {
     "label": "TechCrunch",
     "url": "https://techcrunch.com/2025/09/19/meta-connect-2025-what-to-expect-and-how-to-watch"
    },
    {
     "label": "gizmodo.com",
     "url": "https://www.gizmodo.com/the-meta-ray-ban-displays-dumb-smart-glasses-ai-is-dumb-no-more-2000791752"
    },
    {
     "label": "meta.com · meta ray ban display ai glasse",
     "url": "https://www.meta.com/blog/meta-ray-ban-display-ai-glasses-connect-2025"
    },
    {
     "label": "meta.com · ces 2026 meta ray ban display ",
     "url": "https://www.meta.com/blog/ces-2026-meta-ray-ban-display-teleprompter-emg-handwriting-garmin-unified-cabin-university-of-utah-tetraski"
    },
    {
     "label": "meta.com · 866944989643926",
     "url": "https://www.meta.com/help/ai-glasses/866944989643926"
    }
   ],
   "why": "The first consumer smart glasses with a built-in heads-up display, launched at $799 with the Meta Neural Band, the productized sEMG wristband from Reality Labs' Nature research. It is the first shipping product on the glasses-plus-neural-interface path, and in 2026 it gained Muse Spark, Neural Handwriting and app integrations.",
   "try": [
    {
     "label": "Product page (Meta Store)",
     "url": "https://www.meta.com/ai-glasses/meta-ray-ban-display"
    },
    {
     "label": "Launch announcement (Meta Newsroom)",
     "url": "https://about.fb.com/news/2025/09/meta-ray-ban-display-ai-glasses-emg-wristband"
    }
   ],
   "params": ""
  },
  {
   "id": "ai-glasses",
   "name": "Meta AI Glasses lineup",
   "full_name": "Meta AI Glasses lineup (Ray-Ban Meta Gen 2, Oakley Meta HSTN & Vanguard)",
   "tag": "Camera-and-voice glasses from Ray-Ban and Oakley, now running Muse Spark",
   "year": 2025,
   "year_label": "2025-2026",
   "cat": "reality",
   "size": 1,
   "status": "closed",
   "open": false,
   "lineage": [
    "rayban-display"
   ],
   "desc": "Meta's camera-and-voice AI glasses family with EssilorLuxottica: Ray-Ban Meta Gen 2 and the athlete-focused Oakley Meta line — HSTN (summer 2025) and Vanguard ($499, October 21, 2025) with 3K video capture. Through 2026 the Muse Spark-powered Meta AI assistant rolled out across the lineup in the US and Canada, adding smarter voice conversations and visual answers.",
   "facts": [
    "Oakley Meta Vanguard shoots 3K video through a 12MP, 122-degree wide-angle lens and is built for athletes (Garmin/Strava integrations).",
    "Ray-Ban Meta had sold over 2 million pairs by early 2025, making it the breakout smart-glasses product.",
    "The 2026 Muse Spark update added interruptible multilingual voice chat and shopping assistance."
   ],
   "latest": "Muse Spark assistant rollout across Ray-Ban Meta and Oakley Meta (2026)",
   "links": [
    {
     "label": "TechCrunch",
     "url": "https://techcrunch.com/2025/09/17/meta-unveils-its-new-oakley-meta-vanguard-smart-glasses-for-athletes"
    },
    {
     "label": "androidcentral.com",
     "url": "https://www.androidcentral.com/apps-software/meta/metas-muse-spark-arrives-on-ai-glasses-gen-1-ray-ban-display-waits-for-now"
    },
    {
     "label": "techcabal.com",
     "url": "https://techcabal.com/2025/09/18/ray-ban-meta-glasses-and-every-product-from-meta-connect-2025"
    }
   ],
   "why": "The delivery vehicle for Meta AI outside phones: camera-and-voice glasses built with EssilorLuxottica, expanded in 2025 with Ray-Ban Meta Gen 2 and the athlete-focused Oakley Meta HSTN and Vanguard with 3K video capture. The 2026 Muse Spark rollout put a frontier assistant on people's faces across the lineup in the US and Canada.",
   "try": [
    {
     "label": "Meta AI glasses lineup (Meta Store)",
     "url": "https://www.meta.com/ai-glasses"
    },
    {
     "label": "Oakley Meta Vanguard launch (TechCrunch)",
     "url": "https://techcrunch.com/2025/09/17/meta-unveils-its-new-oakley-meta-vanguard-smart-glasses-for-athletes"
    }
   ],
   "params": ""
  }
 ]
}