metaai·lightalo unofficial · independent

Every figure on the Frontier comparison

The interactive comparison draws charts. This page is the same data as plain tables — every dimension's per-lab value, every dated release and every landmark paper, each with the note that qualifies it, the date it was read and a link to where it came from. It needs no JavaScript and it prints.

14 labs · 13 dimensions · 118 releases · 30 papers. Compiled 2026-09-04. Back to the comparison.

The labs — 14 organisations

The figures on each lab card, in the order the comparison shows them. Ordered by trailing-30-day Hugging Face downloads.

LabFlagshipOpen weightsLicenceHF downloads / moGitHub starsSources
Qwen (Alibaba Cloud)Qwen3.8-2.4T-A95B 2026-08open weights are core strategyMixed — Apache-2.0 on small and mid-size models, custom Qwen licences on the flagships312.6M27.6kHugging Face API — author=Qwen (full paginat ↗ · Qwen3.8-Max License (LICENSE file in the mod ↗ · Artificial Analysis — Comparison of Open Sou ↗ · Qwen — Wikipedia (release history and develo ↗
Google DeepMindGemini 3.8 Flash 2026-09some open weightsMixed — Apache-2.0 on Gemma 4, bespoke Gemma Terms of Use on Gemma 1-3 and the specialist variants136.8M198.8kGemini 3.1 Pro — Model Card, Google DeepMind ↗ · Release notes | Gemini API | Google AI for D ↗ · Gemma 4 — blog.google (2026-04-02) ↗ · Gemma Terms of Use — Google AI for Developer ↗
Meta AI — FAIR + MSLMuse Spark 1.3 2026-09some open weightsMixed — Apache-2.0 on Muse Glimmer, bespoke Llama Community Licence on Llama, 11 distinct licences across the facebook org92.6M59.6kLlama 4 Community License Agreement (Meta, p ↗ · Introducing Muse Glimmer: An Open Agentic Mo ↗ · Muse Spark 1.3: Meta reaches the frontier (A ↗ · Muse Glimmer: Benchmarks and analysis (Artif ↗
OpenAIGPT-6 Astra 2026-09some open weightsApache-2.0 on every model-weight release since 202577.3M121.3kOpenAI API Changelog (developers.openai.com) ↗ · OpenAI docs — GPT-6 Astra model page ↗ · ARC Prize — GPT-6 Astra results (independent ↗ · OpenAI Deployment Safety Hub ↗
NVIDIANVIDIA-Nemotron-3-Ultra-550B-A55B 2026-06open weights are core strategyOpenMDW-1.1 on Nemotron 3 (permissive, but not OSI-listed)48.9MHugging Face API — nvidia/NVIDIA-Nemotron-3- ↗ · OpenMDW License v1.1 ↗ · SEC XBRL — NVIDIA revenue, FY2026 ↗ · NVIDIA — Wikipedia (founding) ↗
Microsoft ResearchPhi-4-reasoning-vision-15B 2026-01some open weightsMIT and Apache-2.042.2MHugging Face API — microsoft org, Phi models ↗ · Hugging Face API — author=microsoft ↗ · GitHub API — deepspeedai/DeepSpeed ↗ · Microsoft Research — Wikipedia (founding) ↗
DeepSeekDeepSeek-V4-Pro-0813 2026-08open weights are core strategyMIT23M104.4kHugging Face API — deepseek-ai/DeepSeek-V4-P ↗ · Artificial Analysis — Comparison of Open Sou ↗ · GitHub API — deepseek-ai/DeepSeek-V3 ↗ · DeepSeek — Wikipedia (founding and structure ↗
Mistral AIMistral-Large-3-675B-Instruct-2512 2025-11some open weightsMixed — Apache-2.0 on Mistral Large 3 and Leanstral, custom on Mistral Medium 3.510.6M10.8kHugging Face API — mistralai/Mistral-Large-3 ↗ · Hugging Face API — author=mistralai ↗ · Artificial Analysis — Comparison of Open Sou ↗ · Mistral AI — Wikipedia (founding) ↗
Z.ai (formerly Zhipu AI)GLM-5.3 2026-08open weights are core strategyMixed — MIT on GLM-5.2 and GLM-5.3-Flash, custom GLM-5.3 License on the flagship10.1MGLM-5.3 License (LICENSE file in the model r ↗ · Hugging Face API — zai-org/GLM-5.3-Flash ↗ · Artificial Analysis — Comparison of Open Sou ↗ · Arena text leaderboard ↗
MiniMaxMiniMax-M3 2026-06some open weightsCustom MiniMax community licences (not OSI-approved), one with territorial exclusions7.6MHugging Face — MiniMaxAI/MiniMax-H3 (licence ↗ · Hugging Face API — MiniMaxAI/MiniMax-M3 ↗ · Artificial Analysis — Comparison of Open Sou ↗ · MiniMax — Wikipedia (founding) ↗
Allen Institute for AI (Ai2)Olmo-3-7B-Instruct 2025-11open weights are core strategyApache-2.06.4M6.7kAI2 model card — allenai/Olmo-3-7B-Instruct ↗ · Hugging Face datasets API — author=allenai ↗ · Hugging Face API — author=allenai ↗ · Allen Institute for AI — Wikipedia (founding ↗
Moonshot AIKimi K3 2026-06open weights are core strategyCustom Kimi K3 License (modified-MIT style, not OSI-approved)4.9M11.1kHugging Face API — moonshotai/Kimi-K3 ↗ · Kimi K3 License (LICENSE file in the model r ↗ · Artificial Analysis — Comparison of Open Sou ↗ · Moonshot AI — Wikipedia (founding) ↗
xAI / SpaceXAIGrok 4.6some open weightsUnverified10.1kArtificial Analysis — Models leaderboard ↗ · Hugging Face API — author=xai-org ↗ · Frontier AI Safety Policies — METR index ↗ · xAI — Wikipedia (founding, SpaceX acquisitio ↗
AnthropicClaude Fable 5.1 2026-09no open weightsNone — no model weights released under any licence0173.7kClaude Fable 5.1 — Claude Docs model page ↗ · Our position on open-weights models — Anthro ↗ · Anthropic's Responsible Scaling Policy ↗ · Donating the Model Context Protocol and esta ↗

Qwen (Alibaba Cloud): 322 of Qwen's 465 repos are Apache-2.0 and none is gated. But its flagships carry scale gates structurally similar to Meta's old 700M-MAU clause — and the Qwen Community License 1.0 on Qwen3.8-Flash-Next is stricter than the Max licence on one clause, having no revenue floor at all on the Model-as-a-Service requirement. Google DeepMind: Gemma 4 (April 2026) moved to Apache 2.0 with no MAU cap. The older Gemma Terms, revised the day before, explicitly exclude Gemma 4 and retain a binding Prohibited Use Policy plus a right for Google to 'restrict (remotely or otherwise) usage' — clauses no OSI licence has. Whether Gemma 4's separately-published Prohibited Use policy is incorporated into the Apache grant was not established from the licence text. Meta AI — FAIR + MSL: The Llama 4 Community Licence requires a separate discretionary licence from Meta above 700M monthly active users, mandates 'Built with Llama' attribution and a 'Llama' name prefix on derivatives — so it is not OSI open source. Muse Glimmer (August 2026) is plain Apache-2.0 with none of those clauses. Both are live at the same time. OpenAI: gpt-oss-120b/20b, gpt-oss-safeguard, privacy-filter and circuit-sparsity all carry license:apache-2.0 — OSI-approved, no MAU cap, no acceptable-use addendum, no gating. Still open weights and not open source: no training data or code. And OpenAI's single most-downloaded artefact, CLIP, declares no licence on Hugging Face at all. NVIDIA: OpenMDW v1.1 requires retaining the agreement and origin notices on distribution, places third-party rights clearance on the licensee, and terminates the grant if the licensee brings patent or copyright litigation over the model materials. It imposes no restrictions on model outputs. Notable that NVIDIA chose it over Apache-2.0. Microsoft Research: Microsoft's open releases are genuinely permissive — Phi and Fara1.5-27B under MIT, Mage-VL under Apache-2.0 — which makes the absence of a current frontier open LLM, rather than the licence, the story. DeepSeek: MIT is OSI-approved with no usage restrictions, no MAU cap, no revenue gate and no attribution obligation beyond the copyright notice. DeepSeek is the most permissive of the large Chinese labs, and the only one shipping MIT at the flagship tier — Qwen, Moonshot and Z.ai all use custom licences on their flagships. Mistral AI: Mistral's open-weight posture is no longer uniform: Mistral Large 3 and Leanstral-1.5-119B-A6B are Apache-2.0, while Mistral-Medium-3.5-128B is tagged license:other with no licence name in the metadata, so its exact terms are unverified. Z.ai (formerly Zhipu AI): The GLM-5.3 License is MIT-shaped but conditional: a Model-as-a-Service operator with affiliates above US$10bn of aggregate 12-month revenue must pass a Z.AI security review, scoped by Z.AI, before any commercial use. MiniMax: The most geographically restrictive licence in this survey: the MiniMax H3 Community License grants rights only in the 'Applicable Territory', defined as worldwide excluding the European Union, the United Kingdom, the Republic of Korea and the United States of America. It also imposes a separate acceptable-use policy. Allen Institute for AI (Ai2): Apache-2.0 on the weights, with training data and code released alongside — the closest thing in this survey to open source in the full sense rather than open weights. The model card adds a non-licence note that the model is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Moonshot AI: MIT-shaped but conditional: a separate Moonshot agreement is required for Model-as-a-Service businesses above US$20M of aggregate 12-month revenue, and attribution is required above 100M MAU or US$20M monthly revenue. Internal use and access via Moonshot's own products are carved out. xAI / SpaceXAI: The xai-org Hugging Face account holds 2 public models drawing 10,131 downloads in the trailing 30 days. This pass did not verify which models those are, their licences, or whether any current Grok weights have been released. Anthropic: There is no Anthropic weight licence to compare, because there is no Anthropic weight release. Its Hugging Face org is verified and hosts 0 models and 14 datasets. This is a licence-category absence, not a restrictive licence. Anthropic's stated position (27 July 2026) is that it 'has never advocated for a ban on open-weights models' and that capability-safe open-weights models are 'a public good'.

Does the lab publish model weights at all?

Is releasing weights the strategy, a side-channel, or something the lab has never done? Read 2026-09-04. Measured in posture (none / some / core-strategy).

LabValueNoteSource
DeepSeekCore strategy — MIT at the flagship tierDeepSeek-V4-Pro-0813 (1.65T) and V4-Flash-0731 (304B) are both MIT.source ↗
Qwen (Alibaba Cloud)Core strategy — the download leader465 models, 312.6M downloads in 30 days, 0% gated.source ↗
Moonshot AICore strategy — few models, all enormous19 models; the flagship Kimi K3 is 2.78T total parameters.source ↗
Z.ai (formerly Zhipu AI)Core strategy — MIT at two tiers154 models; GLM-5.2 and GLM-5.3-Flash are MIT, the GLM-5.3 flagship is not.source ↗
NVIDIACore strategy — and the ecosystem's quantiser927 models; its largest download streams are NVFP4 quantisations of other labs' models.source ↗
Allen Institute for AI (Ai2)Core strategy — weights plus training data970 models and 1,286 datasets; OLMo ships data, code and intermediate checkpoints together.source ↗
Meta AI — FAIR + MSLSome — huge open catalogue, closed frontier2,433 public models across three orgs, but the frontier Muse Spark line is closed and no new Llama has shipped since April 2025. Muse Glimmer (Aug 2026) reopened the open-weight LLM line at a deliberately sub-frontier 30B.source ↗
OpenAISome — Apache-2.0, but not refreshed since Aug 2025gpt-oss-120b/20b created 2025-08-04 and last modified 2025-08-26; six weight releases since mid-2025 in total.source ↗
Google DeepMindSome — open models are deliberately smallGemma 4 spans 2.3B to 31B and is positioned as edge/local; the frontier Gemini line stays closed.source ↗
MiniMaxSome — open weights, restricted territory21 models; the H3 licence grants no rights in the US, EU, UK or South Korea.source ↗
Mistral AISome — Apache-2.0 flagship, custom mid-tier75 models; Mistral Large 3 is Apache-2.0, Medium 3.5 is not.source ↗
Microsoft ResearchSome — but the Phi LLM line has stalled538 models; no new Phi language model since 2026-01-23.source ↗
xAI / SpaceXAISome — but almost nothing public2 models on Hugging Face drawing 10,131 downloads in 30 days. Not verified further in this pass.source ↗
AnthropicNone — zero weights, everHugging Face API returns 0 models for the verified Anthropic org. Its commons contribution is protocols, tooling and 14 datasets.source ↗

A lab with no open weights has no download number, and that is a strategy, not a weakness. Anthropic's absence from every open-weight chart on this page says nothing about the quality of its models — and Anthropic's own published position is that it has never advocated banning open-weight models.

Hugging Face downloads, trailing 30 days

Whose open weights are actually being pulled right now? Read 2026-09-04. Measured in downloads in the trailing 30 days.

LabValueNoteSource
Qwen (Alibaba Cloud)312.6MSummed across all 465 models in the Qwen org. Alibaba's separate Alibaba-NLP org adds a further 7.0M across 70 models.source ↗
Google DeepMind136.8Mgoogle org (136.67M across 1,133 models) plus the small deepmind org (80,904 across 6). 42% of the total is one 2020 encoder.source ↗
Meta AI — FAIR + MSL92.6Mfacebook 65.61M + meta-llama 25.76M + meta-models 1.20M. Third overall.source ↗
OpenAI77.3Mopenai 61.36M + openai-community 15.90M. CLIP and Whisper account for most of it, not gpt-oss.source ↗
NVIDIA48.9M927 models. Its single largest repo is a quantisation of Alibaba's Qwen3.6-35B-A3B at 10.17M.source ↗
Microsoft Research42.2M538 models, led by legacy DeBERTa encoders rather than Phi.source ↗
DeepSeek23.0M104 models. Below meta-llama, because DeepSeek's models are too large for most people to run locally.source ↗
Mistral AI10.6M75 models, still led by the May 2024 Mistral-7B-Instruct-v0.3 at 2.70M.source ↗
Z.ai (formerly Zhipu AI)10.1M154 models.source ↗
MiniMax7.6M21 models; the video model MiniMax-H3 alone is 5.09M of it.source ↗
Allen Institute for AI (Ai2)6.4M970 models, most of them small research artefacts. About 2% of Qwen's volume.source ↗
Moonshot AI4.9MOnly 19 models, all very large.source ↗
xAI / SpaceXAI10.1k2 models.source ↗
Anthropic0 — no weights publishedNot a capability signal. Anthropic has never released weights under any licence; its Hugging Face org hosts 14 datasets and no models.source ↗

A download is a CI job, a quantisation, a Docker layer or a curious human — most are not a user. Org totals are dominated by small legacy encoders rather than flagship LLMs (one 2020 model, google/electra-base-discriminator, is 42% of Google's total; facebook/opt-125m and facebook/contriever lead Meta's). Counts credit the namespace hosting the files, so NVIDIA's biggest streams are other labs' models and Meta's own Muse Glimmer repo is out-downloaded by a community GGUF mirror. Gated repos suppress casual downloads, which cuts against Meta. And a zero here can mean a deliberate closed-weights strategy.

Public models on Hugging Face

How large is the lab's published catalogue? Read 2026-09-04. Measured in public model repositories.

LabValueNoteSource
Meta AI — FAIR + MSL2,433facebook 2,359 + meta-llama 70 + meta-models 4. The largest measured catalogue of any frontier lab; a counter-example search across thirteen other publishers found none larger (EleutherAI at 974 is nearest).source ↗
Google DeepMind1,139google 1,133 + deepmind 6.source ↗
Allen Institute for AI (Ai2)970Many are small research artefacts, which is the point of a nonprofit research catalogue.source ↗
NVIDIA927Heavily composed of speech models and quantisations of other labs' work.source ↗
Microsoft Research538source ↗
Qwen (Alibaba Cloud)465322 of them Apache-2.0.source ↗
Z.ai (formerly Zhipu AI)154source ↗
DeepSeek104source ↗
Mistral AI75source ↗
OpenAI46openai 39 + openai-community 7. Most are legacy research artefacts: CLIP, Whisper, Jukebox, Shap-E, ImageGPT.source ↗
MiniMax21source ↗
Moonshot AI19source ↗
xAI / SpaceXAI2source ↗
Anthropic0The org exists and is verified; it hosts 14 datasets and no models.source ↗

Catalogue size is a research-archive statistic, not a product statistic. Two-thirds of Meta's 2,359-model facebook org carries a non-commercial licence, and most of AI2's 970 are small research checkpoints. A large number here says a lab publishes a lot, not that it publishes well.

Public datasets on Hugging Face

Who releases the data, not just the weights? Read 2026-09-04. Measured in public dataset repositories.

LabValueNoteSource
Allen Institute for AI (Ai2)1,286The only organisation in this set that routinely ships training data alongside weights.source ↗
NVIDIA311Publishes extensive Nemotron training corpora (Nemotron-CC-v2, Nemotron-CC-Math-v1, Nemotron-Post-Training-Dataset-v1 and others).source ↗
Meta AI — FAIR + MSL123Counted on the facebook org.source ↗
Microsoft Research111source ↗
Google DeepMind71The separate google-deepmind org publishes 0.source ↗
OpenAI16source ↗
Anthropic14132,221 downloads in the trailing 30 days. Research data and evaluation benchmarks — hh-rlhf leads all-time at 2,014,117 — not weights.source ↗

This is the dimension on which every large lab, Meta included, trails a small nonprofit. Dataset count is also a crude proxy: one well-documented pre-training corpus matters more than fifty evaluation splits, and this count does not weight them.

What the licence on the flagship open model actually permits

Can you use it in a commercial product without asking anyone? Read 2026-09-04. Measured in this page's four-level coding of the licence text (3 = OSI-approved permissive, 2 = permissive-but-not-OSI, 1 = custom licence with scale, revenue or territory gates, 0 = no weights released).

LabValueNoteSource
DeepSeekMIT — no conditions at allDeepSeek-V4-Pro-0813 and V4-Flash-0731. The most permissive flagship licence at the open-weight frontier.source ↗
Allen Institute for AI (Ai2)Apache-2.0 — plus the training data and codeOlmo-3-7B-Instruct ships with Dolma 3, the Dolci post-training sets, three codebases and intermediate checkpoints.source ↗
OpenAIApache-2.0 — ungated, no MAU capEvery OpenAI weight release since 2025 carries license:apache-2.0. Weights only: no training data or code.source ↗
Microsoft ResearchMIT / Apache-2.0Phi and Fara1.5-27B are MIT; Mage-VL is Apache-2.0.source ↗
Meta AI — FAIR + MSLApache-2.0 on Muse Glimmer — but Llama 4 is notMeta runs both at once. Muse Glimmer (2026-08-10) is plain Apache-2.0 with no MAU cap, no branding requirement and no acceptable-use clause in the grant. The Llama 4 Community Licence still requires a separate Meta licence above 700M MAU, 'Built with Llama' attribution and a 'Llama' name prefix on derivatives — and every Llama repo is gated. Scored on the current flagship open model.source ↗
Google DeepMindApache-2.0 on Gemma 4 — but Gemma 1-3 are notGemma 4 moved to Apache 2.0 in April 2026. The Gemma Terms, revised the day before and explicitly excluding Gemma 4, keep a binding Prohibited Use Policy and Google's right to 'restrict (remotely or otherwise) usage'. Whether Gemma 4's separately-published Prohibited Use policy is incorporated into the Apache grant was not established from the licence text.source ↗
Mistral AIApache-2.0 on Mistral Large 3But Mistral-Medium-3.5-128B is tagged license:other with no licence name in its metadata, so its terms are unverified.source ↗
NVIDIAOpenMDW-1.1 — permissive, not OSI-listedRetain-notices obligation, licensee bears third-party rights clearance, patent-retaliation termination clause, no restriction on outputs.source ↗
Z.ai (formerly Zhipu AI)Custom GLM-5.3 License on the flagshipMIT-shaped but conditional: Model-as-a-Service operators above US$10bn of 12-month revenue must pass a Z.AI security review. GLM-5.2 and GLM-5.3-Flash genuinely are MIT. Arena's leaderboard mislabels the flagship as MIT.source ↗
Qwen (Alibaba Cloud)Custom Qwen3.8-Max License on the flagshipAttribution above 100M MAU or US$20M monthly revenue; separate Qwen licence for Model-as-a-Service businesses above US$50M aggregate revenue. Structurally the shape of Meta's old 700M-MAU clause. 322 of 465 Qwen repos are still Apache-2.0.source ↗
Moonshot AICustom Kimi K3 LicenseSeparate Moonshot agreement required for Model-as-a-Service above US$20M of aggregate 12-month revenue; prominent 'Kimi K3' display above 100M MAU.source ↗
MiniMaxCustom community licence with territorial exclusionsThe MiniMax H3 Community License grants rights only outside the European Union, the United Kingdom, South Korea and the United States — the most geographically restrictive licence in this survey.source ↗
AnthropicNo weights, so no licenceA licence-category absence, not a restrictive licence. There is nothing to compare against Llama's MAU clause or Gemma's Terms because there is no Anthropic weight release.source ↗
xAI / SpaceXAIUnverified2 models on Hugging Face; their licences were not checked in this pass.source ↗

The 0-3 scale is THIS PAGE'S coding of licence texts we read, not a figure any source publishes — the licence names and clauses under it are primary-sourced, the number is editorial. And an OSI-approved licence on the weights is still not 'open source' in the full sense while training data and code are withheld: that is true of Muse Glimmer, Gemma 4 and gpt-oss alike. Only AI2's OLMo ships the whole pipeline.

How much of the lab's open-weight volume needs permission first

Can you just download it, or must you ask and wait for manual approval? Read 2026-09-04. Measured in percentage of trailing-30-day download volume behind a manual Hugging Face access gate.

LabValueNoteSource
Qwen (Alibaba Cloud)0%None of Qwen's 465 repos is gated.source ↗
OpenAI0%None of OpenAI's 39 repos is gated.source ↗
Google DeepMind10.65%352 of 1,133 repos gated — legacy Gemma 1-3 plus EmbeddingGemma, TranslateGemma and the Health AI line (MedGemma, MedSigLIP, MedASR, Path Foundation, Derm Foundation, PaliGemma). Gemma 4 is ungated.source ↗
Meta AI — FAIR + MSL~32% overall — but 100% of LlamaDerived across Meta's three orgs. Every one of the 70 meta-llama repos reads gated='manual', so 100% of that org's 25.76M monthly downloads require requesting access and manual approval. The facebook org gates 168 repos (6.46% of volume, including SAM 3). The new meta-models org, where Muse Glimmer lives, gates nothing.source ↗

Gating is a real barrier that no 'open models' table has ever shown, and it is a single boolean field in a public API. It cuts both ways: it suppresses casual downloads, so it makes Meta's download numbers look SMALLER, not larger. Meta's figure here is derived arithmetic over three orgs (6.46% of the facebook org's 65.6M, 100% of meta-llama's 25.76M, 0% of meta-models' 1.20M) rather than a single published number.

Who built the commons everyone else runs on

Which lab gave away a piece of infrastructure the whole field now depends on — and then gave up control of it? Read 2026-09-04. Measured in GitHub stars on the lab's flagship donated / foundation-hosted infrastructure project.

LabValueNoteSource
Google DeepMindTensorFlow — 198,793Created at Google and hosted in its own tensorflow org. JAX, also Google-originated, has moved out of the google org to jax-ml/jax (36,251 stars).source ↗
Meta AI — FAIR + MSLPyTorch — 102,743Created at Meta and transferred to the newly-formed PyTorch Foundation under the Linux Foundation on 12 September 2022, with a founding board of AMD, AWS, Google Cloud, Meta, Microsoft Azure and NVIDIA. Meta also gave the field FAISS (MIT, 40,852 stars) and co-created ONNX with Microsoft in 2017, handing it to LF AI & Data in 2019. The single largest infrastructure donation of the era — and it predates the current model race.source ↗
AnthropicMCP servers — 90,061The Model Context Protocol, announced 25 November 2024 and donated to the Linux Foundation's Agentic AI Foundation on 9 December 2025. Cross-vendor by design: Claude, ChatGPT, VS Code and Cursor all support it. Anthropic also donated its Petri alignment-auditing toolbox to the nonprofit Meridian Labs in May 2026, explicitly to keep it independent of any AI lab.source ↗
Microsoft ResearchDeepSpeed — 43,061Originated at Microsoft, now a PyTorch Foundation project and out of the microsoft GitHub org (deepspeedai/DeepSpeed). Microsoft also co-created ONNX with Meta.source ↗
OpenAITriton, AGENTS.md — co-founder of the Agentic AI FoundationTriton is MIT-licensed, but its copyright header credits Philippe Tillet for 2018-2020 before OpenAI for 2020-2022, and its design originates in a 2019 Harvard MAPL paper. OpenAI is a co-founder of the Agentic AI Foundation alongside Anthropic and Block, contributing AGENTS.md as a founding project. No star count was measured for Triton in this pass.source ↗

Stars are cumulative and never decay, and they measure developer attention rather than usage. More important: the lab that popularised a project is often not the lab that built it. vLLM (90,919 stars), the inference engine most of the open ecosystem runs on, came out of UC Berkeley's Sky Computing Lab, not any frontier lab; Ray also came from Berkeley; and Triton's own LICENSE credits Philippe Tillet for 2018-2020 before OpenAI for 2020-2022. Infrastructure scoreboards that assign every project to a corporate logo systematically overstate labs and erase universities.

Membership of the Agentic AI Foundation

Who is in the room where the agent-era standard is being written? Read 2025-12-09. Measured in membership tier as listed by the Linux Foundation.

LabValueNoteSource
AnthropicPlatinum — co-founderContributed MCP as a founding project.source ↗
OpenAIPlatinum — co-founderContributed AGENTS.md as a founding project.source ↗
Google DeepMindPlatinumsource ↗
Microsoft ResearchPlatinumsource ↗
Meta AI — FAIR + MSLNot listed at any tierGold and Silver tiers list a further 40+ companies including IBM, Salesforce, SAP, Cisco, Shopify, Uber and Hugging Face. Meta appears at no tier. Meta gave the industry PyTorch in 2022 and is absent from the successor standards body in 2025-26.source ↗

Absence from a member list is not a statement of intent. Report the fact; do not infer a motive. The list is as of the foundation's formation announcement and may have changed since.

Citations on this page's sample of landmark papers

Whose published research is the field standing on? Read 2026-09-04. Measured in summed Semantic Scholar citations across a 30-paper hand-picked sample.

LabValueNoteSource
Google DeepMind407,222 across 6 papersAttention Is All You Need (191,116), BERT (120,372), Vision Transformer (69,024), Chain-of-Thought (21,548), Gemini 1.5 (3,943), Gemma (1,219). Two of the three most-cited papers in the sample are Google's, and both predate the model race.source ↗
Meta AI — FAIR + MSL193,543 across 8 papersPyTorch (55,052), Mask R-CNN (32,929), RoBERTa (31,150), LLaMA (21,439), Llama 3 Herd (18,286), Segment Anything (15,343), DINOv2 (10,184), wav2vec 2.0 (9,160). Note how much of it is not language: vision and speech account for over a third.source ↗
OpenAI176,892 across 5 papersGPT-3 (62,810), CLIP (55,028), GPT-4 Technical Report (26,936), InstructGPT/RLHF (23,807), Whisper (8,311).source ↗
Microsoft Research24,998 across 2 papersLoRA (22,546) and Phi-3 (2,452). LoRA alone underpins essentially all parameter-efficient fine-tuning.source ↗
DeepSeek9,574 across 2 papersDeepSeek-R1 (5,595) and DeepSeek-V3 (3,979). Extraordinary for papers from 2024 and 2025.source ↗
Anthropic7,952 across 2 papersTraining a Helpful and Harmless Assistant (4,335) and Constitutional AI (3,617). Anthropic publishes much of its interpretability work through its own Transformer Circuits Thread rather than conventional venues, so a citation count understates it.source ↗
Mistral AI5,958 across 2 papersMistral 7B (3,851) and Mixtral of Experts (2,107).source ↗
Qwen (Alibaba Cloud)4,875 across 1 paperQwen2.5 Technical Report. The separately-measured Qwen3 Technical Report adds 7,645, which is not in this sample.source ↗
Moonshot AI1,062 across 1 paperKimi k1.5. Kimi K2 adds 370.source ↗
Allen Institute for AI (Ai2)727 across 1 paperOLMo: Accelerating the Science of Language Models.source ↗

This is a SUM OVER A HAND-PICKED SAMPLE of 30 landmark papers (at most a third from Meta), not a measure of total research output — a lab that publishes little appears rarely, which is not a judgement of its work. And citation counts are dominated by paper age: 'Attention Is All You Need' (2017) and BERT (2019) will outrank any 2025 paper indefinitely. Compare within publication-year cohorts, or the chart just re-discovers that old papers are old. Semantic Scholar also holds duplicate records (DeepSeek-V3) and returned nothing at all for the original Gemini 1.0 paper.

Frontier safety policy: what is written down, and when it was last updated

What has each lab committed to in public, at what specificity, and how old is the commitment? Read 2026-09-04. Measured in publication date of the most recent version, as indexed by METR.

LabValueNoteSource
AnthropicResponsible Scaling Policy v3.4 — 8 Jul 2026Nine published versions since September 2023, five of them in the first seven months of 2026. Uses named ASL capability tiers; current frontier models are deployed under ASL-3, and ASL-4 has never been activated.source ↗
xAI / SpaceXAIFrontier Artificial Intelligence Framework — 30 Jun 2026Policy name as rendered by METR.source ↗
Google DeepMindFrontier Safety Framework v3.1 — 17 Apr 2026source ↗
Meta AI — FAIR + MSLAdvanced AI Scaling Framework v2.0 — 8 Apr 2026Renamed from the 'Frontier AI Framework'; adds a loss-of-control risk section alongside chemical/biological and cybersecurity. Section 3.3 Table 1 defines three named risk tiers, anchored qualitatively on catastrophic outcomes, with no numeric capability or evaluation triggers published. Deployment approval sits with the Chief AI Officer; Meta's board provides general product and regulatory compliance oversight, not release-by-release safety sign-off. Meta also publishes per-model Safety & Preparedness Reports.source ↗
Microsoft ResearchFrontier Governance Framework — February 2026METR records the month only; the day is not established.source ↗
OpenAIPreparedness Framework v2.0 — 15 Apr 2025The oldest core framework among the large labs, though OpenAI added a separate Frontier Governance Framework in May 2026. Tracked categories: Biological and Chemical; Cybersecurity; AI Self-improvement. OpenAI publishes per-release system cards on a dedicated Deployment Safety Hub — but as of 2026-09-04 no card had appeared for its new flagship, GPT-6 Astra.source ↗
NVIDIAFrontier safety policy — 17 Feb 2025Indexed by METR; the policy's title was not recorded in this pass. Dates to the Paris AI Action Summit window, as do Amazon's, Cohere's and G42's.source ↗

METR maintains this index and states plainly that 'our indexing these documents should not be considered an endorsement of their substance.' Comparing policies compares what companies have promised, in what detail and how recently — not what they do. Recency is not goodness. And a common false contrast should be avoided: Meta's tiers (Critical / High / Moderate-or-lower) and Anthropic's ASL levels are both named tiers rather than numeric scores; the real difference is in how specifically the capability triggers are defined, not in whether thresholds exist.

Capital expenditure — the only audited numbers on the page

Who is actually spending the money, and can you even see it? Read 2026-09-04. Measured in USD billions, purchases of property and equipment, most recent audited fiscal year (period stated per row).

LabValueNoteSource
Microsoft Research$115.9bn — FY2026 (Jul 2025 - Jun 2026)Additions to property and equipment, audited XBRL. Includes all of Azure, not only AI.source ↗
Google DeepMind$91.4bn — FY2025 (calendar)Audited XBRL. H1 2026 alone was $80.598bn, already 88% of the full 2025 figure. Alphabet's 2026 guidance is reported at $195-205bn from its 22 July 2026 earnings call, but that figure is NOT in its SEC exhibit and could not be reopened from any reachable source — carry it as reported, never as filed.source ↗
Meta AI — FAIR + MSL$69.7bn — FY2025 (calendar)Audited XBRL. H1 2026 was $49.113bn. Meta's own 2026 guidance, verbatim from its Q2 2026 release, is '$130-145 billion, narrowed from our prior outlook of $125-145 billion' — on the finance-lease-inclusive basis, so not directly comparable to this line.source ↗
NVIDIA$6.0bn — FY2026 (ended 25 Jan 2026)Purchases of productive assets, audited XBRL, against FY2026 revenue of $215.9bn. NVIDIA is fabless: it sells the compute rather than buying it, so it does not belong on the same axis as the hyperscalers without this note.source ↗
OpenAIFiles nothingA private company with no audited financials. Its $122bn March 2026 raise is FUNDING, not capex, and must never share a column with these figures.source ↗
AnthropicFiles nothingA private company. It has announced compute COMMITMENTS — $50bn to US infrastructure, $30bn of Azure capacity, up to 5 GW with Amazon and 5 GW of TPUs with Google and Broadcom — but those are contracts, not capital expenditure.source ↗
xAI / SpaceXAIFiles nothingA private SpaceX subsidiary.source ↗

Read every row's period before comparing: fiscal years DO NOT align. Meta and Alphabet are calendar-year filers (FY2025); Microsoft's FY2026 ran July 2025 to June 2026; NVIDIA's FY2026 ended 25 January 2026. Capex is company-wide and no filer breaks out AI capex. NVIDIA is a supplier, not a buyer — it sells the compute — so its tiny figure is not a signal of restraint. OpenAI, Anthropic and xAI file nothing at all, so their cells are empty by necessity. Meta's own headline guidance ($130-145bn for 2026) is on a different basis from the GAAP line here, because it includes principal payments on finance leases — which is why press quotes Meta's 2025 capex as ~$72.2bn while the 10-K line reads $69.691bn. Pick one basis and label it.

What the lab ships that is not a chatbot

How much of the work does an LLM leaderboard structurally fail to see? Read 2026-09-04. Measured in inventory of published non-LLM model families (qualitative).

LabValueNoteSource
Google DeepMindVeo, Imagen and Nano Banana, Lyria, Genie 3, Gemini Robotics, AlphaFold, WeatherNext, AlphaEarth, AlphaEvolve, AlphaGenome, SynthID12+ named model families outside the core LLM line on DeepMind's own model index — biology, weather, planetary mapping, algorithm discovery, robotics and world models. The clearest structural difference from every other lab's portfolio.source ↗
Meta AI — FAIR + MSLSAM / SAM 3 (segmentation), DINOv2 and DINOv3 (vision), ESM and ESMFold (proteins), NLLB-200 (translation), MusicGen (music), wav2vec 2.0 and w2v-BERT (speech), V-JEPA 2 (world models), Brain2Qwerty (non-invasive brain-to-text)Meta's actual open footprint is a perception-and-speech footprint, not a chatbot footprint: 11 of its top 12 downloaded models are not generative chat LLMs. Its 2026 science deployments run on the vision models — SAM 3 and DINOv3 in the DOE Genesis Mission at Lawrence Berkeley National Laboratory, and DINOv3 plus Segment Anything on-device in an ARPA-H-funded assistive-robotics programme at the University of Pittsburgh. Both accounts are from Meta's own blog.source ↗
OpenAICLIP, Whisper, gpt-image-2, gpt-realtime and gpt-audio, GPT-Rosalind (life sciences), GPT-5.6-Cyber — and Sora, being discontinuedCLIP (19.94M/30d), Whisper large-v3-turbo (6.84M) and CLIP-large (6.72M) all out-download gpt-oss-20b (6.21M). Sora the app was discontinued and the Sora and Videos APIs shut down 24 September 2026 with no replacement.source ↗
Microsoft ResearchVibeVoice (speech), Fara (agents), BitNet, Mage-VLMicrosoft's active model lines are no longer LLMs. Its download volume is led by legacy DeBERTa encoders and ResNet-50.source ↗
NVIDIASpeech (parakeet, bigvgan) — plus NVFP4 quantisations of other labs' modelsIts single largest download stream is a quantisation of an Alibaba model.source ↗
MiniMaxMiniMax-H3 video generationA 33B video model with 5.09M downloads in 30 days — most of MiniMax's whole org volume.source ↗
Allen Institute for AI (Ai2)OlmoEarth (geospatial)AI2's newest releases outside language are geospatial rather than generative media.source ↗
AnthropicNone publishedClaude is a text-and-vision LLM family. Anthropic's non-LLM output is datasets, protocols (MCP) and interpretability tooling (circuit-tracer, Petri), not models.source ↗

This row is deliberately not a number, because there is no honest one: the labs publish in different formats and count different things. What IS measured is the download shape — eleven of Meta's twelve most-downloaded Hugging Face models are not generative chat LLMs, and neither are OpenAI's top three, which all out-download gpt-oss-20b. Every standard comparison axis (context window, MMLU, Elo, tokens per second) is defined only for language models, so breadth is invisible by construction rather than by evidence.

Silicon, and things people actually wear

Who controls their own compute, and whose AI exists as a physical object? Read 2026-09-04. Measured in hardware position (qualitative).

LabValueNoteSource
Google DeepMindSeven TPU generations, and now sold beyond its own cloudTPU v7 'Ironwood': 2,307 BF16 TFLOPs and 4,614 FP8 TFLOPs per chip, 192 GB of HBM at 7,380 GBps, 9,216-chip pods (largest addressable slice 2,048 chips). Google trains its own frontier models on it. Reporting that Google began selling TPUs for third-party data centres in 2026 is consistently repeated but could not be confirmed from a primary source — treat as low confidence.source ↗
Meta AI — FAIR + MSLMTIA (internal only) — and the only AI hardware on people's facesFour MTIA generations announced in a single day, 11 March 2026: MTIA 300 for ranking and recommendation training, 400/450/500 aimed at GenAI inference into 2027, with hundreds of thousands of chips already deployed. Inference- and recsys-first, and not available outside Meta. Meanwhile Ray-Ban Meta glasses ship at retail — and in August 2026 Meta gave 15,000 pairs free to every blind and visually impaired adult supported by Vision Ireland, built with EssilorLuxottica.source ↗
NVIDIASells the compute rather than buying itFY2026 revenue of $215.9bn on capex of $6.0bn. It is a supplier to almost every other lab on this page, and an investor in several of them.source ↗
AnthropicNo silicon of its own — deliberately multi-vendorTrains and serves Claude across AWS Trainium, Google TPUs and NVIDIA GPUs at once, with agreements for up to 5 GW from Amazon, 5 GW of next-generation TPUs with Google and Broadcom, and GPU capacity from SpaceX. Note an internal tension in Anthropic's own wording: its April 2026 page says 'multiple gigawatts' for the Google/Broadcom deal, its May 2026 page says five.source ↗
OpenAINo shipping consumer hardware as of 2026-09-04Its first device is targeted for H2 2026 and does not exist at retail. OpenAI acquired Jony Ive's io for a reported ~$6.4-6.5bn; the form factor is rumoured. Its announced compute commitments across Nvidia, Oracle, AMD and Broadcom are reported at ~33 GW, but the components as reported sum to about 26 GW, so the headline and its parts are not self-consistent — low confidence, do not print a gigawatt figure without a primary filing.source ↗

Hardware comparisons imply an independence the supply chain does not support. Meta designs its own chips and is reported (by The Information, unconfirmed by either company) to rent Google's; Anthropic runs on Google, Amazon and NVIDIA silicon simultaneously; OpenAI's compute is contracted across Nvidia, Oracle, AMD, Broadcom and AWS. Nobody here is vertically integrated except Google. Note also that Meta's newsroom is silent on whether MTIA could be offered externally — say 'Meta has not offered MTIA outside its own infrastructure', not that it cannot be.

The release race — 118 dated releases

Newest first.

DateLabReleaseKindNoteSource
2026-09-03OpenAIGPT-6 Astraclosed'Our most capable model, built for the hardest end-to-end work.' 1,050,000-token context, knowledge cutoff 30 April 2026, $10 in / $50 out per million tokens — 2.5x GPT-5.6 Sol's price, reversing the cheaper-with-each-generation trend. It adds asynchronous misalignment monitoring that can halt a conversation for review. ARC Prize verified 95.0% on ARC-AGI-2 (max) and 62.71% on ARC-AGI-3 under the standard harness. No system card had appeared on OpenAI's safety hub as of 2026-09-04.source ↗
2026-09-03Google DeepMindLyria 3.5 music models enter public previewclosedBoth variants take text and image inputs and generate 44.1 kHz stereo audio; $0.04 per song for clips, $0.08 for full-length.source ↗
2026-09-02Google DeepMindGemini 3.8 Flashclosed'Our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.' Google's eighth Flash-line release of 2026, and still no 3.5 Pro.source ↗
2026-09-02Meta AI — FAIR + MSLMuse Spark 1.3closedMeta's fourth Muse Spark release in five months. Artificial Analysis independently measured 62 for the limited-preview (max) variant — behind Claude Fable 5.1 (66) and Claude Opus 5 (63), tied with Claude Fable 5 — and 61 for the generally-available (xhigh). AA also records xhigh as having 'the lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index', at $0.55 per Intelligence Index task. $1.25 in / $4.25 out per million tokens, 1M context.source ↗
2026-09-02Meta AI — FAIR + MSLZuckerberg promises open weights for Muse SparkproductAn X post promising open weights 'soon', with no date. SOURCES CONFLICT on the version: The Register gives none; Wikipedia's Muse Spark article says 1.2 is planned as open-weight. As of 2026-09-04 Meta's meta-models org contains only four repos, all Muse Glimmer — no Muse Spark weights of any version exist. The promise is two days old: outstanding, not overdue.source ↗
2026-09-01AnthropicClaude Fable 5.1 and Claude Mythos 5.1closed$10/$50 per million tokens, 1M context. Anthropic's self-reported table shows Fable 5.1 at 52.6% on Terminal-Bench-Science and 1853 on GDPval-AA v2, against GPT-5.6 Sol at 22.4% and 1711 — with the competitor column run by Anthropic itself, which OpenAI has not endorsed. No Meta model appears anywhere in the comparison. Artificial Analysis independently had Fable 5.1 leading its Intelligence Index at 66 on 2026-09-04.source ↗
2026-08-27Google DeepMindGemini Omni 1.1 Flash GAclosedsource ↗
2026-08-26OpenAIAssistants API shut downproductReplaced by the Responses and Conversations APIs after a one-year migration window. The Responses API is now OpenAI's primary agentic surface, and GPT-6 Astra requires it for tool calling.source ↗
2026-08-25Z.ai (formerly Zhipu AI)GLM-5.3 and GLM-5.3-Flashopen-weightsGLM-5.3 (753B) is joint #1 among open-weight models on Artificial Analysis's index at 60 — but is under a bespoke licence with a US$10bn Model-as-a-Service security-review gate, despite Arena labelling it MIT. GLM-5.3-Flash (321B) genuinely is MIT and scores 57, 4th on the same board.source ↗
2026-08-24Qwen (Alibaba Cloud)Qwen3.8-Flash-Nextopen-weights180B parameters under the custom Qwen Community License 1.0 — which, on the Model-as-a-Service clause, is STRICTER than Qwen's flagship licence, having no revenue floor at all.source ↗
2026-08-20Google DeepMindGemma family passes 1 billion cumulative downloadsproductGoogle's own announcement, with over 100,000 community-published Gemma variants. Google does not break the total down by model, year or platform, so it is not comparable to any per-platform figure.source ↗
2026-08-13Google DeepMindGemini 3.7 Flashclosedsource ↗
2026-08-13DeepSeekDeepSeek-V4-Pro-0813open-weights~1.65 trillion total parameters under MIT — no MAU cap, no revenue gate, no attribution obligation. Independently measured at 53, 5th on the open-weight board. Date is the Hugging Face repo creation date.source ↗
2026-08-12Meta AI — FAIR + MSL15,000 Ray-Ban Meta glasses given free through Vision IrelandproductFree AI glasses for every blind and visually impaired adult supported by Vision Ireland, built with EssilorLuxottica, with training funded by Meta. Text reading, object identification, live translation and hands-free calls. Do not extrapolate 15,000 into market share — it is a donation, not a shipment number.source ↗
2026-08-10Meta AI — FAIR + MSLMuse Glimmer 30B under Apache 2.0open-weightsMeta's return to open weights, and the first open-weight LLM it has released under a commercially permissive OSI-approved licence rather than a bespoke Llama Community Licence. 29,776,626,688 parameters in BF16 by safetensors metadata (Meta's card says ~29.6B; the '30B' is the marketing name), 131,072+ context, distilled from the closed Muse Spark. The Hugging Face repo was created 2026-08-09, the day before the announcement. Open WEIGHTS, not open source: no training data, no training code.source ↗
2026-08-10Meta AI — FAIR + MSL'The Future is for Everyone' — Zuckerberg essayproductThe 2026 open-weight commitment, verbatim: 'Now that Meta Superintelligence Labs are up and running, we will resume releasing some open source models soon.' Note the word 'some'. The essay also commits Meta's independent board of directors to approving the safety criteria for releasing models — but it contains no pledge to open the weights of Muse Spark or any specific model. Set it beside the 2024 essay's 'the first frontier-level open source AI model' and let the reader see the shift.source ↗
2026-08-09OpenAIChatGPT Atlas retiredproductTen months after its macOS release. Announced in March 2026 as part of a merge of Atlas, the ChatGPT desktop app and Codex into a single desktop app.source ↗
2026-08-09Meta AI — FAIR + MSLMoEViE familyopen-weightsPublished from the facebook org between 9 and 12 August 2026 — more evidence that Meta's research weight releases never stopped, only the Llama ones did.source ↗
2026-08-08Qwen (Alibaba Cloud)Qwen3.8-2.4T-A95Bopen-weights~2.45 trillion total parameters — the second-largest open-weight model measured. Released under the custom Qwen3.8-Max License, with attribution required above 100M MAU and a separate licence required for Model-as-a-Service businesses above US$50M revenue. Independently measured at 58, 3rd on the open-weight board.source ↗
2026-08-07OpenAIgpt-5.6-cyber, gpt-daybreak-red and gpt-daybreak-blueclosedA gated two-tier cybersecurity access programme: Daybreak Blue for vetted defenders, Daybreak Red for purpose-trained offensive security. On OpenAI's internal Advanced Cybersecurity eval, GPT-5.6-Cyber completes 95% against 1.5% for GPT-5.6 Sol on general access — the gating is real and measurable, but the eval is OpenAI's own. OpenAI used it to find CVE-2026-15903 in Chrome's V8 engine.source ↗
2026-08-05Qwen (Alibaba Cloud)Qwen3.8-27Bopen-weights27.8B dense parameters, Apache-2.0 — Qwen's volume driver at 5.25M downloads in 30 days and 13,828 likes. Its flagship is licensed far more restrictively.source ↗
2026-08-05Meta AI — FAIR + MSLMuse Spark 1.2closedDistributed through the Meta Model API and OpenRouter. A lower-priced 'Contributor' tier lets Meta retain submitted data to improve its products; data submitted through the higher-priced standard tier is not retained. muse-spark-1.2 (xHigh) later reached rank 5 on Arena's text leaderboard.source ↗
2026-08-05Meta AI — FAIR + MSLMuse Code betaproductA terminal coding agent for large repositories, powered by Muse Spark and positioned against OpenAI Codex and Anthropic's Claude Code. It fans work out to sub-agents in isolated worktrees. No benchmark numbers were published at launch — worth saying, given every competitor ships SWE-Bench figures.source ↗
2026-07-31DeepSeekDeepSeek-V4-Flash-0731open-weights304B total parameters, MIT. DeepSeek's most-downloaded current model at 4.43M in 30 days. Date is the Hugging Face repo creation date.source ↗
2026-07-30OpenAIFast mode service tier; GPT-5.6 Luna cut 80%, Terra cut 20%closedFast mode replaced Priority Processing at up to 2.5x standard speed for twice the price. Luna's cut makes it the cheapest model in OpenAI's current frontier family at $0.20/$1.20 per million tokens.source ↗
2026-07-30Google DeepMindGemini Robotics ER 2 previewclosed'Advanced spatial reasoning, agentic code execution, multi-step tool orchestration, video moment finding, progress classification.' Robotics is an area where Meta has no comparable shipped model line — but both Google endpoints remain in preview, so this is a research-access lead rather than a shipped-product one.source ↗
2026-07-28MiniMaxMiniMax-H3open-weightsA 33B VIDEO GENERATION model, not a language model — 5.09M downloads in 30 days. Its community licence grants rights only outside the EU, UK, South Korea and the United States. SOURCES CONFLICT on the date: the Hugging Face repo was created 2026-07-28, the licence text dates the release to 2026-08-02.source ↗
2026-07-28OpenAIgpt-transcribe and gpt-live-transcribeclosedsource ↗
2026-07-28AnthropicMCP specification revision 2026-07-28productThe current protocol version. Core server features Resources, Prompts and Tools; optional extensions for asynchronous Tasks, Skills over MCP, and inline interactive MCP Apps.source ↗
2026-07-27Meta AI — FAIR + MSLRAMMP assistive robotics: DINOv3 and Segment Anything on-deviceproductAn assistive-robotics programme at the University of Pittsburgh funded up to $41.5M by ARPA-H, running Meta's vision models on-device for wheelchair users, with partners including Kinova Robotics, LUCI Mobility, Carnegie Mellon, Cornell, Northeastern and Purdue. The clearest illustration of why open weights matter for a reason benchmarks never capture: on-device, offline, no API dependency. Meta's own blog is the source.source ↗
2026-07-25Microsoft ResearchMage-VLopen-weightsApache-2.0, 4.7B parameters.source ↗
2026-07-24Microsoft ResearchVibeVoice-ASR-BitNetopen-weightssource ↗
2026-07-24AnthropicClaude Opus 5closed$5/$25 per million tokens, 1M context. Anthropic's launch claims were almost all cost-normalised relative comparisons rather than absolute scores — 'Opus 5 outperforms every other model at any given cost, surpassing Fable 5's best result at just over a third of the cost' on OSWorld 2.0, for instance. Anthropic's docs still recommend starting with Opus 5 for most workloads, ahead of its newer Fable 5.1.source ↗
2026-07-21Meta AI — FAIR + MSLGenesis Mission deployment: SAM 3 and DINOv3 at Lawrence Berkeley National LaboratoryproductMeta's vision models power the DOE Genesis Mission's SYNAPS-I imaging pipeline across five national laboratories (Berkeley, Argonne, Brookhaven, Oak Ridge, SLAC). Meta says an analysis that previously required 'a month of expert annotation per time step' now takes about 15 minutes, and that Advanced Light Source detectors went from one image every six seconds to 100,000 images per second. Meta's own blog is the source; the DOE-side framing is Meta's, not DOE's.source ↗
2026-07-21Google DeepMindGemini 3.6 Flash, Gemini 3.5 Flash-Lite GA, Gemini 3.5 Flash CyberclosedThree models in one day, and still no 3.5 Pro. Google said the same week that it had begun its most ambitious pre-training run yet, for Gemini 4.source ↗
2026-07-17Microsoft ResearchFara1.5-27Bopen-weightsMIT. An agent model — part of Microsoft's shift away from general LLMs.source ↗
2026-07-09OpenAIGPT-5.6 Sol, Terra and Luna; ChatGPT WorkclosedThree named capability tiers, all 1,050,000-token context. All three were classed 'High' capability in both Cybersecurity and Biological/Chemical under OpenAI's own Preparedness Framework — the first time smaller members of a family received a High designation. Independently, Artificial Analysis scored Sol at 59 at launch, one point behind Claude Fable 5; ARC Prize verified it at 92.5% on ARC-AGI-2.source ↗
2026-07-09Meta AI — FAIR + MSLMuse Spark 1.1 and the Meta Model APIclosedThe Meta Model API reached public preview, giving developers paid access to Muse Spark. LOW CONFIDENCE on the date: the SiliconANGLE source was not reopened and is corroborated only indirectly by Wikipedia's 9 July date for Muse Spark 1.1. Claims that this was Meta's first paid model access could not be sourced and must not be printed.source ↗
2026-07-07Meta AI — FAIR + MSLMuse Image launched, Muse Video previewedclosedThe first media-generation models from Meta Superintelligence Labs. Both closed; no weights, no licence, no parameter counts, no resolution or duration specs. Muse Image embeds 'Content Seal', Meta's invisible watermarking system. Meta says Muse Image ranked #2 on Arena for text-to-image and Muse Video #3 for text-to-video as of 5 July 2026 — Meta's own claim about a third-party leaderboard, not independently confirmed, and now stale. It drew immediate criticism for generating images incorporating other people's public Instagram photos on an opt-out basis.source ↗
2026-07-06OpenAIgpt-realtime-2.1 and gpt-realtime-2.1-miniclosedsource ↗
2026-07-02Google DeepMindGemma 4 Technical ReportpaperarXiv:2607.02770, revised 24 July 2026, CC BY 4.0. Describes dense and MoE architectures from 2.3B to 31B and a unified encoder-free 12B that ingests raw audio and image patches. Several hundred authors.source ↗
2026-06-30AnthropicClaude Sonnet 5closed$2/$10 per million tokens, 1M context. The launch post published no absolute Sonnet 5 benchmark score in body text — the numbers it quotes are REGRADED Sonnet 4.6 figures. The introductory price was made permanent on 10 August and the scheduled rise to $3/$15 cancelled.source ↗
2026-06-16Z.ai (formerly Zhipu AI)GLM-5.2open-weights753B total parameters, MoE, genuinely MIT. Date is the Hugging Face repo creation date.source ↗
2026-06-13Moonshot AIKimi K3open-weights~2.78 trillion total parameters — the largest safetensors parameter count of any model checked in this survey. Joint #1 open-weight model on Artificial Analysis's Intelligence Index at 60. Custom Kimi K3 License, not OSI-approved. Date is the Hugging Face repo creation date.source ↗
2026-06-12AnthropicUS export-control directive suspends Claude Fable 5 and Mythos 5productAnthropic received the directive at 5:21pm ET and disabled both models for all customers the same day to comply, while stating: 'We disagree that the finding of a narrow potential jailbreak should be cause for recalling a commercial model deployed to hundreds of millions of people.' The controls were lifted 30 June and Fable 5 returned globally on 1 July.source ↗
2026-06-11Meta AI — FAIR + MSLMobileMoE causal language modelsopen-weightsPublished under a FAIR non-commercial research licence, from the facebook org — during the sixteen-month silence in the meta-llama org.source ↗
2026-06-09AnthropicClaude Fable 5 and Claude Mythos 5closed$10/$50 per million tokens. Safeguards route requests related to cybersecurity, biology and chemistry or distillation to Claude Opus 4.8 instead; Anthropic says this triggers in under 5% of sessions.source ↗
2026-06-09Google DeepMindDiffusionGemma-26B-A4Bopen-weightsApache-2.0, 1.17M downloads in 30 days. No Google announcement blog for this model could be located, so only the repo's existence, date, licence and statistics are verified — do not characterise it beyond that.source ↗
2026-06-03NVIDIANVIDIA-Nemotron-3-Ultra-550B-A55Bopen-weights~560B total / 55B active, under the OpenMDW License v1.1 rather than Apache-2.0. Independently measured at 38 — the highest-placed US-headquartered open-weight model on Artificial Analysis's board, above Meta's Muse Glimmer at 35.source ↗
2026-06-03OpenAIGPT-Rosalind-5.5 system card publishedclosedLOW CONFIDENCE. What is directly verified is that a system card with this date appears on OpenAI's Deployment Safety Hub, confirming a life-sciences model exists and was documented. The reported April 2026 launch and the named partner list could not be verified and must not be printed.source ↗
2026-06-02MiniMaxMiniMax-M3open-weights427B total parameters, MoE, under the custom 'minimax-community' licence. Independently measured at 45 on the open-weight Intelligence Index, 8th on that board. Date is the Hugging Face repo creation date.source ↗
2026-05-19Google DeepMindGemini 3.5 Flash GAclosedGemini 3.5 Pro was teased alongside it, with Google saying it was already in internal use and would roll out 'next month'. As of 2026-09-04 it has not shipped.source ↗
2026-05-07Google DeepMindGemini 3.1 Flash-Lite GAclosedsource ↗
2026-05-07Microsoft ResearchPhi-Ground-Anyopen-weightsA GUI-grounding model, not a general LLM — which is the point: Microsoft's new releases are no longer language models.source ↗
2026-05-07OpenAIgpt-realtime-2, gpt-realtime-translate, gpt-realtime-whisperclosedsource ↗
2026-05-07AnthropicPetri donated to Meridian LabsproductAnthropic's open-source alignment auditing toolbox, donated at version 3.0 to a nonprofit explicitly 'to ensure that Petri remains independent of any AI lab'. The UK AI Security Institute has made it a major part of how it evaluates models for propensity to sabotage AI research.source ↗
2026-04-23OpenAIGPT-5.5 and GPT-5.5 ProclosedAnnounced 23 April; the API changelog dates the Chat Completions / Responses / Batch release to 24 April. OpenAI's launch charts claimed wins over Gemini 3.1 Pro and Claude Opus 4.5 — self-reported comparisons, not independent ones.source ↗
2026-04-21OpenAIgpt-image-2closedsource ↗
2026-04-17OpenAIopenai/privacy-filteropen-weightsApache-2.0 PII-detection token classifier — OpenAI's only new open-weight release of 2026 to date. 328,426 downloads in 30 days.source ↗
2026-04-08Meta AI — FAIR + MSLMuse SparkclosedThe first model in Meta's new Muse series from Meta Superintelligence Labs — and it shipped CLOSED-weight, breaking with the Llama precedent. Meta described it as 'small and fast by design', not as a frontier flagship, and said it hoped to open-source future versions. Asked whether Llama would be developed further, a Meta spokesperson said only: 'Our current Llama models will continue to be available as open source.'source ↗
2026-04-08Meta AI — FAIR + MSLAdvanced AI Scaling Framework v2.0productMeta's safety framework, renamed from the 'Frontier AI Framework', adding a loss-of-control risk section. Three named risk tiers defined qualitatively against catastrophic outcomes, with no numeric capability triggers published. Deployment approval sits with the Chief AI Officer.source ↗
2026-04-07AnthropicProject Glasswing and Claude Mythos PreviewclosedA cyberdefence initiative distributing a restricted, fewer-safeguards Claude to vetted defenders at $25/$125 per million tokens, with $100M of model usage credits committed. Launch partners included AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks.source ↗
2026-04-02Google DeepMindGemma 4 under Apache 2.0open-weightsGoogle's own words: 'Gemma 4 is released under a commercially permissive Apache 2.0 license.' Five sizes from 2.3B effective to 31B dense; 128K-256K context; the E2B, E4B and 12B Unified variants take audio as well as image. The bespoke Gemma Terms, revised the day before, explicitly exclude Gemma 4 and continue to govern Gemma 1-3.source ↗
2026-03-31Mistral AIMistral-Medium-3.5-128Bopen-weightsTagged license:other with no licence name in the metadata — Mistral moving OFF Apache-2.0 at the mid tier, in the opposite direction to Google and Meta. Independently measured at 30 on Artificial Analysis's open-weight index, 12th of 14.source ↗
2026-03-30OpenAIopenai/codex-plugin-ccproductApache-2.0. 'Use Codex from Claude Code to review code or delegate tasks' — an official OpenAI interoperability plugin for a competitor's coding agent. 32,732 stars.source ↗
2026-03-24OpenAISora discontinuation announcedproductPeak downloads of 3.33M across iOS and Google Play in November fell to 1.13M by February; approximately $2.1M of lifetime in-app purchase revenue; a $1bn Disney investment-and-licensing deal collapsed with no money changing hands. The API and Videos API shut down 24 September 2026 with no recommended replacement. The app's own shutdown date was never stated.source ↗
2026-03-17OpenAIgpt-5.4-mini and gpt-5.4-nanoclosedsource ↗
2026-03-12OpenAISora API expanded — character references, 20-second generations, 1080pclosedTwelve days before OpenAI announced it was discontinuing Sora.source ↗
2026-03-11Meta AI — FAIR + MSLMTIA 300, 400, 450 and 500 announcedproductFour custom-silicon generations announced in a single day, against an industry norm of one every one to two years. MTIA 300 handles ranking and recommendation training; 400/450/500 target GenAI inference into 2027. Hundreds of thousands of chips already deployed — and none available to anyone outside Meta.source ↗
2026-03-05OpenAIGPT-5.4 and GPT-5.4 Proclosedsource ↗
2026-03-03OpenAIgpt-5.3-chat-latestclosedsource ↗
2026-03-02Google DeepMindFirst google/gemma-4-* repos staged on Hugging Faceopen-weightsRepos created three to four weeks before the public announcement. Repo-creation timestamps should not be quoted as a launch date — but they do show the staging.source ↗
2026-02-28Allen Institute for AI (Ai2)Olmo-Hybrid-Think-SFT-7Bopen-weightsThe newest OLMo-family language model. There is no OLMo 4; AI2's later 2026 releases are geospatial.source ↗
2026-02-26OpenAIopenai/symphonyproductApache-2.0. 'Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding agents.' 27,026 stars.source ↗
2026-02-24OpenAIgpt-5.3-codexclosedsource ↗
2026-02-23OpenAIgpt-realtime-1.5 and gpt-audio-1.5closedsource ↗
2026-02-20Allen Institute for AI (Ai2)Olmo-Hybrid-Instruct-DPO-7Bopen-weightssource ↗
2026-02-19Google DeepMindGemini 3.1 Pro Previewclosed1M input context, 64K output. As of 2026-09-04 this is still the only Pro-tier model of the 3.x line in the API — six months without a Pro refresh.source ↗
2026-02-11Google DeepMindGemini Deep Think advanced version (January 2026)paperReported at roughly 90% on IMO-ProofBench Advanced and roughly 38% on FutureMath Basic in a DeepMind blog post; results graded by unnamed human experts. Google gives no release date or model name for this version — do not treat it as a shipped model.source ↗
2026-01-28Google DeepMindAlphaGenome source code releasedproductThe repository software is Apache-2.0 and its examples CC BY 4.0; the hosted API is free for non-commercial use only, with a commercial offering in early testing. This is an API-access model, not a downloadable weights release.source ↗
2026-01-23Microsoft ResearchPhi-4-reasoning-vision-15Bopen-weightsMIT. The newest Phi language model — and there has been none since. A Hugging Face-wide search for a Phi-5 returns nothing.source ↗
2025-12-11OpenAIcircuit-sparsity interpretability weightsopen-weightsApache-2.0. 489 downloads in 30 days but 209 likes — a research artefact, not a product, and worth counting if you ask which labs publish interpretability artefacts.source ↗
2025-12-09AnthropicMCP donated to the Agentic AI FoundationproductA directed fund under the Linux Foundation, co-founded by Anthropic, Block and OpenAI, with support from Google, Microsoft, AWS, Cloudflare and Bloomberg. At the time: 97M+ monthly SDK downloads and over 10,000 active public MCP servers. Meta is not a member at any tier.source ↗
2025-11-28Mistral AIMistral-Large-3-675B-Instruct-2512open-weightsApache-2.0. Date is the Hugging Face repo creation date; trade reporting gives a 2 December 2025 release. It draws 2,106 downloads in 30 days — three orders of magnitude below the 2024-era Mistral-7B.source ↗
2025-11-27Meta AI — FAIR + MSLomniASR familyopen-weightsApache-2.0 speech models from the facebook org.source ↗
2025-11-19Allen Institute for AI (Ai2)Olmo-3-7B-Instructopen-weightsApache-2.0, shipped with the Dolma 3 pre-training corpus, the Dolci post-training datasets, three codebases and intermediate step checkpoints. The most complete open release in this survey.source ↗
2025-11-07Meta AI — FAIR + MSLSAM 3open-weightsMeta's flagship 2026-era segmentation model — gated, under an undisclosed custom licence ('other' on Hugging Face), and still drawing 2.09M downloads in 30 days. SAM 1 was Apache-2.0 and ungated: Meta's vision line has moved from permissive to restricted over three generations, in the opposite direction to its LLM line.source ↗
2025-11-01Google DeepMindGemini 3 Proclosed1M-token input context, 64K output, sparse mixture-of-experts. Google's own model card gives the month only; the widely-cited 18 November date comes from secondary coverage.source ↗
2025-10-21OpenAIChatGPT Atlas browserproductRetired 9 August 2026, roughly ten months after its macOS release.source ↗
2025-10-01AnthropicClaude Haiku 4.5closed$1/$5 per million tokens, 200K context. Still the cheapest current Claude, and the only one left on the previous tokenizer.source ↗
2025-09-18OpenAIgpt-oss-safeguard-120b and -20bopen-weightsApache-2.0 safety classifiers that reason about a user-supplied written policy instead of emitting a fixed score. Date is the Hugging Face repo creation date; the public announcement date is unverified.source ↗
2025-09-08Meta AI — FAIR + MSLmap-anythingopen-weightsApache-2.0, from the facebook org.source ↗
2025-08-04OpenAIgpt-oss-120b and gpt-oss-20bopen-weightsApache-2.0, ungated, no MAU cap. 117B total / 5.1B active and 21B / 3.6B respectively. Still OpenAI's only open-weight LLMs, and unmodified since 26 August 2025.source ↗
2025-07-21Google DeepMindGemini Deep Think achieves gold-medal standard at the IMOproduct35 of 42 points, five of six problems solved perfectly, graded and certified by IMO coordinators using the same criteria applied to students. OpenAI reported the same 35/42 in 2025 but self-graded rather than officially entered — the score parity is real; only the grading provenance differs.source ↗
2025-05-31Meta AI — FAIR + MSLV-JEPA 2 weights under MITopen-weightsfacebook/vjepa2-vitl-fpc64-256, plain MIT — more permissive than Apache-2.0, and the counter-example that kills any claim that Apache-2.0 is Meta's most permissive licence ever.source ↗
2025-05-29AnthropicCircuit-tracing tools open-sourcedproductAttribution-graph tooling released publicly and demonstrated on Gemma-2-2b and Llama-3.2-1b — Anthropic open-sourcing the microscope but not the specimen, and pointing it at Google's and Meta's weights.source ↗
2025-05-22AnthropicClaude 4, and Claude Code generally availableproductClaude Code shipped GA with VS Code and JetBrains integrations, GitHub Actions support and an SDK. Anthropic activated its ASL-3 protections the same day, with Claude Opus 4 the first commercial model released under them.source ↗
2025-04-28Meta AI — FAIR + MSLLlama-Prompt-Guard-2open-weightsThe last upload of any kind to Meta's meta-llama org. Nothing has been published there in the sixteen months since.source ↗
2025-04-27Qwen (Alibaba Cloud)Qwen3-0.6Bopen-weightsStill the single most-downloaded Qwen model, at 21.4M in the trailing 30 days. Tiny models dominate raw download counts at every lab.source ↗
2025-04-23Meta AI — FAIR + MSLLlama-Guard-4-12Bopen-weightsA safety classifier, published after Llama 4 itself.source ↗
2025-04-11Meta AI — FAIR + MSLPerception Encoder familyopen-weightsApache-2.0. Part of the evidence that Meta never actually stopped publishing open weights — only Llama ones.source ↗
2025-04-05Meta AI — FAIR + MSLLlama 4 Scout and Maverickopen-weightsLicence effective 5 April 2025; the Hugging Face repos were created 1-2 April. The Llama 4 Community Licence requires a separate Meta licence above 700M MAU, 'Built with Llama' attribution and a 'Llama' name prefix on derivatives. Llama 4 Behemoth, the ~2T-parameter flagship previewed alongside, was repeatedly delayed and never released.source ↗
2025-01-01DeepSeekDeepSeek-R1paperarXiv:2501.12948. 5,595 citations — extraordinary for a paper this recent.source ↗
2024-12-01Qwen (Alibaba Cloud)Qwen2.5 Technical ReportpaperarXiv:2412.15115. 4,875 citations.source ↗
2024-11-25AnthropicModel Context Protocol announcedproductMEDIUM CONFIDENCE on the date: the announcement page was not independently reopened in the verification pass. Consistent with the modelcontextprotocol GitHub org's creation date of 2024-09-20.source ↗
2024-07-23Meta AI — FAIR + MSLLlama 3.1 405B and 'Open Source AI Is the Path Forward'open-weightsZuckerberg's essay called it 'the first frontier-level open source AI model' and predicted future Llama models would become the most advanced in the industry. MEDIUM CONFIDENCE: the 2024 essay could not be reopened in the verification pass, so re-verify its quotations before setting them in type.source ↗
2024-07-01Meta AI — FAIR + MSLThe Llama 3 Herd of ModelspaperarXiv:2407.21783. 18,286 citations.source ↗
2024-05-22Mistral AIMistral-7B-Instruct-v0.3open-weightsStill Mistral's most-downloaded model two years later, at 2.70M in the trailing 30 days.source ↗
2024-04-01Microsoft ResearchPhi-3paperarXiv:2404.14219. 2,452 citations. The Phi line has since stalled.source ↗
2024-03-01Google DeepMindGemini 1.5paperarXiv:2403.05530. 3,943 citations.source ↗
2024-03-01Google DeepMindGemmaopen-weightsarXiv:2403.08295, 1,219 citations. Released under the bespoke Gemma Terms of Use, not an OSI licence — that only changed with Gemma 4 in 2026.source ↗
2024-02-01Allen Institute for AI (Ai2)OLMo: Accelerating the Science of Language ModelspaperarXiv:2402.00838. 727 citations. The first full-pipeline open release: weights, data, code and checkpoints.source ↗
2024-01-01Mistral AIMixtral of ExpertspaperarXiv:2401.04088. 2,107 citations.source ↗
2023-10-01Mistral AIMistral 7BpaperarXiv:2310.06825. 3,851 citations; the 7B era it started still drives Mistral's download volume three years later.source ↗
2023-07-17Meta AI — FAIR + MSLDINOv2open-weightsfacebook/dinov2-large, Apache-2.0. Paper arXiv:2304.07193, 10,184 citations.source ↗
2023-04-10Meta AI — FAIR + MSLSegment Anything (SAM)open-weightsfacebook/sam-vit-huge created 2023-04-10 under Apache-2.0 and ungated — the permissive baseline Meta's later vision releases moved away from. Paper arXiv:2304.02643, 15,343 citations.source ↗
2023-03-01OpenAIGPT-4 Technical ReportpaperarXiv:2303.08774. 26,936 citations as of 2026-09-04.source ↗
2023-02-01Meta AI — FAIR + MSLLLaMA: Open and Efficient Foundation Language ModelspaperarXiv:2302.13971. The paper that started the open-weight era; 21,439 citations as of 2026-09-04.source ↗

What the field builds on — 30 landmark papers

Older papers have had longer to accumulate, so this ranks foundations, not current standing.

PaperLabYearCitationsReadSource
Attention Is All You NeedGoogle DeepMind2017191,1162026-09-04arXiv ↗
BERTGoogle DeepMind2019120,3722026-09-04arXiv ↗
Vision TransformerGoogle DeepMind202069,0242026-09-04arXiv ↗
GPT-3OpenAI202062,8102026-09-04arXiv ↗
PyTorchMeta AI — FAIR + MSL201955,0522026-09-04arXiv ↗
CLIPOpenAI202155,0282026-09-04arXiv ↗
Mask R-CNNMeta AI — FAIR + MSL201732,9292026-09-04arXiv ↗
RoBERTaMeta AI — FAIR + MSL201931,1502026-09-04arXiv ↗
GPT-4 Technical ReportOpenAI202326,9362026-09-04arXiv ↗
InstructGPT (RLHF)OpenAI202223,8072026-09-04arXiv ↗
LoRAMicrosoft Research202122,5462026-09-04arXiv ↗
Chain-of-Thought promptingGoogle DeepMind202221,5482026-09-04arXiv ↗
LLaMAMeta AI — FAIR + MSL202321,4392026-09-04arXiv ↗
The Llama 3 Herd of ModelsMeta AI — FAIR + MSL202418,2862026-09-04arXiv ↗
Segment AnythingMeta AI — FAIR + MSL202315,3432026-09-04arXiv ↗
DINOv2Meta AI — FAIR + MSL202310,1842026-09-04arXiv ↗
wav2vec 2.0Meta AI — FAIR + MSL20209,1602026-09-04arXiv ↗
WhisperOpenAI20228,3112026-09-04arXiv ↗
DeepSeek-R1DeepSeek20255,5952026-09-04arXiv ↗
Qwen2.5Qwen (Alibaba Cloud)20244,8752026-09-04arXiv ↗
Training a Helpful and Harmless AssistantAnthropic20224,3352026-09-04arXiv ↗
DeepSeek-V3DeepSeek20243,9792026-09-04arXiv ↗
Gemini 1.5Google DeepMind20243,9432026-09-04arXiv ↗
Mistral 7BMistral AI20233,8512026-09-04arXiv ↗
Constitutional AIAnthropic20223,6172026-09-04arXiv ↗
Phi-3Microsoft Research20242,4522026-09-04arXiv ↗
Mixtral of ExpertsMistral AI20242,1072026-09-04arXiv ↗
GemmaGoogle DeepMind20241,2192026-09-04arXiv ↗
Kimi k1.5Moonshot AI20251,0622026-09-04arXiv ↗
OLMoAllen Institute for AI (Ai2)20247272026-09-04arXiv ↗