metaai·lightalo unofficial · independent

What this site got wrong

This field guide claims that every fact on it is checked against a primary source. That claim is only worth something if the failures are published too. Here are 24 of them: what a page said, why it was wrong, what it says now, and how it was caught.

Corrections from 2026-09-03 to 2026-09-04. Newest first. Only wrong claims are listed — layout fixes, rewrites and new features are not corrections.

/frontier/

It said

Every dimension shown under one line reading "Researched and adversarially fact-checked 2026-09-04".

It says now

Each dimension prints the date its own values were read, and says how its rows are ordered. The plain-table version at /frontier/data/ carries the same stamp.

Why that was wrong. One dimension was not. The standards-body membership table was read on 9 December 2025, nine months earlier, and each dimension's own as_of date was collected and then rendered nowhere. A page whose central promise is that every figure carries the date it was read was printing one date over all of them.

How it was caught. An adversarial review of the comparison page, which asked where the collected as_of fields were displayed.

/ (guided tour)

It said

"Every star is something Meta's AI labs actually shipped — 119 of them" — narrated over whatever the sky was showing.

It says now

A stop that narrates the whole catalog clears the filters first, and both counts are read from the catalog at render time rather than written into the prose.

Why that was wrong. The tour cleared a filter only when its own target star was hidden. Its opening and closing stops talk about the whole catalog, so a search left in the box could leave that sentence describing 119 projects above a sky showing twenty-two.

How it was caught. An adversarial review that started the tour with a search still active.

/play/

It said

Three clues with a block of ▮▮▮ over ordinary prose: "Muse Spark line", "use standard text-based", "speech-native language models".

It says now

Only the first word of a run may begin lowercase, and only when it is hyphenated. The three false positives are gone and the real expansions are still masked.

Why that was wrong. The masker hides a run of capitalised words whose initials spell the answer's acronym, which is how OPT stopped being given away by "Open Pre-trained Transformers". Its pattern accepted a lowercase initial too, so any three consecutive words with the right first letters were blacked out — including, on the Meta Superintelligence Labs clue, the name of a different project.

How it was caught. Re-running the masker over every clue in the catalog and reading what it hid.

/

It said

94 of 119 projects open source

It says now

Every label says "open release" or "released openly", and the About page's badge legend sets out the range from BSD and Apache 2.0 through Llama 4's 700-million-user clause to research-only weights.

Why that was wrong. The catalog's flag means weights or code were published, which is a weaker claim than an open-source licence. Of 50 entries whose licence could be resolved, 36 were not under an open licence and 14 explicitly forbid commercial use. The LLaMA card carried an "open source" badge directly above its own description reading "released to researchers under a non-commercial license".

How it was caught. An adversarial review that checked the badge against each project's actual licence on Hugging Face.

/p/llama-stack/

It said

Latest: v1.3.1 (1 Sep 2026)

It says now

The card gives the last release published under the Llama name and states that the repository has since been renamed. pulse.py records renames, and the same note appears wherever a moved repository's figures are shown.

Why that was wrong. That release belongs to a different repository. GitHub silently follows renames, and llamastack/llama-stack now answers as ogx-ai/ogx. The API returned that repository's release and I published it as llama-stack's own.

How it was caught. The frontier research pass warned about this exact API behaviour; a later check of the collected data caught that I had walked into it anyway.

/p/msl/

It said

A ~6,500-person Applied AI engineering unit stood up in March 2026 under a named executive.

It says now

Removed. The card states the sourced August 2025 four-division structure instead.

Why that was wrong. The claim appears in none of the card's four cited sources. Two research passes had also named two different executives, which was the signal that neither had a source.

How it was caught. A verifier that opened all four cited sources rather than trusting the summary.

/

It said

119 projects · 13 years

It says now

Everything derived from the catalog says 2016 to 2026. The Timeline, whose scope is events rather than shipped projects, still begins in 2013.

Why that was wrong. The catalog's earliest entry is PyTorch in 2016. FAIR was founded in 2013, but no project in the star map predates 2016. The claim appeared in the hero, the time-lapse button, an era filter, the insights year axis and three meta descriptions.

How it was caught. Noticing that the time-lapse's own data-derived label disagreed with the button beside it.

/#c=vision and every constellation story

It said

…while DINOv2 proves frozen label-free features beat supervised pipelines. 0. 2 m. 1 follows in March 2026.

It says now

The splitter slices the original text at a sentence boundary instead of reassembling it. All twelve stories were verified character-exact.

Why that was wrong. The story renderer collected regular-expression matches and joined them, silently discarding every character the pattern failed to cover. Nine of twelve stories lost text; the Vision story lost 407 of its 1,020 characters and printed the fragment above as if it were prose.

How it was caught. An adversarial review of the galaxy, which quoted the garbled sentence back.

/lab/cutout/

It said

Masks agree with the full-precision model at 0.999 IoU on our test photo.

It says now

All three figures are given, alongside the 0.997 embedding cosine.

Why that was wrong. 0.999 was the best of the model's three candidate masks. The other two agreed at 0.928 and 0.973.

How it was caught. Re-reading my own verification output rather than the sentence I had written from it.

/lab/cutout/

It said

~10M parameters in the encoder

It says now

Both counts are given, measured from the shipped files.

Why that was wrong. The EfficientSAM paper's 10M figure is the whole model. Counting the two ONNX graphs the page actually downloads gives 6.16 million in the encoder and 4.06 million in the decoder.

How it was caught. An adversarial review of the Lab demos, which counted the parameters in the served model.

/lab/cutout/

It said

You can switch off Wi-Fi after this page loads and it still works.

It says now

The page says the files download once, when you pick your first photo, and that it works offline after that. The Lab index carries the same correction and adds that a reload still needs the network.

Why that was wrong. The model files download when the first photo is chosen, not on page load, so cutting the network at load breaks the demo.

How it was caught. Reading the code path against the sentence.

/p/hydra/ and five other project pages

It said

Latest: v1.3.6 (August 2024)

It says now

All corrected against the GitHub releases API, and the site now reads each repository's most recent release live every week, shown only for a repository a single project owns.

Why that was wrong. That release shipped in August 2026. Hydra's date was two years early and nevergrad's a year early, making two actively maintained projects look abandoned. ExecuTorch, llama-stack and the torch domain libraries were several releases behind, and Muse Spark's card said 1.2 while its own body copy said 1.3.

How it was caught. An adversarial review that checked every "Latest" field against the releases API.

/p/muse-image/

It said

Every output carries the Content Seal invisible watermark.

It says now

Scoped to what the source says.

Why that was wrong. The cited source scopes the watermark to the Meta AI app and meta.ai, not to every output of the model.

How it was caught. An adversarial review that opened the cited source.

/timeline/

It said

the first time a frontier lab put commercial-use weights on the table

It says now

Described as the most capable commercially usable weights yet, and the first Llama you could build a business on.

Why that was wrong. The site's own catalog lists earlier commercially usable open weights from Meta, so its own data contradicted the superlative.

How it was caught. An adversarial review that checked the claim against the catalog two constellations away.

/timeline/

It said

A 7B self-supervised vision model distilled from 1.7 billion images

It says now

Both stages are described separately.

Why that was wrong. You distil from a model, not from a dataset. DINOv3's 7B model was trained on 1.7 billion images; the smaller variants were distilled from the 7B.

How it was caught. An adversarial review, which noted the project's own card had it right.

/p/muse-spark/ and five others

It said

Six lineage arrows asserting that a later project gave rise to an earlier one.

It says now

Four were reversed — Ego4D was collected with Aria hardware, UMA is built on the fairchem library, the sEMG paper builds on the emg2 datasets — and two were removed, because "now runs on" is not descent. The build now fails if any parent postdates its child.

Why that was wrong. Lineage means descent, so a parent cannot postdate its child. A 2026 model was shown as having led to a 2025 app.

How it was caught. An adversarial review; the check that prevents recurrence was added afterwards.

/p/brain-decoding/

It said

over 1,500 candidates

It says now

Matches the paper.

Why that was wrong. The paper the card cites says more than 1,000, and the site's own constellation story said 1,000 too.

How it was caught. An adversarial review that opened the cited paper.

/

It said

1.2B+ Llama downloads, linked to the Llama 3 star

It says now

Links to the Llama 4 star, which carries the figure, and the tile says the number is from April 2025.

Why that was wrong. The site promises that every headline number links to the star where it is sourced and dated. That star carries neither the figure nor its date, and the figure was seventeen months old with no date shown.

How it was caught. An adversarial review that followed the site's own promise.

/insights/

It said

A model-size chart on which six closed or never-released models were invisible.

It says now

The rule sets stroke width only. All six are visible.

Why that was wrong. A CSS rule overrode each circle's own stroke attribute, so every hollow marker was drawn in the panel colour on the panel background — including Llama 4 Behemoth, the largest model on the chart, whose label pointed at nothing.

How it was caught. An adversarial review of the non-galaxy pages.

/insights/

It said

A most-starred-repositories chart whose bars were not comparable.

It says now

A fixed column on both viewports; every track measures identically.

Why that was wrong. The value column was auto-sized, so a row reading "32k · archived" left a track up to 36 pixels wider than one reading "103k · last push 2026-09-03". The visual ranking disagreed with the numbers printed beside it.

How it was caught. An adversarial review that measured every bar in pixels.

/about/

It said

80 lineage links

It says now

Both read the count from the catalog at load, so they cannot drift again.

Why that was wrong. Six lineage edges had been corrected away, leaving 78. The stat tile and the prose sentence both still said 80.

How it was caught. An adversarial review that compared the page against /insights/.

/p/hydra/, /p/prophet/, /p/pytorch3d/ and twelve others

It said

Frozen GitHub star counts in the prose, beside a live star badge showing a different number.

It says now

The build strips the clause generically, so a future research pass cannot reintroduce one.

Why that was wrong. The site began reading star counts live every week, which made the fifteen counts written into the prose both redundant and visibly wrong — Hydra's said 9k beside a live 10.6k.

How it was caught. An adversarial review that read the badge and the prose in the same viewport.

/play/

It said

A quiz round that could not be won.

It says now

Distractors are drawn from outside the answer's lineage; any word of the answer's name carried by exactly one option is masked; and a clue can no longer expand the answer's acronym.

Why that was wrong. The clue "Open-source safety tooling for open models" is equally true of Purple Llama and of CyberSecEval and LlamaFirewall, which are components of it. Three of the four options answered the clue and two were parts of the right one. Separately, the masker exempted generic-looking words like "image" and "code", so a clue could print a word that identified the answer uniquely, and one clue opened by spelling out the answer's own acronym.

How it was caught. An adversarial review that drove 500 rounds, and my own check of 80 more.

/data/projects.json

It said

33 factual errors across the catalog, including a wrongly attributed Science cover, an inverted attribution on V-JEPA and several stale version numbers.

It says now

Corrected, and every subsequent research pass is followed by an adversarial reader whose job is to break the claims rather than confirm them.

Why that was wrong. The first research sweep wrote them and a single reader confirmed them.

How it was caught. A six-agent adversarial fact-check of the whole catalog.

Most of these were caught the same way: a second reader whose job was to break the claim rather than confirm it, opening the cited source instead of trusting the summary of it. Several were errors in prose I had written from a source I had read correctly, which is the failure mode this page exists to keep visible. If you find another, please send it — it will end up on this page.