metaai·lightalo unofficial · independent
Universe / Language & LLMs / OPT-175B
Language & LLMs · 2022

OPT-175B

The 175B model that published its 3 a.m. crash logs

open source superseded 125M / 350M / 1.3B / 2.7B / 6.7B / 13B / 30B / 66B / 175B params

Latest: OPT 125M-175B suite (May 2022)

Open Pre-trained Transformers (May 3, 2022): a GPT-3-scale 175B model whose weights were shared with researchers, alongside the full codebase — and, famously, the raw 114-page logbook chronicling every crash, loss spike, and hardware failure of the training run. A radical act of transparency at frontier scale.

Why it matters

OPT-175B was the first GPT-3-scale model whose weights reached researchers, but its lasting contribution is transparency: Meta published the full codebase and a 114-page logbook of crashes, loss spikes and 35+ restarts across 992 A100s, at roughly one-seventh of GPT-3's estimated carbon footprint. It set the template LLaMA followed.

Facts

Try it yourself

Lineage

Descends fromfairseq
Led toLLaMAGalactica

See the whole family tree →

Sources

More in Language & LLMs

Llama 2Code LlamaLlama 3 / 3.1 / 3.2 / 3.3Large Concept ModelsByte Latent TransformerLlama Stack

Read the Language & LLMs story on the sky →

✦ Open on the map Explore Language & LLMs Quiz me