OPT-175B
The 175B model that published its 3 a.m. crash logs
Latest: OPT 125M-175B suite (May 2022)
Open Pre-trained Transformers (May 3, 2022): a GPT-3-scale 175B model whose weights were shared with researchers, alongside the full codebase — and, famously, the raw 114-page logbook chronicling every crash, loss spike, and hardware failure of the training run. A radical act of transparency at frontier scale.
Why it matters
OPT-175B was the first GPT-3-scale model whose weights reached researchers, but its lasting contribution is transparency: Meta published the full codebase and a 114-page logbook of crashes, loss spikes and 35+ restarts across 992 A100s, at roughly one-seventh of GPT-3's estimated carbon footprint. It set the template LLaMA followed.
Facts
- The first time a lab published its 3 a.m. on-call notes from training a 175-billion-parameter model — warts, crashes, and all.
- The public 'Chronicles' logbook documents 35+ manual restarts and cascades of GPU failures over ~2 months on 992 80GB A100s.
- Meta estimated the carbon footprint at ~75 tons CO2e versus an estimated 500 for GPT-3 — a 1/7th footprint at the same scale.
- The logbook records 35 training restarts and over 100 hosts cycled due to hardware failures — machines died almost daily on the 992 A100 GPUs.
- 19 authors (Zhang et al.).
Try it yourself
OPT-1.3B on Hugging Face ↗ Read the OPT-175B training logbook (PDF) ↗ Read the paper ↗
Lineage
Sources
arXiv ↗GitHub · metaseq · metaseq ↗GitHub · metaseq · chronicles ↗GitHub · metaseq · OPT175B_Logbook.pdf ↗