metaai·lightalo unofficial · independent
Universe / Language & LLMs / RoBERTa
Language & LLMs · 2019

RoBERTa

BERT, done right

open source superseded 125M (base) / 355M (large) params

Latest: RoBERTa base/large (Jul 2019)

'A Robustly Optimized BERT Pretraining Approach' (July 2019): Meta showed BERT was severely undertrained, and that longer training on more data (160GB of text) with dynamic masking — no architecture change — topped the GLUE leaderboard. For years it was the default encoder for real-world NLP systems.

Why it matters

RoBERTa showed BERT was badly undertrained: same architecture, more data, longer training and dynamic masking topped GLUE. It is the landmark argument that recipes matter as much as architectures, and years later roberta-base still logs nearly ten million monthly Hugging Face downloads as a default encoder for real-world NLP.

Facts

Try it yourself

Lineage

Descends fromfairseq
Led toXLM-R

See the whole family tree →

Sources

More in Language & LLMs

BARTBlenderBotOPT-175BGalacticaLLaMALlama 2

Read the Language & LLMs story on the sky →

✦ Open on the map Explore Language & LLMs Quiz me