metaai·lightalo unofficial · independent
Universe / On-Device & Silicon / MobileLLM
On-Device & Silicon · 2024

MobileLLM

Proof that sub-billion models could think

open source 125M / 350M / 600M / 1B / 1.5B params

Latest: MobileLLM-1.5B + ParetoQ quantized variants (1/1.58/2/3/4-bit); 125M/350M/600M/1B/1.5B line completed Nov 2024

FAIR's sub-billion-parameter architecture study proving that deep-and-thin transformers with embedding sharing, grouped-query attention, SwiGLU, and block-wise layer sharing beat wide-shallow designs on-device. Published at ICML 2024. Checkpoints from 125M to 1.5B released with full weights and training code — all non-commercial (CC-BY-NC 4.0 weights, FAIR NC code).

Why it matters

Reset how sub-billion language models are designed: deep-and-thin transformers with embedding sharing, grouped-query attention and block-wise layer sharing beat wide-shallow designs at equal size. Published at ICML 2024 with full weights and training code, it became the reference architecture for on-device models, including Meta's R1, Pro and Flash lines.

Facts

Try it yourself

Lineage

Led toMobileLLM-R1 / R1.5MobileLLM-ProMobileLLM-Flash

See the whole family tree →

Sources

More in On-Device & Silicon

GPU superclusters: 350K H100s → Prometheus & HyperionQuantized Llama 3.2Meta LLM CompilerMTIA

Read the On-Device & Silicon story on the sky →

✦ Open on the map Explore On-Device & Silicon Quiz me