MobileLLM-R1 / R1.5
Sub-billion reasoners matching Qwen3-0.6B on a seventh of the tokens
Latest: MobileLLM-R1.5 140M/360M/950M (Nov 24, 2025); R1 released Sept 12, 2025
Open-recipe sub-billion reasoning models (140M/360M/950M, base + final) for math, code, and science. R1-950M matches or beats Qwen3-0.6B on MATH, MMLU, and LiveCodeBench despite fewer than 5T total training tokens vs Qwen3's 36T. R1.5 (Nov 2025) adds on-policy knowledge distillation. FAIR Noncommercial Research License — must be labeled non-commercial.
Why it matters
Showed reasoning ability need not wait for billion-parameter scale: the 950M model matches or beats Qwen3-0.6B on MATH, MMLU and LiveCodeBench despite fewer than 5T training tokens versus Qwen3's 36T. R1.5 in November 2025 added on-policy knowledge distillation, and the fully open recipe lets others reproduce it.
Facts
- R1.5's on-policy KD added 10-35 points on hard reasoning benchmarks: R1.5-950M scores 39.9 on AIME'24 vs Qwen3-0.6B's 11.3, and the 360M jumped from 28.4 to 63.4 on MATH.
- Fully open training recipe: data sources, code, and all checkpoints published.
Try it yourself
MobileLLM-R1.5-950M on Hugging Face ↗ R1 / R1.5 collection on Hugging Face ↗ Read the paper ↗
Lineage
Sources
arXiv ↗GitHub · MobileLLM-R1 ↗Hugging Face · MobileLLM R1 950M ↗Hugging Face · MobileLLM R1.5 950M ↗Hugging Face · mobilellm r1 68c4597b104fac45f ↗