metaai·lightalo unofficial · independent
Universe / On-Device & Silicon / MobileLLM-Flash
On-Device & Silicon · 2026

MobileLLM-Flash

Architecture search with real phone latency in the loop

closed / product 350M / 650M / 1.4B params

Latest: MobileLLM-Flash 350M/650M/1.4B (paper Mar 16, 2026; ACL Industry Track 2026)

Latest MobileLLM generation (350M/650M/1.4B) designed via hardware-in-the-loop architecture search under real mobile latency constraints, with attention skipping for long-context acceleration up to 8k. Paper (Mar 2026) accepted to ACL 2026 Industry Track. Weights had not been publicly released as of Sept 2026.

Why it matters

Designs the model around the phone rather than the benchmark: hardware-in-the-loop architecture search under real mobile latency constraints, plus attention skipping to accelerate contexts up to 8k. Accepted to the ACL 2026 Industry Track; as of September 2026 the weights had not been released, so only the paper is available.

Facts

Try it yourself

Lineage

Descends fromMobileLLM

See the whole family tree →

Sources

More in On-Device & Silicon

MobileLLM-R1 / R1.5MobileLLM-ProGPU superclusters: 350K H100s → Prometheus & HyperionQuantized Llama 3.2Meta LLM CompilerMTIA

Read the On-Device & Silicon story on the sky →

✦ Open on the map Explore On-Device & Silicon Quiz me