metaai·lightalo unofficial · independent
Universe / On-Device & Silicon / Quantized Llama 3.2
On-Device & Silicon · 2024

Quantized Llama 3.2

Quantized Llama 3.2 (1B/3B) + ExecuTorch

Llama 3.2 on a phone — 56% smaller, up to 4x faster

open source 1B / 3B params

Latest: Llama 3.2 1B/3B Instruct QLoRA_INT4_EO8 and SpinQuant_INT4_EO8 (Oct 24, 2024)

Meta's first official quantized Llama releases (Oct 24, 2024): Llama 3.2 1B/3B Instruct in two flavors — Quantization-Aware Training with LoRA adaptors for accuracy, and SpinQuant post-training quantization for portability. 2-4x faster with 56% smaller size and 41% less memory, running on phones via PyTorch's ExecuTorch on Qualcomm, MediaTek, and Arm. Llama 3.2 Community License — commercial use permitted.

Why it matters

Meta's first official quantized Llama releases made small-model on-device deployment a supported path rather than a community hack: QAT with LoRA adaptors for accuracy and SpinQuant for portability, 2-4x faster with 56% smaller size and 41% less memory, running via ExecuTorch on Qualcomm, MediaTek and Arm under the commercial Llama 3.2 license.

Facts

Try it yourself

Lineage

Descends fromLlama 3 / 3.1 / 3.2 / 3.3ExecuTorch

See the whole family tree →

Sources

More in On-Device & Silicon

MobileLLMGPU superclusters: 350K H100s → Prometheus & HyperionMeta LLM CompilerMTIAMobileLLM-R1 / R1.5MobileLLM-Pro

Read the On-Device & Silicon story on the sky →

✦ Open on the map Explore On-Device & Silicon Quiz me