metaai·lightalo unofficial · independent
Universe / Language & LLMs / Multi-token prediction
Language & LLMs · 2024

Multi-token prediction

Several tokens per step: stronger code models, up to 3x faster

open source 7B (released code models) params

Latest: Paper Apr 2024; 7B code models released Jul 2024

'Better & Faster Large Language Models via Multi-token Prediction' (April 2024): training models to predict several future tokens at once via parallel output heads yields better sample efficiency and stronger code models — 13B models solved 12-17% more HumanEval/MBPP problems — while enabling up to 3x faster inference via self-speculative decoding. 7B weights released for research.

Why it matters

Multi-token prediction showed that training a model to predict several future tokens at once — with zero training overhead — yields better sample efficiency and stronger code models, with 13B models solving 12-17% more HumanEval/MBPP problems, plus up to 3x faster self-speculative decoding. The idea flowed straight into industry speculative-decoding heads.

Facts

Try it yourself

Sources

More in Language & LLMs

Llama 3 / 3.1 / 3.2 / 3.3Large Concept ModelsByte Latent TransformerLlama StackCoconutMemory Layers at Scale

Read the Language & LLMs story on the sky →

✦ Open on the map Explore Language & LLMs Quiz me