metaai·lightalo unofficial · independent
Universe / Language & LLMs / Memory Layers at Scale
Language & LLMs · 2024

Memory Layers at Scale

Capacity without FLOPs: a lookup table the model learns to consult

open source base models up to 8B, memory pools up to 128B parameters params

Latest: Paper + code (Dec 2024)

FAIR work scaling trainable key-value memory layers to add capacity without adding FLOPs: a 1.3B model with memory matched models trained on 2-4x more compute, and beat MoE models at equal budget on factual tasks. Part of the same December 2024 FAIR wave as LCM and BLT — three simultaneous bets against the standard dense transformer.

Why it matters

Memory layers add trainable key-value lookup capacity without adding FLOPs: a 1.3B model with memory matched dense models trained on 2-4x more compute and beat MoE at equal budget on factual tasks, scaling to 128B memory parameters. It is one of FAIR's three December 2024 bets against the standard dense transformer.

Facts

Try it yourself

Sources

More in Language & LLMs

Llama 3 / 3.1 / 3.2 / 3.3Large Concept ModelsByte Latent TransformerLlama StackCoconutMulti-token prediction

Read the Language & LLMs story on the sky →

✦ Open on the map Explore Language & LLMs Quiz me