metaai·lightalo unofficial · independent
Universe / Vision / Chameleon
Vision · 2024

Chameleon

Images and text in one token stream, trained from scratch

open source 7B / 34B (hosted on Hugging Face as chameleon-30b) params

Latest: Chameleon 7B/34B (Jun 2024 weight release; image-understanding only)

FAIR's early-fusion mixed-modal foundation model (May 2024): a single token-based transformer trained from scratch on interleaved images and text, able to understand and generate arbitrary interleavings of both. Meta released 7B and 34B weights under a research license in June 2024 — a precursor to natively multimodal product models.

Why it matters

Chameleon was Meta's first early-fusion model: one token-based transformer trained from scratch on interleaved images and text, able to understand and generate either, with the 34B version beating much larger models on interleaved-generation human evals. It is the research precursor of the natively multimodal Llama 4 and Muse Spark lines.

Facts

Try it yourself

Sources

More in Vision

SAM 2SapiensVideo SealSegment AnythingSAM 3DINOv2

Read the Vision story on the sky →

✦ Open on the map Explore Vision Quiz me