Make-A-Video
Text-to-video before it was normal
Latest: Make-A-Video (Sep 2022); superseded by Emu Video, then Movie Gen
Meta's pioneering text-to-video system (September 2022), among the first to generate video from prompts without any paired text-video training data — it learned appearance from text-image pairs and motion from unlabeled video. A research showcase only; never released as a product or open model, but it kicked off the T2V race.
Why it matters
One of the first text-to-video systems, and notable for needing no paired text-video data: it learned appearance from text-image pairs and motion from unlabeled video. It was never released, but it kicked off the text-to-video race that led to Emu Video, Movie Gen and Meta's Muse Video.
Facts
- Announced days before Google's Imagen Video — September 2022 was the month text-to-video arrived; the flying-superhero-dog clip became its calling card.
Try it yourself
Official showcase site (samples) ↗ Read the paper ↗