audio2photoreal
Record your voice, watch a photoreal avatar gesture along
Latest: CVPR 2024 (arXiv Jan 2024); code, dataset and Gradio demo live
Open code and dataset for synthesizing full-bodied photorealistic Codec Avatars — face, body, and hands — that gesture naturally from conversational audio alone. Combines vector-quantization sample diversity with diffusion for high-frequency motion detail. Released January 2024 by Meta Reality Labs researchers with UC Berkeley; published at CVPR 2024.
Why it matters
Synthesizes full-body photorealistic avatars, including face and hands, that gesture naturally from conversational audio alone, combining vector-quantized sample diversity with diffusion for high-frequency motion. Published at CVPR 2024 with code, a four-participant conversational dataset and a Colab demo under CC-BY-NC 4.0, it is a public window into Codec Avatar animation.
Facts
- Ships a Gradio demo where you record your own voice and watch a photoreal avatar gesture along.
- Built on dyadic (two-person) conversation data so avatars react to interpersonal dynamics, not just their own speech.
Try it yourself
Run the demo in Colab ↗ Code and dataset on GitHub ↗ Read the paper ↗
Lineage
Sources
arXiv ↗GitHub · audio2photoreal ↗