metaai·lightalo unofficial · independent
Universe / Vision / ImageBind
Vision · 2023

ImageBind

One embedding space to bind six senses

open source ImageBind-Huge (ViT-H/14 image encoder, ~630M) params

Latest: ImageBind (May 2023)

The first model to bind six modalities — images, text, audio, depth, thermal, and IMU motion data — into one joint embedding space, trained only on image-paired data (May 2023). Enables cross-modal retrieval and arithmetic, like finding sounds that match an image, without any dataset pairing all modalities together.

Why it matters

ImageBind was the first model to bind six modalities — images, text, audio, depth, thermal and IMU — into one embedding space using only image-paired data, so cross-modal retrieval and embedding arithmetic work between senses that were never paired in training. It prefigured the natively multimodal models Meta later shipped.

Facts

Try it yourself

Sources

More in Vision

Segment AnythingDINOv2CoTrackerSAM 2ChameleonSapiens

Read the Vision story on the sky →

✦ Open on the map Explore Vision Quiz me