Aligning Text, Images & Videos is still a struggle
AI has made extraordinary progress in multimodal learning, particularly in aligning text and images into shared embedding spaces. Models like CLIP and GPT-4V(ision) are prime examples of this success, showing how architectures can bridge the gap between vision and language. But when we add videos to the mix, the landscape changes dramatically. Despite existing architectures and innovations, aligning text, images, and videos into a common embedding space is still imperfect and incomplete. ...