Learning aligned image-text representations in a shared embedding space is crucial for vision-language tasks. While contrastive loss is highly effective for representation learning, it typically binds one image to one text caption, limiting the ability to capture individual properties from the text. Diagnostic datasets and benchmarks are vital for developing and understanding multi-property vision-language representations but are currently underexplored. We introduce Prop-Clevr, a diagnostic benchmark designed to evaluate the ability of models to capture multiple properties in vision-language representations. We propose a novel training objective, multi-target contrastive loss, which aligns image embeddings with multiple properties in the corresponding text. Empirical experiments on Prop-Clevr in image classification and text-image retrieval tasks demonstrate the effectiveness of our training objective in producing high-quality embeddings that capture individual properties in the images and texts. We provide the assets, codes, and tools that allow for high customization of Prop-Clevr, enabling the creation of new benchmarks to study and diagnose vision-language and related tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-target Contrastive Objective for Learning Property-Aware Vision-Language Representation

  • Dieu-Hien Nguyen,
  • Nguyen-Khang Le,
  • Le Minh Nguyen

摘要

Learning aligned image-text representations in a shared embedding space is crucial for vision-language tasks. While contrastive loss is highly effective for representation learning, it typically binds one image to one text caption, limiting the ability to capture individual properties from the text. Diagnostic datasets and benchmarks are vital for developing and understanding multi-property vision-language representations but are currently underexplored. We introduce Prop-Clevr, a diagnostic benchmark designed to evaluate the ability of models to capture multiple properties in vision-language representations. We propose a novel training objective, multi-target contrastive loss, which aligns image embeddings with multiple properties in the corresponding text. Empirical experiments on Prop-Clevr in image classification and text-image retrieval tasks demonstrate the effectiveness of our training objective in producing high-quality embeddings that capture individual properties in the images and texts. We provide the assets, codes, and tools that allow for high customization of Prop-Clevr, enabling the creation of new benchmarks to study and diagnose vision-language and related tasks.