<p>Image–text matching (ITM) is a cornerstone of vision-language understanding, enabling precise correspondences between visual and textual content. However, existing ITM methods often struggle with the precise selection of local image regions and neglect the hierarchical semantic structure of text. To address these limitations, we propose a multi-level semantic consistency alignment approach (MSCA). Our method integrates a visual anchor alignment module and a semantic layering module. The visual anchor alignment module extracts the most representative image regions as visual anchors to solve the issue of imprecise local information selection. The semantic layering module, on the other hand, captures the hierarchical structure of the text and divides the semantic information into high-level core semantics, mid-level auxiliary semantics, and low-level boundary semantics, so as to comprehensively enhance the semantic integrity of the text features. Extensive experiments on the Flickr30K and MS-COCO datasets demonstrate that our method significantly outperforms existing approaches, especially in complex semantic reasoning and detail matching. The proposed method demonstrates a wide range of applications and research potential. Our code is available at <a href="https://github.com/zhuliqi0309/MSCA.git">https://github.com/zhuliqi0309/MSCA.git</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing image–text matching through multi-level semantic consistency alignment

  • Liqi Zhu,
  • Dezhi Han,
  • Xiang Shen,
  • Chongqing Chen,
  • Kuan-Ching Li

摘要

Image–text matching (ITM) is a cornerstone of vision-language understanding, enabling precise correspondences between visual and textual content. However, existing ITM methods often struggle with the precise selection of local image regions and neglect the hierarchical semantic structure of text. To address these limitations, we propose a multi-level semantic consistency alignment approach (MSCA). Our method integrates a visual anchor alignment module and a semantic layering module. The visual anchor alignment module extracts the most representative image regions as visual anchors to solve the issue of imprecise local information selection. The semantic layering module, on the other hand, captures the hierarchical structure of the text and divides the semantic information into high-level core semantics, mid-level auxiliary semantics, and low-level boundary semantics, so as to comprehensively enhance the semantic integrity of the text features. Extensive experiments on the Flickr30K and MS-COCO datasets demonstrate that our method significantly outperforms existing approaches, especially in complex semantic reasoning and detail matching. The proposed method demonstrates a wide range of applications and research potential. Our code is available at https://github.com/zhuliqi0309/MSCA.git.