With the advent of the digital age, multimodal data, including images, text, audio, and video, has become a key asset for both enterprises and societal development. There are complex relationships between these different modalities, yet traditional retrieval methods often focus on a single modality, neglecting the complementary and enhancing effects between modalities. This paper proposes a Conflict Mitigation Cross-Modal Search (CM-CMS) technology that leverages deep learning to achieve effective alignment and retrieval of cross-modal data. The core of CM-CMS lies in the effective alignment and retrieval between cross-modal data through deep learning techniques. It first converts images and text into deep feature representations through automatic feature extraction, and then separates modal features into common and unique hash codes through a modality decoupling mechanism, reducing inconsistencies between modalities. The experimental section targets two datasets, MIRFlickr25K and NUS-WIDE10.5K, and validates the effectiveness of CM-CMS in multimodal data retrieval tasks through comparative experimental results with six cutting-edge methods, as well as an analysis of the sensitivity to hyperparameters. The experimental results show that CM-CMS achieved the best mean average precision (MAP) scores in all cross-modal retrieval tasks, with a significant performance improvement compared to existing technologies. Furthermore, the visualization results of common and specific subspaces demonstrate the effectiveness of the CM-CMS method in decoupling. In summary, this study not only improves the accuracy of multimodal data retrieval but also enhances the model’s robustness to data inconsistencies, providing a new technical solution for the field of multimodal data retrieval, showing significant performance improvements, and offering new directions for future research.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CM-CMS: Conflict Mitigation Cross-Modal Search

  • Liu Xin,
  • Qiu Kaiyi,
  • Ma Hongbo,
  • He Wei,
  • Liu Jie

摘要

With the advent of the digital age, multimodal data, including images, text, audio, and video, has become a key asset for both enterprises and societal development. There are complex relationships between these different modalities, yet traditional retrieval methods often focus on a single modality, neglecting the complementary and enhancing effects between modalities. This paper proposes a Conflict Mitigation Cross-Modal Search (CM-CMS) technology that leverages deep learning to achieve effective alignment and retrieval of cross-modal data. The core of CM-CMS lies in the effective alignment and retrieval between cross-modal data through deep learning techniques. It first converts images and text into deep feature representations through automatic feature extraction, and then separates modal features into common and unique hash codes through a modality decoupling mechanism, reducing inconsistencies between modalities. The experimental section targets two datasets, MIRFlickr25K and NUS-WIDE10.5K, and validates the effectiveness of CM-CMS in multimodal data retrieval tasks through comparative experimental results with six cutting-edge methods, as well as an analysis of the sensitivity to hyperparameters. The experimental results show that CM-CMS achieved the best mean average precision (MAP) scores in all cross-modal retrieval tasks, with a significant performance improvement compared to existing technologies. Furthermore, the visualization results of common and specific subspaces demonstrate the effectiveness of the CM-CMS method in decoupling. In summary, this study not only improves the accuracy of multimodal data retrieval but also enhances the model’s robustness to data inconsistencies, providing a new technical solution for the field of multimodal data retrieval, showing significant performance improvements, and offering new directions for future research.