Unsupervised Text-to-Image Person Re-identification Framework with Enhanced Focus Attention and Multi-dimensional Feature Interaction Attention
摘要
Text-to-image re-identification (Re-ID) is the process of using textual descriptions to assist in image recognition and matching. The main challenge of this technology is to ensure semantic consistency between textual descriptions and image content, which is crucial for the accuracy of matching and also requires a large number of high-quality database annotations. In this paper, we propose an enhanced focus attention module (EFAM) and a multi-dimensional feature interaction attention module (MFIAM) based on an unsupervised text-to-image re-identification framework (CLIP-ReID), aiming to address the issue of incorrect image-text semantic matching due to insufficient image feature discrimination. EFAM obtains feature vertex weights through a graph attention layer to highlight key areas in the person's image, while MFIAM enhances the discriminability of features on the different dimensions of the input tensor through three branches. The incorporation of these modules makes the model more accurate and efficient in handling cross-modal retrieval tasks in complex scenarios. Experiments are conducted on three challenging text-to-image person re-identification datasets. The results show that the proposed method outperforms existing mthods on multiple evaluation metrics, demonstrating the effectiveness and robustness of the model.