GCA-Net: Global Contextual Attention Network with Lightweight Hierarchical Alignment for Text-Guided Fashion Image Retrieval
摘要
Text-guided fashion image retrieval faces critical challenges in local spatial context modeling and global feature redundancy removal, leading to suboptimal performance in complex user queries such as “long sleeve shirt with floral patterns but without collar”. To address these, we propose the Global Contextual Attention Network (GCA-Net), a novel framework featuring two key innovations:(1) a Contextual self-Attention module (CoA) module that captures long-range dependencies by integrating static local context and dynamic global attention;(2) a Lightweight Global Attention Module (LGAM) that enhances cross-modal alignment through channel-spatial hierarchical interactions, and resolves feature redundancy in fused representations. Combined with a hybrid loss balancing fine/coarse-grained matching, GCA-Net demonstrates competitive results on the FashionIQ benchmark. (Recall@50: 63.08%, + 1.69% over MUR) and Shoes (Recall@50: 81.54%, + 0.22% over Css-Net), demonstrating superior robustness in handling attribute-dense descriptions.