Dual-Modal Structural Decoupling with Semantic Relation Distillation for Text-Based Person Search
摘要
Text-based person search (TBPS) aims to identify the target person based on a given textual description query. Existing data augmentation methods for text-based person search struggle to balance cross-modal correspondence and identity preservation. Global mixing suffers from degraded fine grained visual-textual interactions, which in turn compromise identity consistency, while local augmentation leads to a loss of structural consistency across modalities. We propose a novel Semantic Relation Distillation (SRD) framework that addresses structural distortion and identity ambiguity limitations through dual-path structural decoupling. Our approach first designs cross-modal distillation framework that disentangles attribute specific features and achieves hierarchical alignment via semantic constrained recombination. Secondly, structure-aware augmentation via semantic-consistent local interpolation and context preserved replacement. Lastly, we design Semantic Relation Similarity (SRS), a structured alignment metric that unifies scene graph matching with multimodal CLIP embeddings to model cross-modal structural alignment through hierarchical relation propagation and importance-aware graph matching. Extensive experiments on CUHK-PEDES, ICFG-PEDES and RSTPReid validate the effectiveness of our method, yielding a 1.48% gain in Rank-1 accuracy compared to state-of-the-art methods, while also exhibiting stronger cross-dataset generalization capability.