DSFAT: a dual-stream framework assisted by textual information for person re-identification in real scenes
摘要
Person re-identification (Re-ID) is a technique employed to recognize a target pedestrian based on source information and video sequences from non-overlapping camera fields of view. This technique has garnered increasing attention due to its extensive application prospects in intelligent video surveillance within the internet of things and other contexts. In this paper, we propose a multi-task learning approach that incorporates text information to enhance recognition accuracy. We employ text as auxiliary data, utilizing a dual-stream transformer’s encoder to extract both image and text features. To further improve the model’s interaction and feature learning capabilities, we introduce a cross-modal interaction encoder (CIE) and a feature sharing learning (FSL) network. The CIE facilitates fine-grained alignment between text and image features, while the FSL network learns modality-invariant feature representations. Our method, DSFAT, demonstrates superior performance in person re-identification tasks when compared to state-of-the-art methods, as validated on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets.