A Novel Multi-modal Traffic Target Recognition Method Based on Channel Attention
摘要
Currently, existing target recognition methods do not adequately address multi-modal traffic target recognition tasks. This paper proposes a novel multi-modal traffic target recognition method based on a channel attention mechanism. Building upon a cross-modal pedestrian retrieval model with a semantic self-matching network, this study introduces an innovative channel attention mechanism and designs a multi-view network incorporating channel attention. This approach refines the network’s weight allocation while automatically extracting partial-level text features for the corresponding visual regions, thereby enhancing the capture of relationships between targets and other objects in the background, and establishing better correspondences between them. Additionally, a combined ranking with strong and weak supervision terms is employed to provide extra guidance for the selected targets, effectively reducing intra-class variance of text features. Finally, this paper constructs the INRIAPerson urban traffic target dataset, on which the matching map improves by over 2% compared to multi-modal target recognition methods such as SSAN and MMMA.