Adaptive Dual Cross-Attention Network for Multispectral Object Detection
摘要
Multispectral object detection is a challenging task due to the inherent sensitivity of convolutional neural networks (CNNs) to modality misalignment, which is often exacerbated by the limited receptive fields of CNNs. To address this, recent fusion strategies have shifted toward Transformer architectures, which can effectively model long-range dependencies through cross-attention mechanisms. However, these strategies typically rely on rigid modality assignments, where specific modalities—such as visible images and thermal images—are fixed as either the reference or sensed modality. This rigid assignment can limit the model’s ability to generalize across different scenarios, where the roles of modalities may vary based on the scene or environmental conditions. In this paper, we propose the Adaptive Dual Cross-Attention Network (ADCA-Net), a novel dual-stream framework that dynamically adjusts modality roles during the fusion process. Experiments show that we achieve 86.2% and 97.6% mAP@50 metrics on DVTOD and LLVIP datasets, respectively, surpassing the state-of-the-art (SOTA) method by 1.2% and 0.5%.