MFCA: Multimodal Object Detection Based on Feature Calibration and Aggregation
摘要
Multimodal object detection has recently attracted much attention. By integrating feature from various modalities, the detection models’ accuracy can be significantly enhanced. Despite this, current approaches predominantly emphasize the straightforward fusion of multimodal features, often neglecting the redundant information, such as noise, that can be present within individual modal features. This can lead to models that fail to effectively capture important features of the object, leading to inadequate detection results. To address this problem, we propose a multimodal object detection model based on feature calibration and aggregation, termed MFCA, which comprises a Spatial Calibration Module (SCM) and a Channel Aggregation Module (CAM). The aim is to focus on the information of a mono-modal and supplement it with the valid information from another modality. SCM filters out the redundant information from the other modality and incorporates the valid information into this modality, which preventing errors in features due to incorporation of redundant information. Then, CAM replaces the unimportant feature channels in one modality using another modality to enrich the detailed semantic information of the features. By jointly fusing multimodal features across different dimensions on the same scale, the ability of model to perceive efficient features is enhanced. Experiments performed on the KAIST, FLIR, and VEDAI datasets validate the effectiveness of the proposed scheme, demonstrating that it outperforms other state-of-the-art methods.