Cross-Modal Hybrid Loss with Enhanced Feature Feedback for RGB-T Crowd Counting
摘要
RGB-T crowd counting integrates visible (RGB) and thermal (T) images to estimate crowd density. This task faces two critical challenges: modality misalignment and weak feature representation. To tackle these challenges, we present an innovative approach by introducing a hybrid loss function with enhanced cross-modal features as feedback. Our approach comprises two key components: a feature enhancement module that iteratively refines and amplifies key features from both modalities and a semantic alignment hybrid loss, which progressively aligns RGB and thermal representations across network hierarchies. Extensive experiments on RGB-T crowd counting benchmarks substantiate the effectiveness of our approach, achieving an average improvement in accuracy of 6% on RGBT-CC and 14% on DroneRGBT compared to previous methods, reaching state-of-the-art performance in RGB-T crowd counting.