Multi-scale Confidence-Aware Feature Fusion Network for Crowd Counting
摘要
Estimating crowd size is a significant area of study within computer vision. Its purpose is to use video or image data to calculate the number of crowds. Deep learning and neural networks have become widely utilized in the realm, offering robust methodologies for analysis and prediction. However, factors such as occlusion, crowd distribution, and chaotic scenes pose obstacles to crowd counting research. To tackle these challenges, this paper proposes a multi-scale confidence-aware feature fusion network to achieve efficient and accurate crowd counting. In this structure, we learn the scale variation of multi-scale features. Due to the different information features that are not fully focused on at multiple scales, we use the original features to learn a confidence score for each scale. Finally, we use cross attention to fuse the learned weighted feature map with the original feature map, fusing features that may be lost in multiple scales or details that cannot be noticed in the initial feature map, resulting in a better feature map and more accurate results.