Estimating crowd size is a significant area of study within computer vision. Its purpose is to use video or image data to calculate the number of crowds. Deep learning and neural networks have become widely utilized in the realm, offering robust methodologies for analysis and prediction. However, factors such as occlusion, crowd distribution, and chaotic scenes pose obstacles to crowd counting research. To tackle these challenges, this paper proposes a multi-scale confidence-aware feature fusion network to achieve efficient and accurate crowd counting. In this structure, we learn the scale variation of multi-scale features. Due to the different information features that are not fully focused on at multiple scales, we use the original features to learn a confidence score for each scale. Finally, we use cross attention to fuse the learned weighted feature map with the original feature map, fusing features that may be lost in multiple scales or details that cannot be noticed in the initial feature map, resulting in a better feature map and more accurate results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale Confidence-Aware Feature Fusion Network for Crowd Counting

  • Zhengpeng Zhao,
  • Gengshen Wu

摘要

Estimating crowd size is a significant area of study within computer vision. Its purpose is to use video or image data to calculate the number of crowds. Deep learning and neural networks have become widely utilized in the realm, offering robust methodologies for analysis and prediction. However, factors such as occlusion, crowd distribution, and chaotic scenes pose obstacles to crowd counting research. To tackle these challenges, this paper proposes a multi-scale confidence-aware feature fusion network to achieve efficient and accurate crowd counting. In this structure, we learn the scale variation of multi-scale features. Due to the different information features that are not fully focused on at multiple scales, we use the original features to learn a confidence score for each scale. Finally, we use cross attention to fuse the learned weighted feature map with the original feature map, fusing features that may be lost in multiple scales or details that cannot be noticed in the initial feature map, resulting in a better feature map and more accurate results.