<p>An accurate scene understanding is conducive to the seamless integration of virtual objects into real-world environments in real-time graphics and augmented reality, particularly for small-scale or distant objects under complex imagery conditions. Recent advances in small object scene recognition have mainly focused on backbone modifications or neck enhancements, but these strategies often fail to capture fine-grained local variations and lack sufficient cross-scale fusion, leading to rendering artifacts like incorrect occlusion and unstable virtual object placement, which breaks immersion. Therefore, we propose a novel recognition framework that embeds graphics-inspired fine-grained representation and extraction into visual feature modeling. Specifically, the proposed Attention-Guided Pyramid Feature Fusion module (AGPFF) incorporates 3D attention within a spatial pyramid structure, enabling fine-grained detail preservation. For feature aggregation, we adopt a dynamic up-sampling method called DySample to reconstruct high-frequency details, while the SPPCSPC module ensures robust cross-scale feature fusion to size-varying targets. Furthermore, an Efficient Multi-scale Attention (EMA) module strengthens the interplay between global context and local cues, and an auxiliary high-resolution prediction head is introduced to improve recognition in cluttered or occluded scenarios. Experimental results demonstrate that the proposed approach maintains an optimal accuracy–speed trade-off, achieving 48.4% mAP and 68.5% mAP50 on the VisDrone2019-DET-val set while outperforming both baseline and prior arts. Extensive experiments on TinyPerson, UA-DETRAC and Virtual KITTI 2 benchmarks demonstrate our method’s strong generalization in cross-domain scenarios of computer graphics and vision.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing scene understanding via attention-guided pyramid fusion recognition under complex imagery conditions

  • Chuanyan Hao,
  • Qinghua Qin,
  • Yiming Han,
  • Hao Zhang,
  • Wanru Song

摘要

An accurate scene understanding is conducive to the seamless integration of virtual objects into real-world environments in real-time graphics and augmented reality, particularly for small-scale or distant objects under complex imagery conditions. Recent advances in small object scene recognition have mainly focused on backbone modifications or neck enhancements, but these strategies often fail to capture fine-grained local variations and lack sufficient cross-scale fusion, leading to rendering artifacts like incorrect occlusion and unstable virtual object placement, which breaks immersion. Therefore, we propose a novel recognition framework that embeds graphics-inspired fine-grained representation and extraction into visual feature modeling. Specifically, the proposed Attention-Guided Pyramid Feature Fusion module (AGPFF) incorporates 3D attention within a spatial pyramid structure, enabling fine-grained detail preservation. For feature aggregation, we adopt a dynamic up-sampling method called DySample to reconstruct high-frequency details, while the SPPCSPC module ensures robust cross-scale feature fusion to size-varying targets. Furthermore, an Efficient Multi-scale Attention (EMA) module strengthens the interplay between global context and local cues, and an auxiliary high-resolution prediction head is introduced to improve recognition in cluttered or occluded scenarios. Experimental results demonstrate that the proposed approach maintains an optimal accuracy–speed trade-off, achieving 48.4% mAP and 68.5% mAP50 on the VisDrone2019-DET-val set while outperforming both baseline and prior arts. Extensive experiments on TinyPerson, UA-DETRAC and Virtual KITTI 2 benchmarks demonstrate our method’s strong generalization in cross-domain scenarios of computer graphics and vision.