<p>Pedestrian detection is a fundamental task in intelligent transportation systems and autonomous driving. Deep learning methods provide effective solutions but still face problems such as variable scale of objects, small objects and pedestrian occlusion, which bring great challenges in this field. This study introduces an enhanced version of the dense pedestrian detection method, RT-DETR-MSS, which is built upon the RT-DETR framework. We make three primary contributions. First, the Multi-Scale Fusion Module is proposed, which enhances multi-scale feature fusion and captures fine-grained spatial information, improving overall performance. Second, we design the Small Object Enhance Structure, which integrates low-level fine-grained features into high-level features to improve small object detection. Third, we introduce the GSConv and Slim-Neck architecture VoV-GSCSP module. This method simplifies the network, optimizes computation, and preserves accuracy. Experiments were performed on the CrowdHuman and WiderPerson datasets to assess the effectiveness of our approach. Compared to the baseline RT-DETR model, our method improved mAP0.5 by 1.7<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="530_2025_1746_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> and 0.8<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="530_2025_1746_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation>, and mAP0.5:0.95 by 2.2<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="530_2025_1746_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> and 1.2<InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="530_2025_1746_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> on the CrowdHuman and WiderPerson datasets, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An improved Multi-Scale Fusion and Small Object Enhancement method for efficient pedestrian detection in dense scenes

  • Yalin Song,
  • Peng Qian,
  • Kexin Zhang,
  • Shichong Liu,
  • Rui Zhai,
  • Ran Song

摘要

Pedestrian detection is a fundamental task in intelligent transportation systems and autonomous driving. Deep learning methods provide effective solutions but still face problems such as variable scale of objects, small objects and pedestrian occlusion, which bring great challenges in this field. This study introduces an enhanced version of the dense pedestrian detection method, RT-DETR-MSS, which is built upon the RT-DETR framework. We make three primary contributions. First, the Multi-Scale Fusion Module is proposed, which enhances multi-scale feature fusion and captures fine-grained spatial information, improving overall performance. Second, we design the Small Object Enhance Structure, which integrates low-level fine-grained features into high-level features to improve small object detection. Third, we introduce the GSConv and Slim-Neck architecture VoV-GSCSP module. This method simplifies the network, optimizes computation, and preserves accuracy. Experiments were performed on the CrowdHuman and WiderPerson datasets to assess the effectiveness of our approach. Compared to the baseline RT-DETR model, our method improved mAP0.5 by 1.7 \(\%\) % and 0.8 \(\%\) % , and mAP0.5:0.95 by 2.2 \(\%\) % and 1.2 \(\%\) % on the CrowdHuman and WiderPerson datasets, respectively.