<p>In recent years, object detection in traffic scenes has gained significant traction and has become a crucial component of intelligent transportation systems. To bolster the robustness of object detection in traffic scenes, we propose an effective multiscale feature fusion network (EMFF-Net) inspired by the You Only Look Once (YOLO) architecture. We propose a pyramid pooling architecture with progressive fusion of features to detect as many objects in the scene as possible. To mitigate the impact of the constructed feature pyramid losing shallow feature information during feature extraction, we propose a cross-layer global feature fusion module. Experimental results demonstrate that our method effectively tackles the challenges yielding impressive detection performance. Our proposed model demonstrates competitive performance in robust object detection tasks within traffic scenes when compared to state-of-the-art methods on public datasets. Specifically, our model achieves a remarkable 96.7% <i>mAP@0.5</i> and 73.0% <i>mAP@0.5:0.95</i> on the KITTI dataset, marking improvements of 1.5% and 3.4%, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EMFF-Net: effective multiscale feature fusion network for traffic object detection

  • Zhong Qu,
  • Shize Fan,
  • Xuehui Yin

摘要

In recent years, object detection in traffic scenes has gained significant traction and has become a crucial component of intelligent transportation systems. To bolster the robustness of object detection in traffic scenes, we propose an effective multiscale feature fusion network (EMFF-Net) inspired by the You Only Look Once (YOLO) architecture. We propose a pyramid pooling architecture with progressive fusion of features to detect as many objects in the scene as possible. To mitigate the impact of the constructed feature pyramid losing shallow feature information during feature extraction, we propose a cross-layer global feature fusion module. Experimental results demonstrate that our method effectively tackles the challenges yielding impressive detection performance. Our proposed model demonstrates competitive performance in robust object detection tasks within traffic scenes when compared to state-of-the-art methods on public datasets. Specifically, our model achieves a remarkable 96.7% mAP@0.5 and 73.0% mAP@0.5:0.95 on the KITTI dataset, marking improvements of 1.5% and 3.4%, respectively.