EMFF-Net: effective multiscale feature fusion network for traffic object detection
摘要
In recent years, object detection in traffic scenes has gained significant traction and has become a crucial component of intelligent transportation systems. To bolster the robustness of object detection in traffic scenes, we propose an effective multiscale feature fusion network (EMFF-Net) inspired by the You Only Look Once (YOLO) architecture. We propose a pyramid pooling architecture with progressive fusion of features to detect as many objects in the scene as possible. To mitigate the impact of the constructed feature pyramid losing shallow feature information during feature extraction, we propose a cross-layer global feature fusion module. Experimental results demonstrate that our method effectively tackles the challenges yielding impressive detection performance. Our proposed model demonstrates competitive performance in robust object detection tasks within traffic scenes when compared to state-of-the-art methods on public datasets. Specifically, our model achieves a remarkable 96.7% mAP@0.5 and 73.0% mAP@0.5:0.95 on the KITTI dataset, marking improvements of 1.5% and 3.4%, respectively.