Multi-scale feature fusion network for real-time semantic segmentation of urban street scenes: enhancing detail retention and accuracy
摘要
Efficient and accurate semantic segmentation of urban street scenes is critical for autonomous driving applications. Traditional single-branch networks excel in extracting semantic information but often lose detailed features. Conversely, dual-branch networks preserve detailed features while maintaining two separate branches, impacting real-time performance. To address these limitations, we propose the multi-scale feature fusion network (MFFNet), which integrates the strengths of both approaches. MFFNet retains copies of feature maps at different scales during inference and employs a multi-scale feature interaction module (MFIM) and a weight-based feature fusion module (WFFM) to facilitate effective feature fusion while preserving fine image details. Experimental results on the Cityscapes and CamVid datasets demonstrate that using RTX 3090 for speed evaluation with input resolutions of 1024*2048 and 960*720, MFFNet achieves a mean Intersection over Union (mIoU) of