A novel lightweight FusionMamba-YOLO algorithm with the hybrid Mamba-CNN architecture for real-time scene understanding in assistive robotic navigation
摘要
In this work, a novel lightweight FusionMamba-YOLO algorithm is introduced to address the challenges of real-time scene understanding in the navigation of quadruped robots. Specifically, the proposed algorithm is based on a novel Mamba-YOLO hybrid architecture. First, the FusionMamba backbone network employs a dual-stream interactive fusion mechanism, integrating global four-directional Mamba scanning with Mamba operations based on local windows. This design enables the effective combination of local feature information while efficiently processing global information. Second, a novel small object enhancement pyramid (SOEP) module is designed to dynamically fuse low-level P2 features with high-level semantic features. Consequently, the SOEP module is a lightweight structure that significantly enhances small object detection capabilities without increasing computational overhead. Third, a new pedestrian traffic light instance (PTL-Instance) dataset is designed. This dataset extends the original pedestrian traffic light (PTL) dataset by incorporating data from complex traffic scenes in Beijing with instance-level annotations for traffic signals and pedestrian crossings. The experimental results demonstrate that the proposed algorithm achieves superior performance, with a mean average precision (mAP) of 88.5% and a mask mean average precision (MaskmAP) of 91.4%, while requiring a computational load of only 9.5 GFLOPs and 2.42M parameters. These findings indicate that the proposed framework provides a robust and efficient solution for autonomous navigation in intelligent guide robots.