Point-line visual–inertial SLAM with enhanced line detection and robust loop closure
摘要
While most existing simultaneous localization and mapping (SLAM) systems rely on point features due to their computational efficiency and robustness in unstructured environments, they often encounter challenges in dynamic or texture-sparse scenarios. To address these limitations, recent studies have incorporated line features, which provide additional geometric constraints and are particularly effective in structured environments. However, commonly used line detection algorithms, such as LSD and EDLines, tend to generate numerous short, fragmented segments, degrading pose estimation accuracy. Moreover, most loop closure detection approaches rely on bag-of-words (BoW) techniques, which lack adaptability to unseen environments and perform poorly under severe appearance changes. To overcome these challenges, we propose IPL-VI-SLAM, an improved point-line visual–inertial SLAM framework with three key innovations: (1) a modified enhanced line segment detector (MELSD) is introduced, which incorporates fusion of similar line and uniform distribution of line features to detect more accurate, continuous, and stable line segments, (2) a robust loop closure detection strategy based on multilayer perceptrons (MLP) is proposed to better handle appearance variations and complex environmental changes, (3) an improved point-line visual–inertial SLAM framework that effectively combines MELSD and the MLP-based loop closure to achieve higher robustness and localization accuracy. In addition, the experiments are conducted on multiple sequences of the EuRoC and the TUM-VI datasets and the experimental results indicate that: (1) the given MELSD algorithm significantly enhances the quality of the extracted line features, (2) the MLP-based loop closure strategy outperforms traditional BoW-based methods, (3) the provided IPL-VI-SLAM system surpasses other state-of-the-art VI-SLAM systems by achieving an average localization accuracy improvement of 46.60%, 30.37%, 61.80%, 33.73%, 45.54%, and 26.66% on the EuRoC dataset compared to VINS-Mono, Open-VINS, LET-VINS, PL-VINS, UV-SLAM, and EPLF-VINS, respectively. For the TUM-VI datasets, it achieves improvements of 25.47%, 8.13%, 22.54%, and 10.12% over VINS-Mono, PL-VINS, and EPLF-VINS, respectively, ranking the second only to Open-VINS.