XFeat-VINS: Integrating Efficient Feature Extraction with Visual-Inertial State Estimation for Robust Localization and Mapping
摘要
In the realm of visual-inertial systems, achieving accurate state estimation is paramount for a variety of applications, including robotic navigation, autonomous vehicles, and augmented reality. Despite the advancements made by some studies that integrate monocular cameras with Inertial Measurement Units (IMUs), there remains an opportunity to enhance the computational efficiency and robustness of their feature extraction and tracking methodologies. Our research introduces an innovative VINS system, XFeat-VINS, which integrates XFeat as its core feature extraction and tracking engine. XFeat is a compact and high-performance convolutional neural network tailored for devices with constrained resources, ensuring high precision while substantially increasing processing velocity. The system refines the feature extraction and matching processes, as well as the state estimation, by harnessing XFeat’s feature descriptors. Furthermore, it incorporates a loop closure detection mechanism based on these descriptors, thereby fortifying the system’s global consistency and resilience. The efficacy of XFeat-VINS has been substantiated through extensive testing on pertinent datasets and in authentic environmental conditions.