VTMF \(^2\) N: Towards Accurate Visual-Tactile Slip Detection via Multi-modal Feature Fusion in Robotic Grasping
摘要
Perceiving whether the touched object slips is essential for stable robotic grasping and adaptively fusing visual images and tactile information is the effective way to achieve accurate slip detection. Previous studies tend to separately extract visual-tactile information and simply fuse them with feature concatenation, lacking the deep investigation of multi-modal fusion strategy and leading to restricted slip detection accuracy. To solve this problem, we propose a novel Visual-Tactile Multi-modal Feature Fusion Network, termed VTMF \(^2\) N, with efficient Channel Shuffle and Cross Attention (CSCA) module, to extract and fuse slipping-related features from the spatial and temporal aspects. Specifically, the CSCA module uses channel shuffle across multi-modality to exchange visual and tactile information for slipping-specific feature interaction, followed by cross attention mechanism to simultaneously learn cross-modality information for slipping-specific feature enhancement. Comparative experiments on public datasets both indicate that our VTMF \(^2\) N achieves SOTA performance and when used as a plug-in module, the proposed CSCA is able to improve the accuracy of two baseline networks (CNN-LSTM, CNN-MSTCN). The visual-tactile robotic grasping experiment is also conducted to prove the validity of the proposed method under real-world scenarios.