Perceiving whether the touched object slips is essential for stable robotic grasping and adaptively fusing visual images and tactile information is the effective way to achieve accurate slip detection. Previous studies tend to separately extract visual-tactile information and simply fuse them with feature concatenation, lacking the deep investigation of multi-modal fusion strategy and leading to restricted slip detection accuracy. To solve this problem, we propose a novel Visual-Tactile Multi-modal Feature Fusion Network, termed VTMF \(^2\) N, with efficient Channel Shuffle and Cross Attention (CSCA) module, to extract and fuse slipping-related features from the spatial and temporal aspects. Specifically, the CSCA module uses channel shuffle across multi-modality to exchange visual and tactile information for slipping-specific feature interaction, followed by cross attention mechanism to simultaneously learn cross-modality information for slipping-specific feature enhancement. Comparative experiments on public datasets both indicate that our VTMF \(^2\) N achieves SOTA performance and when used as a plug-in module, the proposed CSCA is able to improve the accuracy of two baseline networks (CNN-LSTM, CNN-MSTCN). The visual-tactile robotic grasping experiment is also conducted to prove the validity of the proposed method under real-world scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VTMF \(^2\) N: Towards Accurate Visual-Tactile Slip Detection via Multi-modal Feature Fusion in Robotic Grasping

  • Qi’an Tang,
  • Lu Chen,
  • Jingyang Liu,
  • Huaiyao Wang

摘要

Perceiving whether the touched object slips is essential for stable robotic grasping and adaptively fusing visual images and tactile information is the effective way to achieve accurate slip detection. Previous studies tend to separately extract visual-tactile information and simply fuse them with feature concatenation, lacking the deep investigation of multi-modal fusion strategy and leading to restricted slip detection accuracy. To solve this problem, we propose a novel Visual-Tactile Multi-modal Feature Fusion Network, termed VTMF \(^2\) N, with efficient Channel Shuffle and Cross Attention (CSCA) module, to extract and fuse slipping-related features from the spatial and temporal aspects. Specifically, the CSCA module uses channel shuffle across multi-modality to exchange visual and tactile information for slipping-specific feature interaction, followed by cross attention mechanism to simultaneously learn cross-modality information for slipping-specific feature enhancement. Comparative experiments on public datasets both indicate that our VTMF \(^2\) N achieves SOTA performance and when used as a plug-in module, the proposed CSCA is able to improve the accuracy of two baseline networks (CNN-LSTM, CNN-MSTCN). The visual-tactile robotic grasping experiment is also conducted to prove the validity of the proposed method under real-world scenarios.