A single multi-task deep neural network with a multi-scale feature aggregation mechanism for manipulation relationship reasoning in robotic grasping
摘要
In stacked object scenes, effectively perceiving the positional relationships between objects is crucial for enabling robots to grasp safely. To further improve the robot’s ability to understand these positional relationships, this paper proposes a multi-task grasping detection model for stacking object scenes, based on multi-scale feature fusion. We design a multi-scale feature aggregation (MSFA) module to fuse features, enhancing the feature maps’ ability to represent objects of varying sizes. Additionally, we leverage features from multiple views to predict object positional relationships. In particular, the local intersection region feature (IRF) significantly enhances the model’s capability to identify positional relationships between objects while improving computational efficiency. Experimental results on the Visual Manipulation Relationship Dataset (VMRD) demonstrate that our model surpasses current state-of-the-art methods. Furthermore, we conducted grasping experiments in real-world stacking scenarios, which verified the model’s effectiveness and generalization capability.