Aligned with the principles of Industry 5.0, Human-Centric Smart Manufacturing (HSM) emphasizes human well-being by integrating AI-driven solutions to enhance human-robot collaboration (HRC). In recent years, Convolutional Neural Networks (CNNs) have played a key role in advancing machine learning, computer vision, and robotics. Vision Transformers (ViTs) have emerged as powerful models for various vision tasks, sparking interest in their broader applications. However, their adoption in robotics remains relatively unexplored. This study explores the application of Vision Transformers for classifying human gestures and objects, enhancing human-robot collaboration (HRC) through advanced image recognition techniques. By employing a multimodal late fusion approach, visual data from multiple streams are integrated, leveraging complementary information to improve decision-making accuracy. Experimental results show that the accuracy of ViT-based unimodal classification models and the performance of robot task execution using a fusion model highlight the significance of multimodal fusion models in human-robot collaboration. Furthermore, this research is important for the realization of multimodal-based robotic learning, planning, and decision-making, contributing to advancements in Human-Centric Smart Manufacturing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vision Transformer-Based Multimodal Fusion of Gesture and Object Classification in Human-Robot Collaboration

  • Selma Subašić,
  • Lejla Banjanović-Mehmedović,
  • Haris Subašić,
  • Isak Karabegović,
  • Ermin Husak

摘要

Aligned with the principles of Industry 5.0, Human-Centric Smart Manufacturing (HSM) emphasizes human well-being by integrating AI-driven solutions to enhance human-robot collaboration (HRC). In recent years, Convolutional Neural Networks (CNNs) have played a key role in advancing machine learning, computer vision, and robotics. Vision Transformers (ViTs) have emerged as powerful models for various vision tasks, sparking interest in their broader applications. However, their adoption in robotics remains relatively unexplored. This study explores the application of Vision Transformers for classifying human gestures and objects, enhancing human-robot collaboration (HRC) through advanced image recognition techniques. By employing a multimodal late fusion approach, visual data from multiple streams are integrated, leveraging complementary information to improve decision-making accuracy. Experimental results show that the accuracy of ViT-based unimodal classification models and the performance of robot task execution using a fusion model highlight the significance of multimodal fusion models in human-robot collaboration. Furthermore, this research is important for the realization of multimodal-based robotic learning, planning, and decision-making, contributing to advancements in Human-Centric Smart Manufacturing.