This paper studies the design and implementation methods of a human-robot interaction system for social robots. The system consists of three modules: user information processing, dialogue generation, and speech synthesis. By comprehensively utilizing multimodal data such as speech and images, the system achieves personalized services. We analyze typical scenarios and categorize dialogue, one of the most important functions in human-robot interaction, into three main types: service-oriented dialogue, consultation-oriented dialogue, and casual conversation. To effectively address the different requirements of these three types of dialogues, we propose three different dialogue generation strategies: using a state transition-based method for high-precision service functions, leveraging large models combined with a local knowledge base for domain knowledge consultation, and directly using large language models for general casual conversation. For the core component, the intent recognition module, we chose to fine-tune the RoBERTa model after evaluating various models and introduced BiLSTM to deeply integrate the feature information encoded by RoBERTa. This approach effectively captures the dependencies between data, improving accuracy by 2.2% compared to the original model, thereby enhancing the naturalness and efficiency of the user experience.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Design of a Human-Robot Interaction System for Social Robots Assisted by Large Models

  • Xiao Xiao,
  • Yingchao Tang,
  • Chuying Guan,
  • Lanfang Dong

摘要

This paper studies the design and implementation methods of a human-robot interaction system for social robots. The system consists of three modules: user information processing, dialogue generation, and speech synthesis. By comprehensively utilizing multimodal data such as speech and images, the system achieves personalized services. We analyze typical scenarios and categorize dialogue, one of the most important functions in human-robot interaction, into three main types: service-oriented dialogue, consultation-oriented dialogue, and casual conversation. To effectively address the different requirements of these three types of dialogues, we propose three different dialogue generation strategies: using a state transition-based method for high-precision service functions, leveraging large models combined with a local knowledge base for domain knowledge consultation, and directly using large language models for general casual conversation. For the core component, the intent recognition module, we chose to fine-tune the RoBERTa model after evaluating various models and introduced BiLSTM to deeply integrate the feature information encoded by RoBERTa. This approach effectively captures the dependencies between data, improving accuracy by 2.2% compared to the original model, thereby enhancing the naturalness and efficiency of the user experience.