Multi-turn Instruction Invocation on Human-Robot Interaction by Large Language Models
摘要
The developments of human-robot interaction (HRI) and Large Language Models (LLMs) have paved the way for a wide range of robotics applications spanning from industrial automation to service robotics. Although large language models (LLMs) have demonstrated impressive capabilities, their application in robotics is hindered by a critical limitation: the absence of real-world memory and common sense. This deficiency makes it challenging for robots to comprehend multi-turn instructional commands. For instance, a command such as ‘Remind me to take medicine tomorrow morning’ could lead to ambiguity, as the model may struggle to determine whether this is indeed an instruction and whether additional arguments are required. Moreover, there is a substantial gap in the literature regarding comparative studies on the efficacy of prompt engineering versus supervised fine-tuning for tasks that involve the invocation of robot instructions based on LLMs. Addressing this gap is essential for advancing the integration of LLMs into practical robotics systems and improving human-robot interaction capabilities. In this study, we present a novel multi-turn instruction invocation framework designed to address the challenges of multi-turn instruction invocation and other dialogue-related tasks in robotics. Using a real-world robot dataset, we conduct a comprehensive evaluation of various large-scale models to assess their performance in terms of instruction invocation. This systematic comparison enables us to identify the strengths and limitations of existing approaches and provide insights into the development of more effective robotics systems.