Human–robot interaction through joint robot planning with large language models
摘要
Large language models (LLMs) have demonstrated remarkable zero-shot generalisation capabilities, expanding their utility beyond natural language processing into various applications. Leveraging extensive web knowledge, these models generate meaningful text data in response to user-defined prompts, introducing a novel mode of interaction with software applications. Recent investigations have extended the generalisability of LLMs into the domain of robotics, addressing challenges in existing robot learning techniques such as reinforcement learning and imitation learning. This paper explores the application of LLMs for robot planning as an alternative approach to generate high-level robot plans based on prompts provided to the language model. The proposed methodology facilitates continuous user interaction and adjustment of task execution plans in real time. A pre-trained LLM is utilised for collaborative human–robot planning using natural language, complemented by vision language models (VLMs) responsible for generating scene descriptions incorporated into the prompt for contextual grounding. Evaluation of the system is conducted within the VIMABench benchmark simulated environment. Further practicality assessment involves experimentation with the system on a robotic arm engaged in tabletop manipulation activities. The observed results reveal that soliciting human feedback for zero-shot plans generated by LLMs in robot manipulation yields an 8.6% performance increase in simulated evaluations and a 14% increase in physical evaluations compared to the baseline, showcasing the efficacy of the proposed approach.