Robot Vision-and-Language Navigation Action Decision-Making Based on Diffusion Policy
摘要
Robot vision-and-language navigation is a representative task in the field of embodied intelligence. In the navigation environment, the robot observes the scene, follows the natural language instructions, makes decisions by realize the fusion and matching of the above multimodal information, and performs corresponding actions to complete the task. The most of the previous work focused on solving the input problem of the robot and achieved some good results. However, regarding the output problem of the robot, the traditional training methods are based on imitation learning to fit the data or added reinforcement learning to randomly explore and obtain rewards to achieve action decision-making, which leads to insufficient utilization of the dataset. In this paper, in order to alleviate the distribution shift problem caused by the behavior cloning and better learn the distribution of state-action pairs in the dataset, the robot vision-and-language navigation action decision-making policy is represented as a conditional denoising diffusion process, which can be applied to different input perception feature extraction models. Specifically, we use the robot vision-and-language multimodal feature information as the condition of the diffusion model, and use the state-action pairs in the dataset as the supervised learning signal to guide the diffusion policy to generate a series of correct navigation actions. A large number of experimental results on the standard Room-to-Room benchmark show that the diffusion policy can help improve the navigation performance of the robot vision-and-language navigation task.