Hierarchical Language-Conditioned Robot Learning with Vision-Language Models
摘要
Language-conditioned robotic imitation learning plays a vital role at the convergence of robotics, natural language processing, and computer vision. It facilitates the creation of robots that can comprehend their surroundings and carry out intricate tasks in response to human language instructions. However, existing methods exhibit poor generalization to unseen environments, limiting their practical application. To tackle this problem, we introduce an innovative framework that utilizes pre-trained vision-language models (VLMs) for robotic manipulation tasks. Our approach employs a hierarchical decomposition strategy, dividing robot control learning into two levels: high-level vision-language comprehension and task planning, and low-level action execution. The VLMs are slightly fine-tuned using imitation learning on language-conditioned manipulation datasets. We evaluate our method on the CALVIN benchmark for zero-shot long-horizon language-conditioned tasks, the average completed task length improves more than 2.5 times compared to a strong baseline method HULC. This demonstrates significant improvements in generalization capability. Our research presents a simple and effective solution, successfully integrating pre-trained VLMs into robotic control systems, offering valuable insights for future research and practical applications.