BEAC: Imitating Complex Exploration and Task-oriented Behaviors for Invisible Object Nonprehensile Manipulation
摘要
Imitation learning (IL) is a promising approach for acquiring policies to automate robotic tasks from human demonstrations, leveraging the high cognitive capabilities of humans. However, applying IL to nonprehensile manipulation tasks involving invisible objects under partial observability, such as excavating buried rocks, remains challenging. In such settings, the demonstrator must make complex action decisions, including exploratory actions to locate the object and task-oriented actions to accomplish the task, while inferring the object’s hidden state. This often leads to inconsistent demonstrations and imposes a high cognitive load. For these problems, insights from cognitive science suggest that encouraging demonstrators to follow simple, pre-designed exploration rules can help mitigate the problems of action inconsistency and high cognitive load. Accordingly, when performing IL from demonstrations guided by such exploration rules, it is crucial to imitate not only the demonstrator’s task-oriented behavior but also his/her mode-switching behavior (between exploration and task-oriented behavior) under partial observability. Based on the above considerations, this paper proposes a novel IL framework, called Belief Exploration-Action Cloning (BEAC), which employs a switching policy structure that integrates a pre-designed exploration policy with a task-oriented action policy trained on belief states estimated from past history. Through simulation and real-robot experiments, we demonstrate that BEAC achieves superior task performance, higher accuracy in mode and action prediction, and reduced demonstrator cognitive load, as confirmed by a user study.