Average Merging and Prompting Training for Policy Generalization
摘要
Recent research has demonstrated that the flexible creation of generic policies can be achieved by merging multiple task-specific, individually trained policies, although this approach has been limited to tasks within the distribution. The present study investigates whether a prompt-based transformer architecture can enable merged multi-task models to generalize to unseen tasks. In this study, multi-task models are obtained without the requirement for centralized training by averaging and merging the attention layer parameters of single models trained on three different tasks. These models are then fine-tuned based on the prompting trajectories of the test tasks. Additionally, a language pretraining model is introduced as an initialization during the individual model training phase. Experiments conducted in four environments of Mujoco and Meta-World demonstrate that the merge-fine-tuning model performs at least as well as the single-task model in unseen test tasks while significantly reducing the number of convergence steps. This work facilitates the extension of generic models to out-of-distribution environments and provides a novel approach for building generic strategies capable of generalizing to unseen tasks.