Recent research has demonstrated that the flexible creation of generic policies can be achieved by merging multiple task-specific, individually trained policies, although this approach has been limited to tasks within the distribution. The present study investigates whether a prompt-based transformer architecture can enable merged multi-task models to generalize to unseen tasks. In this study, multi-task models are obtained without the requirement for centralized training by averaging and merging the attention layer parameters of single models trained on three different tasks. These models are then fine-tuned based on the prompting trajectories of the test tasks. Additionally, a language pretraining model is introduced as an initialization during the individual model training phase. Experiments conducted in four environments of Mujoco and Meta-World demonstrate that the merge-fine-tuning model performs at least as well as the single-task model in unseen test tasks while significantly reducing the number of convergence steps. This work facilitates the extension of generic models to out-of-distribution environments and provides a novel approach for building generic strategies capable of generalizing to unseen tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Average Merging and Prompting Training for Policy Generalization

  • Dongdong Zhao,
  • Song Yao,
  • Shi Yan

摘要

Recent research has demonstrated that the flexible creation of generic policies can be achieved by merging multiple task-specific, individually trained policies, although this approach has been limited to tasks within the distribution. The present study investigates whether a prompt-based transformer architecture can enable merged multi-task models to generalize to unseen tasks. In this study, multi-task models are obtained without the requirement for centralized training by averaging and merging the attention layer parameters of single models trained on three different tasks. These models are then fine-tuned based on the prompting trajectories of the test tasks. Additionally, a language pretraining model is introduced as an initialization during the individual model training phase. Experiments conducted in four environments of Mujoco and Meta-World demonstrate that the merge-fine-tuning model performs at least as well as the single-task model in unseen test tasks while significantly reducing the number of convergence steps. This work facilitates the extension of generic models to out-of-distribution environments and provides a novel approach for building generic strategies capable of generalizing to unseen tasks.