Advancements in large language models (LLMs) have expanded their applications while raising concerns about safety and ethics, with post-training alignment emerging as a significant development to guide their behavior. In this paper, we introduce refined Direct Preference Optimization (rDPO), a method for improving the behavioral alignment of LLMs without the need for human-annotated data. The method involves creating synthetic data using self-critique prompting by a teacher LLM and then utilizing a generalized DPO loss function to distil to a student LLM. The loss function incorporates an additional external reward model to improve the quality of synthetic data, making rDPO robust to potential noise in the synthetic dataset. rDPO is shown to be effective in a diverse set of behavioral alignment tasks, such as improved safety, robustness against role-playing, and reduced sycophancy. Our results confirm that rDPO is more data-efficient than vanilla DPO, the current technique for aligning LLMs with offline data. Code available at github.com/vicgalle/refined-dpo

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs

  • Víctor Gallego

摘要

Advancements in large language models (LLMs) have expanded their applications while raising concerns about safety and ethics, with post-training alignment emerging as a significant development to guide their behavior. In this paper, we introduce refined Direct Preference Optimization (rDPO), a method for improving the behavioral alignment of LLMs without the need for human-annotated data. The method involves creating synthetic data using self-critique prompting by a teacher LLM and then utilizing a generalized DPO loss function to distil to a student LLM. The loss function incorporates an additional external reward model to improve the quality of synthetic data, making rDPO robust to potential noise in the synthetic dataset. rDPO is shown to be effective in a diverse set of behavioral alignment tasks, such as improved safety, robustness against role-playing, and reduced sycophancy. Our results confirm that rDPO is more data-efficient than vanilla DPO, the current technique for aligning LLMs with offline data. Code available at github.com/vicgalle/refined-dpo