Distractors are incorrect answer options designed to mislead or confuse test-takers in multiple-choice reading comprehension questions. In real-world exam settings, creating distractors for English reading comprehension questions is complex and varied, with subjective and diverse evaluation standards. Developing a distractor generation technique that meets real-world requirements is a highly challenging task with significant research value. To address these challenges, we introduce DGRL (Distractors Generation based on Reinforcement Learning from preference feedback), a method using cutting-edge large language models trained through reinforcement learning to generate multiple distractors for real-world human exam. First, the distractor generation model is fine-tuned through supervised fine-tuning (SFT) on a reading comprehension question dataset. Then, using preference feedback reinforcement learning, we build and train a reward model to evaluate the quality of individual distractors. Combining the reward model with a diversity evaluation metric, we design an objective function and further train the fine-tuned model using reinforcement learning. Experiments show that the DGRL, after SFT and reinforcement learning, can generate multiple high-quality distractors that meet the real-world requirements in one go, serving as a valuable reference and aid in real-world question-setting for human exam.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

High-Quality Distractors Generation for Human Exam Based on Reinforcement Learning from Preference Feedback

  • Ruofan Wang,
  • Yuru Jiang,
  • Yuyang Tao,
  • Mengyuan Li,
  • Xia Wang,
  • Shili Ge

摘要

Distractors are incorrect answer options designed to mislead or confuse test-takers in multiple-choice reading comprehension questions. In real-world exam settings, creating distractors for English reading comprehension questions is complex and varied, with subjective and diverse evaluation standards. Developing a distractor generation technique that meets real-world requirements is a highly challenging task with significant research value. To address these challenges, we introduce DGRL (Distractors Generation based on Reinforcement Learning from preference feedback), a method using cutting-edge large language models trained through reinforcement learning to generate multiple distractors for real-world human exam. First, the distractor generation model is fine-tuned through supervised fine-tuning (SFT) on a reading comprehension question dataset. Then, using preference feedback reinforcement learning, we build and train a reward model to evaluate the quality of individual distractors. Combining the reward model with a diversity evaluation metric, we design an objective function and further train the fine-tuned model using reinforcement learning. Experiments show that the DGRL, after SFT and reinforcement learning, can generate multiple high-quality distractors that meet the real-world requirements in one go, serving as a valuable reference and aid in real-world question-setting for human exam.