<p>The validity of multiple-choice questions (MCQs) in reading comprehension assessments relies heavily on the quality of the distractors. However, the manual design of these distractors is both time-consuming and costly, prompting researchers to turn to computer technology for the automatic generation of distractors. This task involves the process of taking a reading comprehension article, a question and its corresponding correct answer as input, with the goal of generating distractors that are related to the answer, semantically consistent with the question, and traceable within the article. Initially, heuristic rule-based approaches were employed, to generate only word-level or phrase-level distractors. Recent studies have shifted towards using sequence-to-sequence neural networks for sentence-level distractor generation. Despite these advancements, these methods face two key challenges: difficulty in capturing long-distance semantic relationships within the context, leading to overly general or context-independent distractors, and the tendency for the generated distractors to be semantically similar. To address these limitations, this paper proposes a Transformer-Enhanced Hierarchical Encoding with Multi-Decoder (THE-MD) network, composed of a hierarchical encoder and multiple decoders. Specifically, the encoder employs the Transformer architecture to encode the context and capture long-range semantic information, thereby generating more contextually relevant distractors. The decoder utilizes multiple decoding strategies and a dissimilarity loss function to collaboratively generate diverse distractors. The experimental results show that the THE-MD model outperforms existing baselines on both automatic and manual evaluation metrics. On the RACE and RACE++ datasets, the model increased the BLEU-4 scores to 7.45 and 10.60, and the ROUGE-L scores to 22.96 and 34.88, while also demonstrating excellent performance in fluency and coherence metrics. These improvements highlight their potential to enhance the generation of MCQ distractors in educational assessments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer-enhanced hierarchical encoding with multi-decoder for diversified MCQ distractor generation

  • Xiaohui Dong,
  • Zhengluo Li,
  • Haoming Su,
  • Jixiang Xue,
  • Xiaochao Dang

摘要

The validity of multiple-choice questions (MCQs) in reading comprehension assessments relies heavily on the quality of the distractors. However, the manual design of these distractors is both time-consuming and costly, prompting researchers to turn to computer technology for the automatic generation of distractors. This task involves the process of taking a reading comprehension article, a question and its corresponding correct answer as input, with the goal of generating distractors that are related to the answer, semantically consistent with the question, and traceable within the article. Initially, heuristic rule-based approaches were employed, to generate only word-level or phrase-level distractors. Recent studies have shifted towards using sequence-to-sequence neural networks for sentence-level distractor generation. Despite these advancements, these methods face two key challenges: difficulty in capturing long-distance semantic relationships within the context, leading to overly general or context-independent distractors, and the tendency for the generated distractors to be semantically similar. To address these limitations, this paper proposes a Transformer-Enhanced Hierarchical Encoding with Multi-Decoder (THE-MD) network, composed of a hierarchical encoder and multiple decoders. Specifically, the encoder employs the Transformer architecture to encode the context and capture long-range semantic information, thereby generating more contextually relevant distractors. The decoder utilizes multiple decoding strategies and a dissimilarity loss function to collaboratively generate diverse distractors. The experimental results show that the THE-MD model outperforms existing baselines on both automatic and manual evaluation metrics. On the RACE and RACE++ datasets, the model increased the BLEU-4 scores to 7.45 and 10.60, and the ROUGE-L scores to 22.96 and 34.88, while also demonstrating excellent performance in fluency and coherence metrics. These improvements highlight their potential to enhance the generation of MCQ distractors in educational assessments.