Systematic Control of Multiple-Choice Item Difficulty Through LLM-Based Distractor Generation
摘要
Advancements in large language models (LLMs) have expanded their use in automatic item generation. In particular, distractor generation (DG), the process of creating distractors for multiple-choice items, has benefited from the flexibility of LLMs to generate test items across various domains. However, despite the growing demand for such applications, systematic methods for controlling difficulty in DG remain underdeveloped. This study explores whether difficulty levels can be systematically adjusted by modifying specific linguistic elements, such as word count, morphosyntactic complexity, and vocabulary complexity, in LLM-generated distractors while keeping the correct answer options unchanged. Furthermore, from the perspective of educational assessment, we evaluated whether the generated items are accurate tools for measuring human achievement by computing item parameters based on item response theory using response data. Our findings confirm that difficulty adjustment in DG can be achieved by systematically modifying distractor elements and that such control is essential for ensuring that items function as intended in real educational contexts. Moreover, our results demonstrate that the items generated using our approach can be validated within the framework of test theory, reinforcing their applicability in educational assessments.