<p>Large Language Models (LLMs) are transforming industrial-organizational psychology and human resource management, with one of their most promising applications being automatic item generation (AIG) for psychological test development. Although recent advances in LLM-based AIG—particularly for non-cognitive assessments such as personality— show significant potential, ensuring rigorous quality control remains a persistent challenge. This study introduces a novel AIG framework, the LLM-based Multi-agent AIG system (LM-AIG), where each agent is responsible for different stages of item development, including item generation, content review, linguistic evaluation, bias assessment, and item revision. The LM-AIG also incorporates human feedback to enhance item quality. We implemented the LM-AIG framework using the open-source tool AutoGen to generate items assessing attitudes toward the use of AI in the workplace. To evaluate the quality of the generated items, we conducted an empirical study based on structured ratings from human raters, assessing construct relevance, linguistic clarity, appropriate language level, contextual specificity, and potential bias. This paper further discusses the role of human-in-the-loop mechanisms within the LM-AIG system and outlines future research directions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AI-powered Automatic Item Generation for Psychological Tests: A Conceptual Framework for an LLM-based Multi-Agent AIG System

  • Philseok Lee,
  • Mina Son,
  • Zihao Jia

摘要

Large Language Models (LLMs) are transforming industrial-organizational psychology and human resource management, with one of their most promising applications being automatic item generation (AIG) for psychological test development. Although recent advances in LLM-based AIG—particularly for non-cognitive assessments such as personality— show significant potential, ensuring rigorous quality control remains a persistent challenge. This study introduces a novel AIG framework, the LLM-based Multi-agent AIG system (LM-AIG), where each agent is responsible for different stages of item development, including item generation, content review, linguistic evaluation, bias assessment, and item revision. The LM-AIG also incorporates human feedback to enhance item quality. We implemented the LM-AIG framework using the open-source tool AutoGen to generate items assessing attitudes toward the use of AI in the workplace. To evaluate the quality of the generated items, we conducted an empirical study based on structured ratings from human raters, assessing construct relevance, linguistic clarity, appropriate language level, contextual specificity, and potential bias. This paper further discusses the role of human-in-the-loop mechanisms within the LM-AIG system and outlines future research directions.