<p>The rise of artificial intelligence (AI) presents a unique opportunity to enhance research capacity, yet rigorous evaluation of AI in psychological science is essential. This study explores the application of ChatGPT version 3.5 for developing survey items and compares their psychometric properties to those of a validated instrument measuring the construct of grit. In Study 1, an exploratory factor analysis with 180 college students revealed that AI-generated items replicated the two-factor structure of the Short Grit Scale and demonstrated high internal consistency reliability (Factor 1, α = 0.94; Factor 2, α = 0.93) with moderate to strong correlations with the Short Grit Scale. In Study 2, a confirmatory factor analysis with a larger sample of 366 participants confirmed the two-factor structure with strong factor loadings ranging from 0.78 to 0.88 and acceptable model fit indices (CFI = 0.97, TLI = 0.95, RMSEA = 0.09, SRMR = 0.04). Additionally, hierarchical regression analysis showed that AI-generated items predicted academic performance and explained more of the variance in grade point average than the Short Grit Scale. This study provides early insights into AI’s potential to support researchers in the lengthy scale development process.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Artificial intelligence in scale development: evaluating AI-generated survey items against gold standard measures

  • John Terry,
  • Gerald Strait,
  • Steve Alsarraf,
  • Emily Weinmann,
  • Allison Waychoff

摘要

The rise of artificial intelligence (AI) presents a unique opportunity to enhance research capacity, yet rigorous evaluation of AI in psychological science is essential. This study explores the application of ChatGPT version 3.5 for developing survey items and compares their psychometric properties to those of a validated instrument measuring the construct of grit. In Study 1, an exploratory factor analysis with 180 college students revealed that AI-generated items replicated the two-factor structure of the Short Grit Scale and demonstrated high internal consistency reliability (Factor 1, α = 0.94; Factor 2, α = 0.93) with moderate to strong correlations with the Short Grit Scale. In Study 2, a confirmatory factor analysis with a larger sample of 366 participants confirmed the two-factor structure with strong factor loadings ranging from 0.78 to 0.88 and acceptable model fit indices (CFI = 0.97, TLI = 0.95, RMSEA = 0.09, SRMR = 0.04). Additionally, hierarchical regression analysis showed that AI-generated items predicted academic performance and explained more of the variance in grade point average than the Short Grit Scale. This study provides early insights into AI’s potential to support researchers in the lengthy scale development process.