<p>The present study investigates how GPT language models approximate human sentence-processing dynamics in Korean subject honorification. Drawing upon four open-access datasets from prior psycholinguistic experiments—acceptability judgements, politeness ratings, self-paced reading, and eye-tracking reading—we compare human performance with surprisal estimates from three Korean-capable GPT variants (KoGPT-2, Ko-GPT-Trinity, and mGPT). Across tasks, the models reliably detect coarse morphosyntactic violations, assigning high surprisal to clear honorific-agreement mismatches that elicit low acceptability and processing slowdowns in humans. However, GPT surprisal fails to capture graded, context-dependent aspects of behaviour, including the optionality of the subject honorific suffix, animacy-sensitive politeness judgements, and delayed spillover effects in reading. Correlations between human measures and model surprisal are consistently weak and highly condition-specific. These findings suggest that GPT language models provide a useful but limited proxy for human sentence processing: they track whether inputs are broadly well-formed but underrepresent the socio-pragmatic and discourse-level constraints that shape processing dynamics in Korean. This calls for more cognitively and culturally appropriate approaches to computational metrics for explainable AI in language sciences.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

To what extent do GPT language models capture human sentence processing dynamics? Evidence from Korean subject honorification

  • Gyu-Ho Shin,
  • Nayoung Kwon,
  • Seongmin Mun

摘要

The present study investigates how GPT language models approximate human sentence-processing dynamics in Korean subject honorification. Drawing upon four open-access datasets from prior psycholinguistic experiments—acceptability judgements, politeness ratings, self-paced reading, and eye-tracking reading—we compare human performance with surprisal estimates from three Korean-capable GPT variants (KoGPT-2, Ko-GPT-Trinity, and mGPT). Across tasks, the models reliably detect coarse morphosyntactic violations, assigning high surprisal to clear honorific-agreement mismatches that elicit low acceptability and processing slowdowns in humans. However, GPT surprisal fails to capture graded, context-dependent aspects of behaviour, including the optionality of the subject honorific suffix, animacy-sensitive politeness judgements, and delayed spillover effects in reading. Correlations between human measures and model surprisal are consistently weak and highly condition-specific. These findings suggest that GPT language models provide a useful but limited proxy for human sentence processing: they track whether inputs are broadly well-formed but underrepresent the socio-pragmatic and discourse-level constraints that shape processing dynamics in Korean. This calls for more cognitively and culturally appropriate approaches to computational metrics for explainable AI in language sciences.