AI-generated familiarity estimates are a useful new source of information about word knowledge in Simplified Chinese
摘要
This study evaluated the usefulness of AI-generated estimates of word familiarity for predicting word difficulty in Simplified Chinese, building on previous research in alphabetic languages. We found that familiarity estimates produced using large language models (LLMs) showed moderate-to-strong correlations with human familiarity ratings. These LLM estimates were the most effective predictors of both word naming and lexical decision times, surpassing traditional metrics such as word frequency and human familiarity ratings, while the latter still provided modest, non-overlapping variance. GPT-4o with English instructions produced superior results compared to the Chinese-centered models currently available. The results imply that LLM familiarity estimates are a valuable resource for Chinese psycholinguistics, supporting work across experimental design, modeling, and norming. We release familiarity estimates for 27,624 words for unrestricted research and educational use.