<p>Large language models are increasingly applied to political discourse, but their ability to detect culturally grounded gender sensitivity remains underexplored. We introduce KOGENT, a benchmark dataset of 1,222 transcripts from the Korean National Assembly, annotated for gender sensitivity across 6,024 utterances. Each utterance is labeled as high or low in gender sensitivity, based on contextual indicators of bias, discrimination, or inclusion, and tagged for the target group. KOGENT spans Korean legislative sessions from 1948 to 2024. Annotation reliability was ensured through dual coding and adjudication, yielding high intercoder agreement. When tasked with labeling utterances by gender sensitivity, GPT-4.1 achieved F1-scores of 87.5% (zero-shot) and 91.2% (18-shot), while GPT-4o reached 90.4% and 91.1%, respectively. While incorporating in-domain examples enhanced model performance, limitations in distinguishing between criticisms and reinforcements of inequality, culturally specific terminology, and extended contexts were observed for both models. Our results demonstrate KOGENT’s utility as a robust benchmark for analyzing gender sensitivity in Korean political speech and evaluating multilingual LLMs’ sociocultural alignment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A benchmark dataset for evaluating gender sensitivity in Korean political discourse with large language models

  • Sunkyoung Park,
  • Eunbi Cho,
  • Chan Young Jung,
  • Woo Chang Kang,
  • Taegyoon Kim,
  • Eunah Park,
  • Sanghoun Song

摘要

Large language models are increasingly applied to political discourse, but their ability to detect culturally grounded gender sensitivity remains underexplored. We introduce KOGENT, a benchmark dataset of 1,222 transcripts from the Korean National Assembly, annotated for gender sensitivity across 6,024 utterances. Each utterance is labeled as high or low in gender sensitivity, based on contextual indicators of bias, discrimination, or inclusion, and tagged for the target group. KOGENT spans Korean legislative sessions from 1948 to 2024. Annotation reliability was ensured through dual coding and adjudication, yielding high intercoder agreement. When tasked with labeling utterances by gender sensitivity, GPT-4.1 achieved F1-scores of 87.5% (zero-shot) and 91.2% (18-shot), while GPT-4o reached 90.4% and 91.1%, respectively. While incorporating in-domain examples enhanced model performance, limitations in distinguishing between criticisms and reinforcements of inequality, culturally specific terminology, and extended contexts were observed for both models. Our results demonstrate KOGENT’s utility as a robust benchmark for analyzing gender sensitivity in Korean political speech and evaluating multilingual LLMs’ sociocultural alignment.