<p>This study explores the potential of using large language models (LLMs) for automating fine-grained speech act annotation by assessing GPT-4o’s and DeepSeek’s performance in this task. This fine-grained annotation refers to the annotation of speech acts within the framework of local grammar, which annotates both speech act utterances and pragmatically meaningful syntactic units of a speech act utterance. Zooming in on the speech act of thanking and drawing on data taken from the British National Corpus, our investigation found that both models achieved high accuracy – 90.29% for GPT-4o and 92.95% for DeepSeek respectively, indicating that LLMs can approach human performance in domains that have traditionally relied on manual annotation. The subsequent detailed marker-by-marker analyses revealed that each model exhibits strengths and vulnerabilities; specifically, GPT-4o excelled with frequent, informal and context-dependent markers, while DeepSeek performed better with explicit and formal markers. Overall, the study shows that LLMs have great potential to facilitate complex tasks such as fine-grained speech act annotation, which not only means that LLMs can be a valuable methodological resource but also highlights the possibility of developing a human-LLM collaboration framework for speech act research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large Language Models for Automating Fine-grained Speech Act Annotation: A Critical Evaluation of GPT-4o and DeepSeek

  • Hang Su,
  • Jun Ye

摘要

This study explores the potential of using large language models (LLMs) for automating fine-grained speech act annotation by assessing GPT-4o’s and DeepSeek’s performance in this task. This fine-grained annotation refers to the annotation of speech acts within the framework of local grammar, which annotates both speech act utterances and pragmatically meaningful syntactic units of a speech act utterance. Zooming in on the speech act of thanking and drawing on data taken from the British National Corpus, our investigation found that both models achieved high accuracy – 90.29% for GPT-4o and 92.95% for DeepSeek respectively, indicating that LLMs can approach human performance in domains that have traditionally relied on manual annotation. The subsequent detailed marker-by-marker analyses revealed that each model exhibits strengths and vulnerabilities; specifically, GPT-4o excelled with frequent, informal and context-dependent markers, while DeepSeek performed better with explicit and formal markers. Overall, the study shows that LLMs have great potential to facilitate complex tasks such as fine-grained speech act annotation, which not only means that LLMs can be a valuable methodological resource but also highlights the possibility of developing a human-LLM collaboration framework for speech act research.