The Language of Nuance: Exploring the Limits of Large Language Models in Handling Ambiguity
摘要
Recently launched large language models (LLMs), such as ChatGPT, Copilot, and Gemini, have demonstrated impressive capabilities in understanding and responding to human queries across various domains. However, thorough evaluation is needed, as human language involves ambiguities at multiple levels-phonetic, lexical, syntactic, pragmatic, and discourse. These ambiguities are not merely inherent but contribute to communicative efficiency, posing significant challenges for LLMs. To address this, 24 pragmatic and discourse ambiguous sentences were curated and tested across ChatGPT 4.0, Copilot, and Gemini-an area often overlooked in studies on Generative AI tools due to its contextual complexity. The results indicate that ambiguity is not just an inherent feature but a key element of human language, enhancing communication efficiency through nuanced contextual understanding-a phenomenon still difficult for AI tools to fully comprehend. Notably, ChatGPT showed better results than other tools, achieving 75% accuracy. These findings highlight the need to integrate more nuanced linguistic rules and frameworks into LLMs to improve contextual interpretation and ambiguity handling.