Large Language Models (LLMs) are predominantly trained on English data, leading to significant performance challenges for low-resource languages. This study focuses on Urdu as a case study to explore how LLMs process low-resource languages. We find that when prompted in a low-resource language, LLMs primarily reason internally in English, and this internal reasoning is more coherent than the generated text in the target language. This contrast highlights a gap between the model’s ability to comprehend low-resource languages and its ability to generate text in them. By analyzing these mechanisms, this work underscores the need for targeted improvements to enhance LLM performance for low-resource languages such as Urdu.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Understanding vs. Generation: LLMs Are Better at Comprehending Low-Resource Languages like Urdu Than Generating Text in Them

  • Taaha Saleem Bajwa

摘要

Large Language Models (LLMs) are predominantly trained on English data, leading to significant performance challenges for low-resource languages. This study focuses on Urdu as a case study to explore how LLMs process low-resource languages. We find that when prompted in a low-resource language, LLMs primarily reason internally in English, and this internal reasoning is more coherent than the generated text in the target language. This contrast highlights a gap between the model’s ability to comprehend low-resource languages and its ability to generate text in them. By analyzing these mechanisms, this work underscores the need for targeted improvements to enhance LLM performance for low-resource languages such as Urdu.