Large Language Models (LLMs) have garnered significant attention due to their powerful natural language understanding and reasoning abilities. In this paper, we investigate the capability of 14 large language models to answer update-to-date information and conduct an analysis through a specific empirical study. The analysis revealed that most models lack the ability to immediately integrate fresh new knowledge, with only a handful, including Baidu's ERNIE Bot and Moonshot AI's Kimi, accurately reflecting updates on real-time events. These models, however, demonstrated a preference for capturing information in their native language, showing a weakness in cross-language knowledge transfer. A significant number of models, regardless of origin, implemented content censorship for politically sensitive questions, opting for refusal to answer or providing neutral responses with complete evidence chains. The study emphasizes the necessity for technological advancements to overcome these limitations, particularly in enhancing real-time knowledge acquisition and multi-language proficiency. It also highlights the potential of certain models to provide a superior user experience, indicating areas for future development in the field of large language models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Can Large Language Model Services Answer the Update-to-Date Information? An Empirical Study

  • Lei Wu,
  • Daiwen Sun,
  • Yu Shang,
  • Xingyi Yu,
  • Yu Zhu,
  • Yangzhao Yang

摘要

Large Language Models (LLMs) have garnered significant attention due to their powerful natural language understanding and reasoning abilities. In this paper, we investigate the capability of 14 large language models to answer update-to-date information and conduct an analysis through a specific empirical study. The analysis revealed that most models lack the ability to immediately integrate fresh new knowledge, with only a handful, including Baidu's ERNIE Bot and Moonshot AI's Kimi, accurately reflecting updates on real-time events. These models, however, demonstrated a preference for capturing information in their native language, showing a weakness in cross-language knowledge transfer. A significant number of models, regardless of origin, implemented content censorship for politically sensitive questions, opting for refusal to answer or providing neutral responses with complete evidence chains. The study emphasizes the necessity for technological advancements to overcome these limitations, particularly in enhancing real-time knowledge acquisition and multi-language proficiency. It also highlights the potential of certain models to provide a superior user experience, indicating areas for future development in the field of large language models.