With the rapid development of artificial intelligence technology, LLM (Large Language Model) has made significant progress in the field of natural language processing. However, as the scale of these models continues to expand and their application domains broaden, systematic evaluation of their performance has become increasingly important. Although there are many evaluations of large models in general domains, there are very few assessments in specific domains such as the cognitive domain. To accurately evaluate the capabilities of large models in this field, this paper constructs a LLM evaluation system based on cognitive domain, LLM-EC, comprehensively and objectively analyzes and compares the performance of several current open-source models through multi-dimensional and multi-metric evaluation methods, aiming to provide a scientific basis for model selection in practical applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Systematic Evaluation Framework for Cognitive Domain Performance of Large Language Models

  • Chao Yao,
  • Danyang Liu,
  • Qingkai Wei,
  • Jiawei Li,
  • Yifan Zhou,
  • Shiwen Wang,
  • Chaoying Zheng

摘要

With the rapid development of artificial intelligence technology, LLM (Large Language Model) has made significant progress in the field of natural language processing. However, as the scale of these models continues to expand and their application domains broaden, systematic evaluation of their performance has become increasingly important. Although there are many evaluations of large models in general domains, there are very few assessments in specific domains such as the cognitive domain. To accurately evaluate the capabilities of large models in this field, this paper constructs a LLM evaluation system based on cognitive domain, LLM-EC, comprehensively and objectively analyzes and compares the performance of several current open-source models through multi-dimensional and multi-metric evaluation methods, aiming to provide a scientific basis for model selection in practical applications.