Systematic Survey of Large Language Model Evaluation: Methods, Benchmarks, and Comprehensive Taxonomic Framework
摘要
Large Language Models (LLMs) have rapidly reshaped research and practice in natural language processing, but methods for evaluating these systems have developed more slowly. In this survey we review work on LLM evaluation published between 2020 and 2025, drawing on 182 studies identified through a PRISMA 2020–guided search of five major sources (ACL Anthology, IEEE Xplore, ACM Digital Library, arXiv, and Google Scholar) and subsequent quantitative and qualitative analysis. On the basis of this literature, we propose a three-dimensional framework that groups evaluation methods by methodology (intrinsic, extrinsic, and human-based), by scope (task-specific, general-purpose, and domain-specific), and by the Evaluation Framework Approach employed (automated metrics, benchmark suites, and interactive/hybrid assessment). We use this framework to characterize current practice and to highlight three recurring difficulties: bias introduced by evaluation design choices, rapid saturation of popular benchmarks, and weak correspondence between standard metrics and performance in real use cases. We then examine how widely used evaluation frameworks such as HELM, LM-Evaluation Harness, and TruLens partially address, but in some respects also reproduce, these limitations, particularly in multilingual, culturally diverse, and deployed settings. The survey concludes with practical recommendations for designing and reporting LLM evaluations, guidance for practitioners selecting or building evaluation pipelines, and a research agenda that emphasizes closer alignment between evaluation protocols and the conditions under which LLMs are actually used.