From assistance to accountability: a systematic scoping review of LLMs in responsible scholarly review, research integrity, and scientometrics
摘要
The growing use of large language models (LLMs) across research workflows, including literature review, evidence synthesis, peer-review support, citation validation, and research evaluation, has intensified debates about research integrity, accountability, transparency, and the boundary between legitimate assistance and inappropriate delegation. While LLMs may enhance the efficiency and semantic depth of scientometric and scholarly review practices, their use also raises concerns about hallucinated evidence, fabricated references, biased evaluation, unclear authorship, inadequate disclosure, and the erosion of human responsibility. This study conducts a workflow-oriented systematic scoping review with qualitative thematic synthesis to map and critically synthesize the emerging literature on LLMs in scientometrics and scholarly review processes, rather than a comprehensive review of AI and research integrity as a whole. Following PRISMA-informed procedures, searches were conducted in Scopus, the Web of Science Core Collection, and IEEE Xplore for English-language publications published from 2023 to 2025. After eligibility-based screening and backward citation searching, 74 studies were included as an analytical sample. The review maps five conceptual and methodological approaches: human-in-the-loop and hybrid architectures, retrieval-augmented generation, autonomous and multi-agent research systems, semantic and ontological integration, and prompt engineering with structured evaluation frameworks. It also identifies five major application clusters, review automation and intelligent screening, automated review generation and agentic research workflows, quality assessment and peer-review support, data extraction and semantic classification, and citation analysis and source validation, alongside a governance-oriented category addressing disclosure, confidentiality, accountability, and responsible-use frameworks. The scoping synthesis suggests that retrieval-grounded, domain-adapted, and human-supervised systems offer more reliable pathways than unsupported general-purpose generation, especially for citation-sensitive, evaluative, and high-stakes scholarly tasks. Quantitative performance indicators reported in individual studies were treated descriptively as part of mapping evaluation practices, rather than as directly comparable outcomes for pooled estimation. As a normative implication of the scoping synthesis, the review argues that responsible LLM use should be framed as accountable human–machine collaboration grounded in traceable evidence, transparent disclosure, domain-specific validation, fairness auditing, reproducibility, and non-delegable human responsibility.