Citer: Leveraging Large Language Models and Explainable AI for Analyzing Scientific Collaboration Networks
摘要
Citer represents an innovative approach to analyzing scientific collaboration networks through the integration of large language models (LLMs) and explainable AI techniques. This paper introduces Citer’s comprehensive methodology, which combines weighted graph representations, fine-tuned LLMs, and advanced information retrieval techniques within a Retrieval-Augmented Generation (RAG) pipeline. The system provides deep insights into collaboration patterns, identifies potential research partnerships, and facilitates the creation of topic-specific chatbots for in-depth literature analysis. By synthesizing various strands of research in network analysis, natural language processing, and explainable AI, Citer offers a powerful, transparent, and efficient tool for understanding and navigating the complex landscape of scientific collaboration. Our results demonstrate significant improvements in query classification, impactful word identification, and collaboration network analysis compared to traditional methods. We utilized datasets from arXiv and S2ORC, employing models such as Qwen-2.5B, DistilBERT, and LLaMA 32B in a RAG pipeline to enhance the system’s capabilities. Citer processed over 21.5 million research papers from the arXiv and S2ORC datasets, fine-tuning the Qwen-2.5B model using LoRA to achieve optimal performance in a Retrieval-Augmented Generation (RAG) pipeline. Through this, Citer generated over 150,000 potential collaboration links. The fine-tuning of the RAG model improved contextually relevant response generation by 17%. Moreover, the constructed knowledge graph identified key research clusters, with over 10,000 highly influential nodes enabling detailed analysis of interdisciplinary collaborations. The system also showed a 27% improvement in query classification accuracy compared to traditional methods.