TermRAG: Data Resource Summarization via Term Retrieval Augmented Generation
摘要
A comprehensive Data Resource System could provide semantic standardization across data from various sources, facilitating cross-business recognition and serving as the cornerstone in addressing the challenge of data silos. However, existing systems, e.g., the Urban Knowledge System (UKS), primarily focus on offering data abstractions, still challenging in refining the semantics of data. These difficulties arise from the openness and the redundancy of textual semantics, which can result in differences among a nuanced distinction of textual terms. To address these challenges, we propose the Term Retrieval Augmented Generation (TermRAG) to precisely refine the semantics of data and expand the terms. Specifically, TermRAG extracts the correlated expressions from historical data, and proactively expands them with foresight by the Large Language Models (LLMs). With the alignment of the human understanding and the generating processes, it greatly improves the generated quality and provides crucial insights to experts. We benchmark our approach against other baselines under three real-world datasets, with results highlighting our method’s superiority. The refined terms have been successfully integrated into Beijing’s municipal government applications, significantly boosting data interoperability efficiency.