Similarity Measurement Model for Standard Clauses Based on RAG
摘要
With the continuous advancement of the economy and society, the digitization of standards is a crucial direction for the further development of standardization. However, for the power industry standards, the digitization process is still in its initial stage, requiring significant human effort in the writing and processing of standards. Additionally, due to the complexity and specialization of the text in power industry standards clauses, traditional machine learning models struggle to accurately understand the semantic information of these clauses, meeting the needs of standard digitization development. To address this issue, this paper focuses on the practical application scenarios of writing power industry standard clauses and analyzing the differences in similar clauses. By introducing pre-trained large language models and a retrieval-augmented generation framework, we construct a vector database of power industry standard clauses and propose a similarity measurement model for standard clauses based on RAG. Finally, through specific application tests, we verify that the model constructed in this paper can accurately match power industry standard knowledge information based on user input and automatically perform difference analysis, providing a new solution for the digitization of power industry standards.