Multi-pretrained-Model-Based Ensemble Semantic Similarity Ranking of Chinese-English Legal Sentence Pairs
摘要
The development of deep learning and large language models has further enhanced the importance of high-quality language data. Recently, with the support of deep learning algorithms, codes, and evaluation methods, natural legal language processing has made rapid progress. Currently, natural legal language processing still faces the problem of scarcity of bilingual legal language data, and more importantly, the quality of existing data is uneven. We address the quality evaluation of Chinese-English legal sentence pairs, explore a new method based on multiple pretrained models, and propose an ensemble semantic similarity ranking algorithm. The experimental results show that our method can comprehensively utilize the semantic measurement capabilities of multiple large language models, and the corresponding implementation is an industrial-level straightforward and efficient algorithm.