<p>We develop a scalable LLM-embedding-based framework for semantic analysis of large-scale scholarly corpora. In this framework, modern LLM-derived and pretrained scientific-domain embedding models transform paper abstracts into dense semantic representations, while approximate nearest-neighbor (ANN) retrieval enables efficient corpus-scale similarity computation. Using 307,215 statistics-related papers, we compare six embedding models, including <i>text-embedding-3-small</i> and <i>text-embedding-3-large</i> from OpenAI together with scientific-domain and contrastive models. We evaluate whether <i>Annoy</i>, used as the scalability layer, preserves reliable semantic neighborhoods in these embedding spaces. We utilize the resulting embedding framework to construct a journal-level similarity network that reflects content-based proximity and introduce an LLM-embedding-based semantic novelty measure, defined by semantic distance from prior work. Computational experiments demonstrate semantic discrimination, retrieval fidelity, scalability, and robustness of the proposed framework. This approach supports large-scale scientometric analysis and provides a reproducible semantic perspective on scientific knowledge organization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A scalable framework for scholarly similarity and novelty measurement with LLM-derived semantic embeddings

  • Jin Wang,
  • Yi Ding,
  • Rui Pan

摘要

We develop a scalable LLM-embedding-based framework for semantic analysis of large-scale scholarly corpora. In this framework, modern LLM-derived and pretrained scientific-domain embedding models transform paper abstracts into dense semantic representations, while approximate nearest-neighbor (ANN) retrieval enables efficient corpus-scale similarity computation. Using 307,215 statistics-related papers, we compare six embedding models, including text-embedding-3-small and text-embedding-3-large from OpenAI together with scientific-domain and contrastive models. We evaluate whether Annoy, used as the scalability layer, preserves reliable semantic neighborhoods in these embedding spaces. We utilize the resulting embedding framework to construct a journal-level similarity network that reflects content-based proximity and introduce an LLM-embedding-based semantic novelty measure, defined by semantic distance from prior work. Computational experiments demonstrate semantic discrimination, retrieval fidelity, scalability, and robustness of the proposed framework. This approach supports large-scale scientometric analysis and provides a reproducible semantic perspective on scientific knowledge organization.