Integrating Multi-granularity Text Semantics into Dense Math Information Retrieval Via Domain Adaptation
摘要
The use of advanced dense retrieval models to interpret and process the formulas is essential for enhancing math information retrieval (MathIR). Traditional MathIR methods, however, often fail to fully leverage both the formulas and their associated textual contexts, which limits the precise comprehension of complex mathematical queries and impedes retrieval efficacy. To overcome this challenge, we propose an innovative framework that integrates multi-granularity text semantics into dense math information retrieval through domain adaptation techniques. By employing a multi-granularity analysis, our approach captures the semantic nuances of formula-associated text across various structural levels. Furthermore, domain adaptation strategies enable our model to align closely with the distinct linguistic features of scientific texts, effectively reducing the semantic discrepancies between specialized literature and general language. We also investigate both symmetric and asymmetric semantic search methods within this framework, illustrating how different semantic representations and matching mechanisms can enhance the relevance and precision of retrieval outcomes. Our experimental results demonstrate that our framework exhibits superior performance in processing intricate mathematical queries, successfully bridging the semantic gap between scientific literature and general language text, and offering a promising new direction for mathematical information retrieval.