A recent line of work has investigated the use of corpus graphs to improve the latency-vs-effectiveness envelope of information retrieval systems. The key idea is to build a document-to-document similarity graph offline, allowing additional relevance signals to be exploited during query processing. However, these graphs are inherently expensive to build, requiring a quadratic “all-pairs similarity” computation. In this work, we examine the problem of building corpus graphs using bag-of-words models, and explore heuristics to build high quality graphs at a fraction of the total cost of exhaustive algorithms. We demonstrate that simple mechanisms such as document titles, expanded surrogate queries, and high impact terms can yield effective graphs at a fraction of the cost of their exhaustive counterparts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Approximate Bag-of-Words Top-k Corpus Graphs

  • Lachlan Dunn,
  • Luke Gallagher,
  • Joel Mackenzie

摘要

A recent line of work has investigated the use of corpus graphs to improve the latency-vs-effectiveness envelope of information retrieval systems. The key idea is to build a document-to-document similarity graph offline, allowing additional relevance signals to be exploited during query processing. However, these graphs are inherently expensive to build, requiring a quadratic “all-pairs similarity” computation. In this work, we examine the problem of building corpus graphs using bag-of-words models, and explore heuristics to build high quality graphs at a fraction of the total cost of exhaustive algorithms. We demonstrate that simple mechanisms such as document titles, expanded surrogate queries, and high impact terms can yield effective graphs at a fraction of the cost of their exhaustive counterparts.