<p>With the rapid expansion of the Internet and the ever-growing volume of data, search engines face increasing difficulties in providing users with the most pertinent content for their queries. The typical search process, where a user inputs a query, and the system returns a list of pages, often falls short in delivering highly relevant results, with existing ranking methods not always aligning with user expectations. This paper addresses these challenges by developing a novel web page retrieval method to help users obtain more relevant content. The proposed method utilizes a Google custom search engine to respond to the user query, collect the search results along with its metadata from the Google dataset, and store it in JSON file via the Google API for the reranking process. This paper proposes a novel query expansion model that leverages the generative abilities of large language models, specifically ChatGPT, for interactive and automated query expansion to enhance the accuracy of the research results. The proposed model uses two metrics, namely cosine similarity and word mover’s distance, to assess the similarity between user queries and retrieve results by utilizing document metadata by considering the syntactic and semantic aspects of the text. The proposed method is very effective, and the results show a marked improvement in the search results compared to the results retrieved using the Bing, DuckDuckGo, and Google page rank algorithms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing web page retrieval performance with advanced query expansion: leveraging ChatGPT and metadata-driven analysis

  • Ali A. Alani,
  • Adil Al-Azzawi

摘要

With the rapid expansion of the Internet and the ever-growing volume of data, search engines face increasing difficulties in providing users with the most pertinent content for their queries. The typical search process, where a user inputs a query, and the system returns a list of pages, often falls short in delivering highly relevant results, with existing ranking methods not always aligning with user expectations. This paper addresses these challenges by developing a novel web page retrieval method to help users obtain more relevant content. The proposed method utilizes a Google custom search engine to respond to the user query, collect the search results along with its metadata from the Google dataset, and store it in JSON file via the Google API for the reranking process. This paper proposes a novel query expansion model that leverages the generative abilities of large language models, specifically ChatGPT, for interactive and automated query expansion to enhance the accuracy of the research results. The proposed model uses two metrics, namely cosine similarity and word mover’s distance, to assess the similarity between user queries and retrieve results by utilizing document metadata by considering the syntactic and semantic aspects of the text. The proposed method is very effective, and the results show a marked improvement in the search results compared to the results retrieved using the Bing, DuckDuckGo, and Google page rank algorithms.