<p>In this paper, we present a methodology to classify dataset entries in datasets, based on their relevance for answering different specific queries. It employs a repeated individualized inference approach to identify entries with significant Shapley values, contributing with accurate answers to queries about other entries in the dataset. This information is captured in three matrices: a general relevance matrix, a Shapley value matrix, and a significant Shapley value matrix. Since usually the information in datasets is non-homogeneously distributed, relevance is often concentrated in a few entries. This is in particular observed in a representative case study.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identifying Highly Relevant Entries in Datasets: A Relevance-Based Classification

  • Fernando Delbianco,
  • Fernando Tohmé

摘要

In this paper, we present a methodology to classify dataset entries in datasets, based on their relevance for answering different specific queries. It employs a repeated individualized inference approach to identify entries with significant Shapley values, contributing with accurate answers to queries about other entries in the dataset. This information is captured in three matrices: a general relevance matrix, a Shapley value matrix, and a significant Shapley value matrix. Since usually the information in datasets is non-homogeneously distributed, relevance is often concentrated in a few entries. This is in particular observed in a representative case study.