Data products have emerged as a powerful paradigm for managing data in both intra-enterprise and federated environments, providing structured data assets that include not only the data itself but also services, metadata, and access policies. However, a key challenge in federated environments is the discovery of relevant data products. Traditional discovery mechanisms are heavily dependent on metadata, which is often inconsistent, incomplete, or not standardized across organizations. This lack of metadata quality significantly limits the effectiveness of discovery, making it difficult for consumers to identify and retrieve the data they need. To address this challenge, we propose a content-based discovery framework that shifts the focus from metadata to the actual content of data products. Our approach uses sampling techniques to extract meaningful data representations and a tabular retrieval model for natural language queries. Directly interacting with data improves discovery accuracy, enabling effective data access in federated environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Content-Based Data Product Retrieval in Federated Environments with LLM and Sampling

  • Matteo Falconi,
  • Pierluigi Plebani

摘要

Data products have emerged as a powerful paradigm for managing data in both intra-enterprise and federated environments, providing structured data assets that include not only the data itself but also services, metadata, and access policies. However, a key challenge in federated environments is the discovery of relevant data products. Traditional discovery mechanisms are heavily dependent on metadata, which is often inconsistent, incomplete, or not standardized across organizations. This lack of metadata quality significantly limits the effectiveness of discovery, making it difficult for consumers to identify and retrieve the data they need. To address this challenge, we propose a content-based discovery framework that shifts the focus from metadata to the actual content of data products. Our approach uses sampling techniques to extract meaningful data representations and a tabular retrieval model for natural language queries. Directly interacting with data improves discovery accuracy, enabling effective data access in federated environments.