Data matching is a persistent challenge in heterogeneous data integration. Traditional data matching methods, relying on schema-based, instance-based, or hybrid approaches, often fall short when aligning disparate data from schema-only with instance-only sources. To address this problem of aligning disparate sources, we present an innovative data matching framework that enables the matching of one data source, where only schema-related information is available, with another data source, where only instance-related information is available. The strength of our framework lies in its ability to combine outputs from multiple Schema-Instances matchers, generating auxiliary information to enhance the alignment between disparate data structures. Our framework is validated using the Valentine Benchmark through extensive experiments. These findings underscore the potential of our approach to advance the integration of diverse data sources.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Alignment of Schema-Only and Instance-Only Data Sources Using Large Language Models

  • Nour Elhouda Kired,
  • Franck Ravat,
  • Jiefu Song,
  • Olivier Teste

摘要

Data matching is a persistent challenge in heterogeneous data integration. Traditional data matching methods, relying on schema-based, instance-based, or hybrid approaches, often fall short when aligning disparate data from schema-only with instance-only sources. To address this problem of aligning disparate sources, we present an innovative data matching framework that enables the matching of one data source, where only schema-related information is available, with another data source, where only instance-related information is available. The strength of our framework lies in its ability to combine outputs from multiple Schema-Instances matchers, generating auxiliary information to enhance the alignment between disparate data structures. Our framework is validated using the Valentine Benchmark through extensive experiments. These findings underscore the potential of our approach to advance the integration of diverse data sources.