The rapid expansion of datasets across various domains is driven by advancements in data collection, storage technologies, and the growing demand for high-quality data. However, this has created a messy and scattered situation. While dataset search platforms have contributed to dataset accessibility, significant challenges remain in dataset search and utilization, including the lack of sufficient datasets, unified and rich descriptive metadata, and a dataset classification model. We address these issues by developing a global dataset repository with a unified metadata structure and an automated dataset classification system. Our repository aggregates datasets from reputable sources and organizes them using a two-level hierarchical categorization system that classifies datasets by data type and application topic. To automate dataset classification, we introduce SelectDataset, a hierarchical model that utilizes large language models (LLMs) for inference augmentation, improving classification accuracy for complex and lengthy textual descriptions. Our approach also incorporates decoupling training to address class imbalance, further enhancing model robustness. Through comprehensive experiments, we demonstrate that SelectDataset offers a more efficient and accurate solution.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SelectDataset: Enabling Dataset Exploration Through Enriched Descriptive Metadata and Hierarchical Model-Based Application Topic Classification

  • Ruixin Yuan,
  • Ning Li,
  • Qiantai Peng,
  • Hengrun Zhang,
  • HuiQun Yu,
  • Guisheng Fan

摘要

The rapid expansion of datasets across various domains is driven by advancements in data collection, storage technologies, and the growing demand for high-quality data. However, this has created a messy and scattered situation. While dataset search platforms have contributed to dataset accessibility, significant challenges remain in dataset search and utilization, including the lack of sufficient datasets, unified and rich descriptive metadata, and a dataset classification model. We address these issues by developing a global dataset repository with a unified metadata structure and an automated dataset classification system. Our repository aggregates datasets from reputable sources and organizes them using a two-level hierarchical categorization system that classifies datasets by data type and application topic. To automate dataset classification, we introduce SelectDataset, a hierarchical model that utilizes large language models (LLMs) for inference augmentation, improving classification accuracy for complex and lengthy textual descriptions. Our approach also incorporates decoupling training to address class imbalance, further enhancing model robustness. Through comprehensive experiments, we demonstrate that SelectDataset offers a more efficient and accurate solution.