<p>High-entropy alloy (HEA) electrocatalysts offer a large design space for the hydrogen evolution reaction (HER), but literature data are sparse, heterogeneous and often difficult to reuse for quantitative modelling. Here we present a curated dataset of 180 multinary alloy catalysts compiled from published sources, with standardized provenance (DOIs), normalized atomic-percent compositions, harmonized metadata, and two HER targets: onset potential and Tafel slope. Each entry is enriched with composition-derived Magpie descriptors, and we provide a mutual-information-selected subset to support reproducible small-data benchmarking. To characterize the dataset structure, we report a composition-only UMAP embedding with K-means partitioning and element-wise sensitivity analyses within clusters. As a technical validation and reference baseline, we train Gaussian-process regression models with uncertainty estimates and evaluate their performance under random and grouped-by-DOI cross-validation. Finally, we use the grouped-by-DOI results to quantify the transferability of composition-derived baselines to unseen literature sources, and provide an uncertainty-aware, model-guided list of 100 candidate alloys intended to guide future standardized HER experiments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Curated dataset of multinary alloy HER catalysts for composition-only modelling with Magpie descriptors and GP baselines

  • Stéphane Gorsse,
  • Tang Bijun,
  • Yaoyao Tang,
  • Ma Mingyu,
  • Zheng Liu

摘要

High-entropy alloy (HEA) electrocatalysts offer a large design space for the hydrogen evolution reaction (HER), but literature data are sparse, heterogeneous and often difficult to reuse for quantitative modelling. Here we present a curated dataset of 180 multinary alloy catalysts compiled from published sources, with standardized provenance (DOIs), normalized atomic-percent compositions, harmonized metadata, and two HER targets: onset potential and Tafel slope. Each entry is enriched with composition-derived Magpie descriptors, and we provide a mutual-information-selected subset to support reproducible small-data benchmarking. To characterize the dataset structure, we report a composition-only UMAP embedding with K-means partitioning and element-wise sensitivity analyses within clusters. As a technical validation and reference baseline, we train Gaussian-process regression models with uncertainty estimates and evaluate their performance under random and grouped-by-DOI cross-validation. Finally, we use the grouped-by-DOI results to quantify the transferability of composition-derived baselines to unseen literature sources, and provide an uncertainty-aware, model-guided list of 100 candidate alloys intended to guide future standardized HER experiments.