<p>With the growing accumulation of scattered single-cell data and the rapid advancement of artificial intelligence (AI), there is a pressing need for a high-quality, well-organized, and AI-ready single-cell data resources to support large-scale model. Here, we present version 2.0 of human Ensemble Cell Atlas (hECA), a cell atlas incorporating both single-cell RNA sequencing (scRNA-seq) and single-cell ATAC sequencing (scATAC-seq) data. It expands the scRNA-seq data collection to 10,831,024 human cells with unified labels, and adds the new modality of scATAC-seq profiles with 1,450,511 cells. The data cover 42 human organs and tissues. To ensure cross-dataset consistency and quality, we standardized gene expression and chromatin accessibility matrices, harmonized cellular metadata, and manually re-annotated cell types based on the unified Hierarchical Annotation Framework (uHAF). The strength of the dataset has been shown in pre-training the large generative cellular AI model scMulan. hECA2.0 provides a well-structured and ready-to-use data resource, serving as a robust data foundation for AI-driven single-cell research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

hECA v2.0: an AI-ready ensemble cell atlas of single-cell RNA and ATAC sequencing data

  • Xi Xi,
  • Yixin Chen,
  • Xinze Wu,
  • Minsheng Hao,
  • Jiaqi Li,
  • Haiyang Bian,
  • Qiuchen Meng,
  • Fanhong Li,
  • Chen Li,
  • Chuxi Xiao,
  • Xiaomin Dong,
  • Renke You,
  • Yifan Xiong,
  • Peng Yang,
  • Zijing Gao,
  • Xuejian Cui,
  • Yan Pan,
  • Zhen Li,
  • Wenrui Li,
  • Zhuofeng Li,
  • Xiaoyang Chen,
  • Yanfei Cui,
  • Hairong Lv,
  • Rui Jiang,
  • Lei Wei,
  • Xuegong Zhang

摘要

With the growing accumulation of scattered single-cell data and the rapid advancement of artificial intelligence (AI), there is a pressing need for a high-quality, well-organized, and AI-ready single-cell data resources to support large-scale model. Here, we present version 2.0 of human Ensemble Cell Atlas (hECA), a cell atlas incorporating both single-cell RNA sequencing (scRNA-seq) and single-cell ATAC sequencing (scATAC-seq) data. It expands the scRNA-seq data collection to 10,831,024 human cells with unified labels, and adds the new modality of scATAC-seq profiles with 1,450,511 cells. The data cover 42 human organs and tissues. To ensure cross-dataset consistency and quality, we standardized gene expression and chromatin accessibility matrices, harmonized cellular metadata, and manually re-annotated cell types based on the unified Hierarchical Annotation Framework (uHAF). The strength of the dataset has been shown in pre-training the large generative cellular AI model scMulan. hECA2.0 provides a well-structured and ready-to-use data resource, serving as a robust data foundation for AI-driven single-cell research.