<p>Structural variants (SVs) contribute significantly to genetic diversity yet present computational challenges during analysis. We introduce SDFA, a standardized decomposition format and toolkit for efficient analysis of SVs in large-scale population genomics. SDFA efficiently stores and retrieves all SV types while providing algorithms for consistent SV merging, memory-efficient annotation, and precise gene feature annotation across large cohorts. SDFA outperforms existing tools, achieving at least 17.64 times faster merging than four tools and 120.93 times faster annotation than three tools, and uniquely handles complex SVs. We validate SDFA on 895,054 SVs from 150,119 individuals in the UK Biobank dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SDFA: a standardized decomposition format and toolkit for efficient analysis of structural variants in large-scale population genomic studies

  • Wenjie Peng,
  • Liubin Zhang,
  • Bin Tang,
  • Junhao Liang,
  • Zhi Liu,
  • Lihang Ye,
  • Yangyang Yuan,
  • Yifei Wang,
  • Ruijie Tan,
  • Nan Lin,
  • Chao Xue,
  • Hui Jiang,
  • Li Fang,
  • Miaoxin Li

摘要

Structural variants (SVs) contribute significantly to genetic diversity yet present computational challenges during analysis. We introduce SDFA, a standardized decomposition format and toolkit for efficient analysis of SVs in large-scale population genomics. SDFA efficiently stores and retrieves all SV types while providing algorithms for consistent SV merging, memory-efficient annotation, and precise gene feature annotation across large cohorts. SDFA outperforms existing tools, achieving at least 17.64 times faster merging than four tools and 120.93 times faster annotation than three tools, and uniquely handles complex SVs. We validate SDFA on 895,054 SVs from 150,119 individuals in the UK Biobank dataset.