Background <p>Deep learning models can predict the functional impact of genetic variants. However, training and applying these sequence models to personalized genome sequences are severely constrained by the computational burden of storing and reading massive datasets.</p> Results <p>We developed GenVarLoader, an accelerated dataloader that creates personalized genomic sequences and functional tracks on-the-fly to overcome these bottlenecks. GenVarLoader stores this personalized genomic data in formats specifically optimized for machine learning. Our tool reads variants up to 26,000 times faster than a pre-subset BCF, loads personalized data up to 130 times faster than tuned FASTA and pyBigWig pipelines, and reduces storage approximately 2,000-fold compared to existing alternatives.</p> Conclusions <p>By eliminating severe data loading bottlenecks, GenVarLoader ensures efficient model training and inference. It provides a scalable solution for integrating biobank-scale personalized genomic data with deep learning applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GenVarLoader: an accelerated dataloader for applying deep learning to personalized genomics

  • David Laub,
  • Aaron Ho,
  • Loukik Raina,
  • Jeff Jaureguy,
  • Adam Klie,
  • Rany M. Salem,
  • Graham McVicker,
  • Hannah Carter

摘要

Background

Deep learning models can predict the functional impact of genetic variants. However, training and applying these sequence models to personalized genome sequences are severely constrained by the computational burden of storing and reading massive datasets.

Results

We developed GenVarLoader, an accelerated dataloader that creates personalized genomic sequences and functional tracks on-the-fly to overcome these bottlenecks. GenVarLoader stores this personalized genomic data in formats specifically optimized for machine learning. Our tool reads variants up to 26,000 times faster than a pre-subset BCF, loads personalized data up to 130 times faster than tuned FASTA and pyBigWig pipelines, and reduces storage approximately 2,000-fold compared to existing alternatives.

Conclusions

By eliminating severe data loading bottlenecks, GenVarLoader ensures efficient model training and inference. It provides a scalable solution for integrating biobank-scale personalized genomic data with deep learning applications.