Background <p>Tandem repeats play critical roles in human disease and phenotypic diversity but are among the most challenging classes of genomic variation to measure accurately. Long-read sequencing has the potential to accurately characterize long and complex tandem repeats. While an increasing number of genotyping methods are available, no systematic effort has been undertaken to evaluate their usability, accuracy, and performance across motifs and allele lengths.</p> Results <p>We reviewed 25 bioinformatic tools and selected seven actively maintained methods for benchmarking using publicly available Oxford Nanopore genome sequencing data from more than 100 individuals. We assessed performance across 43,009 genome-wide tandem repeat loci using four complementary strategies: concordance with haplotype-resolved Human Pangenome Reference Consortium assemblies, Mendelian consistency, cross-tool consistency, and sensitivity to molecularly confirmed pathogenic expansions. Most methods achieved high concordance with assemblies, with higher accuracy using R10 Oxford Nanopore pore chemistry than older R9 chemistry. Accuracy declined with increasing allele length, and most tools performed worse on homopolymers, heterozygous loci, and alleles differing from the reference genome. Assembly concordance and Mendelian consistency did not predict sensitivity to pathogenic expansions, suggesting that these metrics captured distinct aspects of performance.</p> Conclusions <p>No single genotyper performs consistently best across all assessments, but strong contenders emerge in each. Our results demonstrate that length accuracy overestimates tandem repeat genotyping performance. Sequence-level benchmarking is essential for selecting tools best-suited for population studies and clinical diagnostics. This work provides practical guidance for tool selection and highlights key priorities for future long-read tandem repeat genotyping method development.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A comprehensive assessment of tandem repeat genotyping methods for Nanopore long-read genomes

  • Elbay Aliyev,
  • Akshay Avvaru,
  • Wouter De Coster,
  • Garrison M. Arner,
  • Denis M. Nyaga,
  • Sophia B. Gibson,
  • Ben Weisburd,
  • Bida Gu,
  • Claudia Gonzaga-Jauregui,
  • Jonas A. Gustafson,
  • Joy Goffena,
  • Wayne E. Clarke,
  • Evan E. Eichler,
  • Theodore M. Nelson,
  • Anthony A. Snead,
  • Xinxia Peng,
  • Marcelo Ayllon,
  • Nikhita Damaraju,
  • Miranda PG Zalusky,
  • Kendra Hoekzema,
  • David Twesigomwe,
  • Lei Yang,
  • Phillip A. Richmond,
  • Nathan D. Olson,
  • Andrea Guarracino,
  • Qiuhui Li,
  • Angela L. Miller,
  • Zachary B. Anderson,
  • Sophie HR Storz,
  • Anna O. Basile,
  • André Corvelo,
  • Catherine E. Reeves,
  • Adrienne Helland,
  • Rajeeva Lochan Musunuri,
  • Mahler Revsine,
  • Karynne E. Patterson,
  • Cate R. Paschal,
  • Christina Zakarian,
  • Sara Goodwin,
  • Tanner D. Jensen,
  • Esther Robb,
  • W. Richard McCombie,
  • Fritz J. Sedlazeck,
  • Justin M. Zook,
  • Stephen B. Montgomery,
  • Erik Garrison,
  • Mikhail Kolmogorov,
  • Michael C. Schatz,
  • Richard N. McLaughlin Jr,
  • Michael C. Zody,
  • Matthew Loose,
  • Miten Jain,
  • Rob Patro,
  • Caitlin N. Jacques,
  • Mark J. P. Chaisson,
  • Danny E. Miller,
  • Elizabeth Ostrowski,
  • Harriet Dashnow

摘要

Background

Tandem repeats play critical roles in human disease and phenotypic diversity but are among the most challenging classes of genomic variation to measure accurately. Long-read sequencing has the potential to accurately characterize long and complex tandem repeats. While an increasing number of genotyping methods are available, no systematic effort has been undertaken to evaluate their usability, accuracy, and performance across motifs and allele lengths.

Results

We reviewed 25 bioinformatic tools and selected seven actively maintained methods for benchmarking using publicly available Oxford Nanopore genome sequencing data from more than 100 individuals. We assessed performance across 43,009 genome-wide tandem repeat loci using four complementary strategies: concordance with haplotype-resolved Human Pangenome Reference Consortium assemblies, Mendelian consistency, cross-tool consistency, and sensitivity to molecularly confirmed pathogenic expansions. Most methods achieved high concordance with assemblies, with higher accuracy using R10 Oxford Nanopore pore chemistry than older R9 chemistry. Accuracy declined with increasing allele length, and most tools performed worse on homopolymers, heterozygous loci, and alleles differing from the reference genome. Assembly concordance and Mendelian consistency did not predict sensitivity to pathogenic expansions, suggesting that these metrics captured distinct aspects of performance.

Conclusions

No single genotyper performs consistently best across all assessments, but strong contenders emerge in each. Our results demonstrate that length accuracy overestimates tandem repeat genotyping performance. Sequence-level benchmarking is essential for selecting tools best-suited for population studies and clinical diagnostics. This work provides practical guidance for tool selection and highlights key priorities for future long-read tandem repeat genotyping method development.