Background <p>Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in European population.</p> Results <p>Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including quantitative traits, time-to-event traits, ordinal traits, and longitudinal traits. Since no alternative model is fitted, SPAmix primarily focuses on association p values. For each genetic variant, SPAmix uses genotype data and genetic principal components to estimate individual-specific allele frequency, which is subsequently used to calibrate <i>p</i> values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. We also propose SPAmix<sub>local</sub> to incorporate local ancestry to calculate ancestry-specific <i>p</i> values. To maximize the statistical powers, SPAmix<sub>CCT</sub> is proposed to combine the <i>p</i> values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination.</p> Conclusions <p>The SPAmix-based approaches are more accurate than Tractor to address phenotypic variance heterogeneity among ancestries when analyzing quantitative traits and to address an unbalanced case–control ratio when analyzing binary traits. SPAmix<sub>CCT</sub> is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SPAmix: a scalable, accurate, and universal analysis framework for large-scale genetic association studies in admixed populations

  • Yuzhuo Ma,
  • He Xu,
  • Ying Li,
  • Hyesung Kim,
  • Lin-lin Xu,
  • Lin Miao,
  • Peng Xu,
  • Fengbiao Mao,
  • Xu-jie Zhou,
  • Wei Zhou,
  • Seunggeun Lee,
  • Ji-Feng Zhang,
  • Peipei Zhang,
  • Wenjian Bi

摘要

Background

Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in European population.

Results

Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including quantitative traits, time-to-event traits, ordinal traits, and longitudinal traits. Since no alternative model is fitted, SPAmix primarily focuses on association p values. For each genetic variant, SPAmix uses genotype data and genetic principal components to estimate individual-specific allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. We also propose SPAmixlocal to incorporate local ancestry to calculate ancestry-specific p values. To maximize the statistical powers, SPAmixCCT is proposed to combine the p values of SPAmix and SPAmixlocal via Cauchy combination.

Conclusions

The SPAmix-based approaches are more accurate than Tractor to address phenotypic variance heterogeneity among ancestries when analyzing quantitative traits and to address an unbalanced case–control ratio when analyzing binary traits. SPAmixCCT is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.