A fundamental goal of genetics is to identify which and how genetic variants are associated with a trait, often using the regression summary statistics from genome-wide association (GWA) studies. Important methodological challenges are accounting for inflation in GWA effect estimates as well as investigating multiple traits simultaneously. We leverage machine learning approaches for these two challenges, developing a computationally efficient method called ML-MAGES. First, we shrink the inflation in GWA effect sizes caused by non-independence among variants using neural networks. We then cluster variant associations among multiple traits via variational inference. We compare the performance of neural network shrinkage to regularized regression and fine-mapping, which both address inflation but handle focal regions of different sizes. Our method outperforms those frameworks in approximating the true effects in simulated data. Our infinite mixture clustering approach offers a flexible, data-driven way to distinguish different types of associations—trait-specific, shared across traits, or spurious—among multiple traits based on their regularized effects. Clustering applied to our neural network shrinkage results also produces consistently higher precision and recall for distinguishing gene-level associations in simulations. We demonstrate the application of ML-MAGES on association analyses of two quantitative traits and two binary traits from the UK Biobank data. The identified genes from single-trait enrichment tests overlap with those having known relevant biological processes to the traits. Besides trait-specific associations, ML-MAGES identifies several variants with shared multi-trait associations, suggesting putative shared genetic architecture.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ML-MAGES: A Machine Learning Framework for Multivariate Genetic Association Analyses with Genes and Effect Size Shrinkage

  • Xiran Liu,
  • Lorin Crawford,
  • Sohini Ramachandran

摘要

A fundamental goal of genetics is to identify which and how genetic variants are associated with a trait, often using the regression summary statistics from genome-wide association (GWA) studies. Important methodological challenges are accounting for inflation in GWA effect estimates as well as investigating multiple traits simultaneously. We leverage machine learning approaches for these two challenges, developing a computationally efficient method called ML-MAGES. First, we shrink the inflation in GWA effect sizes caused by non-independence among variants using neural networks. We then cluster variant associations among multiple traits via variational inference. We compare the performance of neural network shrinkage to regularized regression and fine-mapping, which both address inflation but handle focal regions of different sizes. Our method outperforms those frameworks in approximating the true effects in simulated data. Our infinite mixture clustering approach offers a flexible, data-driven way to distinguish different types of associations—trait-specific, shared across traits, or spurious—among multiple traits based on their regularized effects. Clustering applied to our neural network shrinkage results also produces consistently higher precision and recall for distinguishing gene-level associations in simulations. We demonstrate the application of ML-MAGES on association analyses of two quantitative traits and two binary traits from the UK Biobank data. The identified genes from single-trait enrichment tests overlap with those having known relevant biological processes to the traits. Besides trait-specific associations, ML-MAGES identifies several variants with shared multi-trait associations, suggesting putative shared genetic architecture.