Logistic regression with covariates is the gold standard for detecting epistasis (statistical genetic interactions) when analysing case-control datasets from genome-wide association studies (GWAS) of diseases. Nevertheless, genome-wide interaction studies (GWAIS) are still performed without covariate correction for performance reasons as the analysis of modern GWAS datasets may lead to several weeks of computation time. However, omitting necessary covariate information causes a substantial statistical error in most studies requiring genetic ancestry adjustment via principal component analysis (PCA). Here, we present a novel approach that uses proxy covariates generated by k-means clustering in combination with contingency tables to reduce the runtime complexity of logistic regression from \(\mathcal {O}(N I)\) to \(\mathcal {O}(N + IK)\) and to minimize the statistical error to a ground truth (GT) implementation that uses per-sample covariate vectors from the PCA. By using GWAS data with 141,621 genetic markers from 3,520 German patients with inflammatory bowel disease (IBD) and 4,288 healthy controls, we demonstrated a 97-fold speed-up with two k-means clusters from PCA covariates compared to the GT implementation. At the same time, we improved the mean relative error (MRE) by more than 55 % when compared to logistic regression without covariate correction. Our developments enable logistic regression-based epistasis analysis with clustered PCA covariates for GWAS datasets on a genome-wide scale.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Logistic Regression with Covariate Clustering in Genome-Wide Association Interaction Studies

  • Volker Neff,
  • Lars Wienbrandt,
  • David Ellinghaus

摘要

Logistic regression with covariates is the gold standard for detecting epistasis (statistical genetic interactions) when analysing case-control datasets from genome-wide association studies (GWAS) of diseases. Nevertheless, genome-wide interaction studies (GWAIS) are still performed without covariate correction for performance reasons as the analysis of modern GWAS datasets may lead to several weeks of computation time. However, omitting necessary covariate information causes a substantial statistical error in most studies requiring genetic ancestry adjustment via principal component analysis (PCA). Here, we present a novel approach that uses proxy covariates generated by k-means clustering in combination with contingency tables to reduce the runtime complexity of logistic regression from \(\mathcal {O}(N I)\) to \(\mathcal {O}(N + IK)\) and to minimize the statistical error to a ground truth (GT) implementation that uses per-sample covariate vectors from the PCA. By using GWAS data with 141,621 genetic markers from 3,520 German patients with inflammatory bowel disease (IBD) and 4,288 healthy controls, we demonstrated a 97-fold speed-up with two k-means clusters from PCA covariates compared to the GT implementation. At the same time, we improved the mean relative error (MRE) by more than 55 % when compared to logistic regression without covariate correction. Our developments enable logistic regression-based epistasis analysis with clustered PCA covariates for GWAS datasets on a genome-wide scale.