Logistic Regression with Covariate Clustering in Genome-Wide Association Interaction Studies
摘要
Logistic regression with covariates is the gold standard for detecting epistasis (statistical genetic interactions) when analysing case-control datasets from genome-wide association studies (GWAS) of diseases. Nevertheless, genome-wide interaction studies (GWAIS) are still performed without covariate correction for performance reasons as the analysis of modern GWAS datasets may lead to several weeks of computation time. However, omitting necessary covariate information causes a substantial statistical error in most studies requiring genetic ancestry adjustment via principal component analysis (PCA). Here, we present a novel approach that uses proxy covariates generated by k-means clustering in combination with contingency tables to reduce the runtime complexity of logistic regression from \(\mathcal {O}(N I)\) to \(\mathcal {O}(N + IK)\) and to minimize the statistical error to a ground truth (GT) implementation that uses per-sample covariate vectors from the PCA. By using GWAS data with 141,621 genetic markers from 3,520 German patients with inflammatory bowel disease (IBD) and 4,288 healthy controls, we demonstrated a 97-fold speed-up with two k-means clusters from PCA covariates compared to the GT implementation. At the same time, we improved the mean relative error (MRE) by more than 55 % when compared to logistic regression without covariate correction. Our developments enable logistic regression-based epistasis analysis with clustered PCA covariates for GWAS datasets on a genome-wide scale.