The Determinant of a Linkage Disequilibrium Matrix Can Identify Disease Genes in Low-Density Genome-Wide Data of Single-Nucleotide Polymorphisms
摘要
Introduction. Methods for partitioning single-nucleotide polymorphisms (SNPs) into blocks followed by haplotype analysis are among the approaches increasing the power of genome-wide association studies (GWASs). However, its applicability and coverage of genomic loci depend on the correlation between SNPs, determined by their density and location. We investigated SNP block partitioning in genomic data represented by a limited number of markers using our previously proposed method based on the determinant of the linkage disequilibrium (LD) matrix to identify loci associated with ischemic stroke (IS). Materials and methods. Genotypic data from patients with IS (N = 923) and the control group (N = 305) were analyzed for 67.925 SNPs. The LD matrix determinant was used to group SNPs and assess their association within a block. Blocks with recovered haplotypes were tested for association with IS using the χ2 test for independence. Functional analysis of the identified candidate genes for IS was performed using the DAVID online service. Results. SNPs were divided into blocks of correlated SNPs using the LD matrix determinant. It was found that the maximum number of blocks, with the greatest variability in their sizes, was observed at similar values of the determinant’s rounding threshold to zero (ε = 0.001) in data with low and high SNP density (883.908 SNPs). Eight blocks associated with IS were identified. These blocks contained 102 genomic loci, three of which—IGLC3, IGLC6, and IGLC7—were overrepresented in the immunoglobin complex of Gene Ontology. Conclusions. It was found that the determinant of LD matrix as a measure of SNP connectivity allows one to identify SNP blocks in genomic data with both high and low SNP density, which expands the scope of applicability of GWAS based on the use of haplotype tests.