Comparative Analysis of Six Correlation Metrics on Identifying DNA Co-Methylation Patterns
摘要
DNA methylation is an important epigenetic event associated with cancers. Different genomic sites tend to be co-methylated. It is unclear which correlation metrics should be used to study co-methylation and how they perform in identifying highly co-methylated (HCM) sites. The impact of different features of the data is also unclear, e.g., outlier, low variance, and data transformation (from B to M = logit(B)). We therefore conducted comparative analyses of six metrics, Pearson, Spearman, Kendall, Hoeffding, Distance, and Maximal Information Coefficient (MIC). Key findings are summarized below. First, the numbers of HCM sites identified by the six metrics were very different when using a fixed cutoff value. Pearson and Distance identified more HCM sites and had strong similarities. However, these metrics were susceptible to outliers and data transformation. They identified more HCM sites when using B values, but these sites tended to have outliers and lower variance. Second, Kendall’s and Hoeffding’s scores were significantly lower than those of the other metrics, leading to fewer HCM pairs being identified. Third, MIC required a large sample size to perform effectively. Although it may detect unique correlation patterns, it is difficult to interpret these patterns biologically. In summary, considering many factors together (e.g., outliers, low variance, data transformation, cutoffs, and runtime), researchers should carefully evaluate the distinct strengths and limitations of each method to select the most appropriate one for their methylation data and the goal of their analysis.