Background <p>Current methods for predicting gene expression from histone modifications rely on arbitrary binary classification thresholds such as the median to distinguish between high and low expression for genes. This approach lacks biological justification, creates dataset-dependent classifications, and ignores the relative regulatory relationships between genes that are often more biologically meaningful than absolute cutoffs.</p> Methods <p>We developed a deep learning method using pairwise ranking to determine the relative gene expression levels between gene pairs based on their histone modification signals. The approach adopted the DirectRanker architecture, which enforces antisymmetry by construction through a twin-subnet design with asymmetric subtraction, and was evaluated against three baseline classifiers (Random Forest, Logistic Regression, and SVM Linear) on identical features and data splits. Ablation testing across all 31 non-empty subsets of five histone marks (H3K4me3, H3K9ac, H3K9me3, H3K27ac, H3K27me3) was conducted to identify their individual contributions. Datasets were strictly partitioned (80% training, 10% validation, 10% testing) with gene-level separation to prevent data leakage, and all experiments were repeated across five independent random seeds. Four datasets were analysed: two normal adult liver cell datasets (Donor&#xa0;3 and Donor&#xa0;4 from NCBI GEO GSE19465) and two HepG2 hepatocellular carcinoma cell line datasets (ENCSR134DWG and GSE76344).</p> Results <p>Multi-mark combinations anchored by active histone marks (H3K27ac, H3K4me3, H3K9ac) consistently drove predictive performance, with the best combinations achieving AUROC values of 0.833−0.867 and test accuracies of 74–78% across datasets. DirectRanker attained AUPRC values of 0.819−0.856, closely matching or exceeding Random Forest on precision–recall performance. Repressive marks (H3K9me3, H3K27me3) consistently underperformed when used in isolation, with single-mark models achieving test accuracies of only 50–66%. DirectRanker produced substantially higher antisymmetry scores (0.961−0.983) than all baseline classifiers (0.787−0.851), confirming logically consistent pairwise predictions. Performance plateaued at four- to five-mark combinations, reflecting correlated predictive signal among the active marks within the promoter-proximal feature space.</p> Conclusions <p>The pairwise ranking framework provides a principled alternative to threshold-based classification, enabling relative gene expression comparisons without arbitrary binarisation cutoffs. While Random Forest achieved marginally higher accuracy, DirectRanker’s structural antisymmetry guarantee makes it the more suitable choice for applications requiring a globally coherent gene ranking. Active promoter marks carry the majority of the predictive signal, while the performance saturation with increasing mark combinations reflects overlapping information content rather than biological redundancy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A deep learning framework for threshold-free relative gene expression ranking from histone modification signals

  • Abdul Munif,
  • Amitava Datta,
  • Zhaoyu Li,
  • Max Ward

摘要

Background

Current methods for predicting gene expression from histone modifications rely on arbitrary binary classification thresholds such as the median to distinguish between high and low expression for genes. This approach lacks biological justification, creates dataset-dependent classifications, and ignores the relative regulatory relationships between genes that are often more biologically meaningful than absolute cutoffs.

Methods

We developed a deep learning method using pairwise ranking to determine the relative gene expression levels between gene pairs based on their histone modification signals. The approach adopted the DirectRanker architecture, which enforces antisymmetry by construction through a twin-subnet design with asymmetric subtraction, and was evaluated against three baseline classifiers (Random Forest, Logistic Regression, and SVM Linear) on identical features and data splits. Ablation testing across all 31 non-empty subsets of five histone marks (H3K4me3, H3K9ac, H3K9me3, H3K27ac, H3K27me3) was conducted to identify their individual contributions. Datasets were strictly partitioned (80% training, 10% validation, 10% testing) with gene-level separation to prevent data leakage, and all experiments were repeated across five independent random seeds. Four datasets were analysed: two normal adult liver cell datasets (Donor 3 and Donor 4 from NCBI GEO GSE19465) and two HepG2 hepatocellular carcinoma cell line datasets (ENCSR134DWG and GSE76344).

Results

Multi-mark combinations anchored by active histone marks (H3K27ac, H3K4me3, H3K9ac) consistently drove predictive performance, with the best combinations achieving AUROC values of 0.833−0.867 and test accuracies of 74–78% across datasets. DirectRanker attained AUPRC values of 0.819−0.856, closely matching or exceeding Random Forest on precision–recall performance. Repressive marks (H3K9me3, H3K27me3) consistently underperformed when used in isolation, with single-mark models achieving test accuracies of only 50–66%. DirectRanker produced substantially higher antisymmetry scores (0.961−0.983) than all baseline classifiers (0.787−0.851), confirming logically consistent pairwise predictions. Performance plateaued at four- to five-mark combinations, reflecting correlated predictive signal among the active marks within the promoter-proximal feature space.

Conclusions

The pairwise ranking framework provides a principled alternative to threshold-based classification, enabling relative gene expression comparisons without arbitrary binarisation cutoffs. While Random Forest achieved marginally higher accuracy, DirectRanker’s structural antisymmetry guarantee makes it the more suitable choice for applications requiring a globally coherent gene ranking. Active promoter marks carry the majority of the predictive signal, while the performance saturation with increasing mark combinations reflects overlapping information content rather than biological redundancy.