Background <p>Current cervical cancer risk stratification relies on FIGO stage and lymph node status, which fail to capture molecular heterogeneity and leave patients over- or under-treated. Conventional transcriptomic biomarker discovery relies solely on differential expression, yielding thousands of candidates that often lack disease specificity. To address these limitations, we developed a computational framework that integrates knowledge graph (KG) embeddings with transcriptomic differential expression to identify cervical cancer-specific prognostic biomarkers.</p> Results <p>Among ten embedding models for prediction of gene-disease associations, ComplEx showed the highest cervical cancer specificity (Top 100 mean <i>Z</i>-score = 4.19, range 3.54–6.71). Intersection of KG-predicted candidates (<i>Z</i> ≥ 1.645, <i>P</i> &lt; 0.05) with DESeq2 differentially expressed genes (adjusted <i>P</i> &lt; 0.05, |log<sub>2</sub>FoldChange|&gt; 1) identified 468 high-confidence candidate genes. Using a nested cross-validated Top-K weighted risk score approach, a 20-gene prognostic signature was finalized. The signature demonstrated an unbiased nested cross-validated C-index of 0.638 ± 0.057 for Top-K selection, while the apparent full-data C-index was 0.760 (bootstrap 95% CI 0.697–0.812) and the mean time-dependent AUC was 0.795. Multivariable Cox regression confirmed the signature as an independent prognostic factor after adjusting for age and FIGO stage (HR = 1.89 per SD, <i>P</i> &lt; 0.001, C-index = 0.803), with the strongest discrimination in Stage I disease (C-index = 0.796). Also, external validation in GSE52903 showed a consistent but non-significant trend toward survival discrimination (log-rank <i>P</i> = 0.055; HR = 1.30).</p> Conclusions <p>This study introduces a <i>Z</i>-score-based specificity framework that integrates knowledge-graph embedding with transcriptomic analysis for cervical cancer biomarker discovery. Unlike conventional KG evaluations that rely solely on link-prediction metrics, the disease-specificity <i>Z</i>-score calibrates each gene against ten control solid tumors to filter pan-cancer noise, thereby enhancing the disease specificity of the transcriptomic analysis. These results support the signature’s hypothesis-generating potential for personalized risk stratification, though prospective validation in independent cohorts with survival endpoints remains essential before clinical application.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A knowledge graph-guided transcriptomic framework identifies a 20-gene prognostic signature for cervical cancer

  • Ziyi Deng,
  • Xuehua Bi,
  • Hanyuan Zhang

摘要

Background

Current cervical cancer risk stratification relies on FIGO stage and lymph node status, which fail to capture molecular heterogeneity and leave patients over- or under-treated. Conventional transcriptomic biomarker discovery relies solely on differential expression, yielding thousands of candidates that often lack disease specificity. To address these limitations, we developed a computational framework that integrates knowledge graph (KG) embeddings with transcriptomic differential expression to identify cervical cancer-specific prognostic biomarkers.

Results

Among ten embedding models for prediction of gene-disease associations, ComplEx showed the highest cervical cancer specificity (Top 100 mean Z-score = 4.19, range 3.54–6.71). Intersection of KG-predicted candidates (Z ≥ 1.645, P < 0.05) with DESeq2 differentially expressed genes (adjusted P < 0.05, |log2FoldChange|> 1) identified 468 high-confidence candidate genes. Using a nested cross-validated Top-K weighted risk score approach, a 20-gene prognostic signature was finalized. The signature demonstrated an unbiased nested cross-validated C-index of 0.638 ± 0.057 for Top-K selection, while the apparent full-data C-index was 0.760 (bootstrap 95% CI 0.697–0.812) and the mean time-dependent AUC was 0.795. Multivariable Cox regression confirmed the signature as an independent prognostic factor after adjusting for age and FIGO stage (HR = 1.89 per SD, P < 0.001, C-index = 0.803), with the strongest discrimination in Stage I disease (C-index = 0.796). Also, external validation in GSE52903 showed a consistent but non-significant trend toward survival discrimination (log-rank P = 0.055; HR = 1.30).

Conclusions

This study introduces a Z-score-based specificity framework that integrates knowledge-graph embedding with transcriptomic analysis for cervical cancer biomarker discovery. Unlike conventional KG evaluations that rely solely on link-prediction metrics, the disease-specificity Z-score calibrates each gene against ten control solid tumors to filter pan-cancer noise, thereby enhancing the disease specificity of the transcriptomic analysis. These results support the signature’s hypothesis-generating potential for personalized risk stratification, though prospective validation in independent cohorts with survival endpoints remains essential before clinical application.