Background <p>Cancer genomes contain many mutations, but only a subset drive tumor development. Accurately pinpointing these driver variants remains challenging. We aim to build an accurate and interpretable model by combining DNA sequence, protein 3D structure, and cancer omics data.</p> Methods <p>We present ModVAR, a multimodal model that integrates DNA sequences, predicted protein tertiary structures, and cancer omics data to classify driver variants. The approach uses pre-trained models (DNAbert2 and ESMFold) and a self-supervised strategy for cancer omics profiles. We evaluate performance on clinically and experimentally validated driver variants with standard classification metrics, examine therapeutic relevance through molecular docking, assess modeling of variants in intrinsically disordered protein regions, and analyze modality contributions.</p> Results <p>Here we show that ModVAR achieves strong accuracy across benchmarks for identifying validated driver variants. It prioritizes variants with potential therapeutic actionability supported by docking analyses, and the inclusion of structural predictions enables effective modeling of variants in intrinsically disordered regions. Interpretation indicates that the protein structure modality contributes most to predictions. At scale, the method produces 3,971,946 publicly available variant annotations.</p> Conclusions <p>ModVAR integrates sequence, structure, and cancer omics signals to aid driver-variant discovery. It provides robust performance across tasks, supports hypothesis generation and target discovery, and supplies a large-scale resource that advances cancer research and personalized therapy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multimodal framework for comprehensive driver variant prediction in cancer

  • Hai Yang,
  • Yijia Chen,
  • Tianyi Zhou,
  • Yingzhuo Wang,
  • Qin Zhou,
  • Ting Xiao,
  • Qian Zhang,
  • Jing Zhang,
  • Dongdong Li,
  • Zhe Wang

摘要

Background

Cancer genomes contain many mutations, but only a subset drive tumor development. Accurately pinpointing these driver variants remains challenging. We aim to build an accurate and interpretable model by combining DNA sequence, protein 3D structure, and cancer omics data.

Methods

We present ModVAR, a multimodal model that integrates DNA sequences, predicted protein tertiary structures, and cancer omics data to classify driver variants. The approach uses pre-trained models (DNAbert2 and ESMFold) and a self-supervised strategy for cancer omics profiles. We evaluate performance on clinically and experimentally validated driver variants with standard classification metrics, examine therapeutic relevance through molecular docking, assess modeling of variants in intrinsically disordered protein regions, and analyze modality contributions.

Results

Here we show that ModVAR achieves strong accuracy across benchmarks for identifying validated driver variants. It prioritizes variants with potential therapeutic actionability supported by docking analyses, and the inclusion of structural predictions enables effective modeling of variants in intrinsically disordered regions. Interpretation indicates that the protein structure modality contributes most to predictions. At scale, the method produces 3,971,946 publicly available variant annotations.

Conclusions

ModVAR integrates sequence, structure, and cancer omics signals to aid driver-variant discovery. It provides robust performance across tasks, supports hypothesis generation and target discovery, and supplies a large-scale resource that advances cancer research and personalized therapy.