<p>The increasingly complex nature of drug discovery requires new computational approaches to reduce cost and development time. This work introduces a novel bidirectional transformer-based architecture that seamlessly maps natural-language drug indications to Simplified Molecular Input Line Entry System (SMILES) encoded molecular structures. The proposed model MedT5-Bi integrates three key contributions that collectively enhance molecular generation from textual descriptions. First, a Molecule-Aware Embeddings (MAEmb) module fuses MolEmbedder token embeddings with structural insights derived from a Graph Neural Network (GNN) to effectively capturing both sequential and topological features of chemical entities. Second, a Dynamic Attention Mechanism (DAM) adaptively switches between softmax and log-linear attention formulations based on input length and complexity, thereby maintaining performance consistency across varying sequence distributions. Third, the system is fine-tuned via reinforcement learning (RL) using a carefully designed composite reward function that jointly optimizes chemical validity, structural similarity, and fingerprint-based metrics. This RL-based training stage aligns generative outputs with desired chemical properties while improving the model’s generalization across diverse indication inputs. Evaluated on a large, augmented ChEMBL dataset. The proposed architecture outperforms existing state-of-the-art model by 16.6<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(-\)</EquationSource> </InlineEquation>24.8% on standard benchmarks including BLEU, ROUGE, Levenshtein distance, Morgan/Tanimoto similarity, and Text2Mol metrics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Medt5-bi: bidirectional translation between drug indications and molecular structures using a chemically-aware transformer

  • Soham Pahari,
  • M. Srinivas

摘要

The increasingly complex nature of drug discovery requires new computational approaches to reduce cost and development time. This work introduces a novel bidirectional transformer-based architecture that seamlessly maps natural-language drug indications to Simplified Molecular Input Line Entry System (SMILES) encoded molecular structures. The proposed model MedT5-Bi integrates three key contributions that collectively enhance molecular generation from textual descriptions. First, a Molecule-Aware Embeddings (MAEmb) module fuses MolEmbedder token embeddings with structural insights derived from a Graph Neural Network (GNN) to effectively capturing both sequential and topological features of chemical entities. Second, a Dynamic Attention Mechanism (DAM) adaptively switches between softmax and log-linear attention formulations based on input length and complexity, thereby maintaining performance consistency across varying sequence distributions. Third, the system is fine-tuned via reinforcement learning (RL) using a carefully designed composite reward function that jointly optimizes chemical validity, structural similarity, and fingerprint-based metrics. This RL-based training stage aligns generative outputs with desired chemical properties while improving the model’s generalization across diverse indication inputs. Evaluated on a large, augmented ChEMBL dataset. The proposed architecture outperforms existing state-of-the-art model by 16.6 \(-\) 24.8% on standard benchmarks including BLEU, ROUGE, Levenshtein distance, Morgan/Tanimoto similarity, and Text2Mol metrics.