Generative deep learning in bioinformatics: a microservices-driven paradigm for big data environment
摘要
The exponential growth of bioinformatics and drug discovery data necessitates computational frameworks capable of efficiently managing large-scale, heterogeneous, and continuously evolving datasets. This study presents a hybrid deep learning framework built upon a microservices-based architecture to enhance scalability, modularity, and efficiency in bioinformatics workflows. The proposed system integrates generative deep learning—specifically, a variational autoencoder (VAE)—to tackle challenges in bioinformatics. Through a microservices-driven design, the framework enables modular deployment, semantic interoperability, and scalable integration of generative models into bioinformatics pipelines. The architecture is model-agnostic and supports flexible validation by encapsulating cheminformatics evaluations, as independent microservices. This allows automated updates and seamless incorporation of new evaluation metrics without disrupting the overall pipeline. Rather than replicating existing cheminformatics validation efforts, the framework provides an extensible foundation for integrating, reusing, and scaling such studies across large datasets in bioinformatics and drug discovery. In contrast to traditional monolithic architectures, the microservices paradigm supports independent deployment, optimization, and scaling of deep learning components—such as molecule generation, toxicity prediction, and pharmacokinetics analysis—while maintaining uninterrupted workflow execution. This design ensures real-time data processing, interoperability across distributed systems, and reduced computational overhead. Experiments conducted on benchmark datasets from DrugBank, ChEMBL, and TCGA demonstrate superior performance compared to monolithic baselines in predictive accuracy, biomarker identification, and computational efficiency. These results underscore the potential of microservices as a software engineering paradigm for advancing bioinformatics and accelerating the drug discovery process.