Molecular Representations for Drug Discovery
摘要
Boosted by the incredible advancement of machine learning algorithms, machine learning-aided drug discovery has quickly become a central topic in the field over the past decade. In this chapter, rather than focusing on algorithms, molecular representations typically seen in the field are revisited. The categories of molecular representations consist of four types of data “modalities,” sequential, topological, spatial, and temporal modalities, inspired by the categorization of primary, secondary, tertiary, and ensemble protein structures. Each modality is reviewed with examples from the literature. Moreover, knowledge graphs are discussed for representation learning, along with multimodal fusion techniques aimed at leveraging the strengths of each modality. With successful stories in other fields, it is foreseeable that by combining the right domain knowledge with the correct combination of modalities and the proper algorithm, the next breakthrough of machine learning aided drug discovery is on the horizon.