Machine learning-driven prediction of organic solar cell performance: a data-centric approach to molecular design
摘要
Organic solar cells (OSCs) offer a promising route toward flexible and sustainable energy technologies, yet predictive modeling of device parameters remains challenging due to the chemical diversity of donor–acceptor systems and morphology-dependent effects. In this work, we present the first systematic demonstration of using autoencoder-compressed molecular fingerprints with tree-based machine learning models to predict key OSC performance metrics—power conversion efficiency (PCE), open-circuit voltage (Voc), short-circuit current (Jsc), and fill factor (FF)—from a broad experimental dataset of 2500 donor–acceptor pairs, including both fullerene and non-fullerene acceptors. These compact models, trained on compressed descriptors of only 32 dimensions, achieved strong predictive accuracy (Pearson
The dataset used in this work comprises approximately 2500 experimentally characterized donor–acceptor pairs from bulk heterojunction OSCs. These include both fullerene and non-fullerene acceptor systems. For each pair, the database provides electronic descriptors, polymerization-related metrics, and the SMILES representations of the donor and acceptor molecules. Molecular fingerprints were computed from SMILES codes using the RDKit and CDK cheminformatics toolkits. A variety of machine learning models were explored, including feedforward neural networks, autoencoders for feature compression, tree-based ensemble methods, and kernel-based regression algorithms. Hyperparameter tuning was carried out using the Optuna and BayesSearchCV libraries to ensure optimal model performance.