<p>Reliable analyte prediction from Raman spectra is critical for bioprocess monitoring, yet spectral variability across domains, such as different media or scale-up conditions, challenges model generalization. Models for analyte prediction are typically trained solely on labeled source-domain data, often limiting their performance under domain shifts. Transfer learning has been successful in other fields, but is underexplored for spectroscopic data in upstream bioprocesses. To help bridge this gap we assess whether unsupervised transfer learning methods can enhance model generalization by leveraging unlabeled target-domain spectra. We benchmark four unsupervised transfer learning approaches (CORAL, JDOT, TCA, OPP-MMD) against two bioprocess-relevant baselines: a fixed standard preprocessing workflow (REF) and an optimized source calibration based on systematic preprocessing screening (OPP). All methods are evaluated on two bioprocess Raman datasets: Dataset A comprises five glucose-spiked media and is complemented by in-silico spectral perturbations designed to mimic commonly occurring Raman noise types, such as baseline shifts, scattering artifacts, and spectral shifts. Dataset B evaluates scale-up transfer from pilot to production bioreactors for predicting glucose, biomass, and phosphate. Across the investigated cases, a consistent pattern emerged: JDOT and CORAL were most beneficial when across-domain differences dominated, whereas under high within-domain variability transfer learning provided no benefit and even underperformed the baseline. Because these transfer learning approaches can be integrated with standard Raman preprocessing pipelines, unsupervised transfer learning is best viewed as a complementary PAT tool rather than a universal replacement.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Limits and benefits of unsupervised transfer learning for Raman spectroscopy-based bioprocess monitoring

  • Bianca Hüpfl,
  • Alexandra Umprecht,
  • Bence Kozma,
  • Peter Ettinger,
  • Oliver Spadiut

摘要

Reliable analyte prediction from Raman spectra is critical for bioprocess monitoring, yet spectral variability across domains, such as different media or scale-up conditions, challenges model generalization. Models for analyte prediction are typically trained solely on labeled source-domain data, often limiting their performance under domain shifts. Transfer learning has been successful in other fields, but is underexplored for spectroscopic data in upstream bioprocesses. To help bridge this gap we assess whether unsupervised transfer learning methods can enhance model generalization by leveraging unlabeled target-domain spectra. We benchmark four unsupervised transfer learning approaches (CORAL, JDOT, TCA, OPP-MMD) against two bioprocess-relevant baselines: a fixed standard preprocessing workflow (REF) and an optimized source calibration based on systematic preprocessing screening (OPP). All methods are evaluated on two bioprocess Raman datasets: Dataset A comprises five glucose-spiked media and is complemented by in-silico spectral perturbations designed to mimic commonly occurring Raman noise types, such as baseline shifts, scattering artifacts, and spectral shifts. Dataset B evaluates scale-up transfer from pilot to production bioreactors for predicting glucose, biomass, and phosphate. Across the investigated cases, a consistent pattern emerged: JDOT and CORAL were most beneficial when across-domain differences dominated, whereas under high within-domain variability transfer learning provided no benefit and even underperformed the baseline. Because these transfer learning approaches can be integrated with standard Raman preprocessing pipelines, unsupervised transfer learning is best viewed as a complementary PAT tool rather than a universal replacement.