<p>Foundation models have demonstrated immense value for scRNA-seq analysis, but their fine-tuning or inference on heterogeneous, privacy-sensitive clinical cohorts is governed by strict data protection policies, which often prohibit centralization. We introduce Clifti-GPT, a privacy-preserving federated framework based on secure multi-party computation (SMPC) that enables collaborative model training and transferable inference, where zero-shot predictions are performed across decentralized clinical repositories by securely aggregating local statistics rather than transferring data embeddings, without sharing patient data, clinical-level statistics, or models. Built upon the scGPT foundation model, Clifti-GPT achieves performance within 4% of centralized scGPT baselines in accuracy, precision, recall, and macro-F1 for cell type classification and reference mapping across six datasets. Furthermore, it demonstrates rapid convergence in terms of communication rounds, reaching 99% of centralized performance on cell type classification in at most two federated rounds on two evaluated datasets, and scales robustly to 30 clients with less than 2% accuracy loss on a large-scale federated cell type classification setting. Our analysis shows that batch effects impact both Clifti-GPT and centralized baseline, while correction leads to similar results across evaluation metrics in heterogeneous settings for both models. Together, these results indicate that Clifti-GPT enables effective fine-tuning and application of single-cell foundation models across distributed clinical datasets in a manner that is GDPR-compatible by design and addresses real-world privacy and institutional data-governance requirements.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Clifti-GPT: privacy-preserving federated fine-tuning and transferable inference of foundation models on clinical single-cell data

  • Mohammad Bakhtiari,
  • Maria Louise Elkjaer,
  • Ali Oğuz Can,
  • Fabian Theis,
  • Mhaned Oubounyt,
  • Jan Baumbach

摘要

Foundation models have demonstrated immense value for scRNA-seq analysis, but their fine-tuning or inference on heterogeneous, privacy-sensitive clinical cohorts is governed by strict data protection policies, which often prohibit centralization. We introduce Clifti-GPT, a privacy-preserving federated framework based on secure multi-party computation (SMPC) that enables collaborative model training and transferable inference, where zero-shot predictions are performed across decentralized clinical repositories by securely aggregating local statistics rather than transferring data embeddings, without sharing patient data, clinical-level statistics, or models. Built upon the scGPT foundation model, Clifti-GPT achieves performance within 4% of centralized scGPT baselines in accuracy, precision, recall, and macro-F1 for cell type classification and reference mapping across six datasets. Furthermore, it demonstrates rapid convergence in terms of communication rounds, reaching 99% of centralized performance on cell type classification in at most two federated rounds on two evaluated datasets, and scales robustly to 30 clients with less than 2% accuracy loss on a large-scale federated cell type classification setting. Our analysis shows that batch effects impact both Clifti-GPT and centralized baseline, while correction leads to similar results across evaluation metrics in heterogeneous settings for both models. Together, these results indicate that Clifti-GPT enables effective fine-tuning and application of single-cell foundation models across distributed clinical datasets in a manner that is GDPR-compatible by design and addresses real-world privacy and institutional data-governance requirements.