<p>To advance precision medicine in pathology, artificial intelligence (AI)-driven foundation models must generalize across diverse datasets, tissues, and clinical tasks. However, their comparative performance and generalizability in computational pathology remain incompletely characterized. Here, we benchmark 32 AI foundation models across four categories, including general vision models (VM), general vision-language models (VLM), pathology-specific vision models (Path-VM), and pathology-specific vision-language models (Path-VLM), using slide- and patch-level tasks from The Cancer Genome Atlas (TCGA), Clinical Proteomic Tumor Analysis Consortium (CPTAC), external benchmarking datasets, and out-of-domain datasets. Across TCGA tasks, Path-VMs consistently rank among the strongest performers. Evaluation across CPTAC and out-of-domain datasets reveals more nuanced generalization behavior, with model rankings showing modest but consistent shifts across datasets and task categories. Pairwise statistical comparisons indicate that differences among top-performing models are often small and task dependent. Path-VMs outperform Path-VLMs and remain competitive with VMs. Model size and pretraining dataset scale do not consistently predict downstream performance. Finally, late decision-level ensembling improves aggregate performance across external datasets and tissue types, highlighting complementary strengths across foundation models. <i>PathBench:</i> <a href="https://pathbench.stanford.edu/">https://pathbench.stanford.edu/</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A benchmark study of vision and pathology foundation models for computational pathology

  • Rohan Bareja,
  • Francisco Carrillo-Perez,
  • Yuanning Zheng,
  • Marija Pizurica,
  • Tarak Nath Nandi,
  • Lu Tian,
  • Jeanne Shen,
  • Ravi Madduri,
  • Olivier Gevaert

摘要

To advance precision medicine in pathology, artificial intelligence (AI)-driven foundation models must generalize across diverse datasets, tissues, and clinical tasks. However, their comparative performance and generalizability in computational pathology remain incompletely characterized. Here, we benchmark 32 AI foundation models across four categories, including general vision models (VM), general vision-language models (VLM), pathology-specific vision models (Path-VM), and pathology-specific vision-language models (Path-VLM), using slide- and patch-level tasks from The Cancer Genome Atlas (TCGA), Clinical Proteomic Tumor Analysis Consortium (CPTAC), external benchmarking datasets, and out-of-domain datasets. Across TCGA tasks, Path-VMs consistently rank among the strongest performers. Evaluation across CPTAC and out-of-domain datasets reveals more nuanced generalization behavior, with model rankings showing modest but consistent shifts across datasets and task categories. Pairwise statistical comparisons indicate that differences among top-performing models are often small and task dependent. Path-VMs outperform Path-VLMs and remain competitive with VMs. Model size and pretraining dataset scale do not consistently predict downstream performance. Finally, late decision-level ensembling improves aggregate performance across external datasets and tissue types, highlighting complementary strengths across foundation models. PathBench: https://pathbench.stanford.edu/