Background <p>Accurate histological subtyping of esophageal cancer is critical for treatment selection and prognostication. Whole-slide images (WSIs) provide rich morphological information, but they are extremely large and lack region-level annotations. We evaluated an interpretable deep learning pipeline for classifying esophageal squamous cell carcinoma (ESCC) and esophageal adenocarcinoma (EAC) using WSIs.</p> Methods <p>This retrospective study included 139 hematoxylin-and-eosin-stained WSIs from 139 patients (58 with EAC and 81 with ESCC). To differentiate EAC from ESCC, tissue tiles were generated and fixed-length embeddings were extracted using the original frozen Google ViT. During model fitting, each slide was represented by a randomly sampled bag of 128 tile embeddings and slide-level prediction was performed using an attention-gated multiple-instance learning model. We employed focal loss and weighted sampling strategies and evaluated internal performance using stratified fivefold cross-validation. The five original MIL checkpoints were applied to a balanced 50-case non-TCGA public cohort with an internal-only Platt mapping and post hoc full-cohort threshold optimization.</p> Results <p>The mean AUC was high in the training sets (0.98 ± 0.01) and internal validation folds (0.95 ± 0.05), with the best validation fold exhibiting an AUC of 0.99 and an accuracy of 96.4%. Neither normalized attention entropy (<i>p</i> = 0.364) nor maximum attention (<i>p</i> = 0.423) differed significantly between histologies. In post hoc full-cohort threshold optimization, internally calibrated external scores achieved apparent 80.0% accuracy and balanced accuracy, with adenocarcinoma sensitivity of 76.0% and ESCC sensitivity of 84.0% .</p> Conclusions <p>This model demonstrated potential as an interpretable slide-level screening approach for ESCC versus EAC classification, but the external operating threshold was domain dependent. Prospective multicenter validation using a prespecified cutoff is required.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-based multiple-instance learning enables whole-slide pathology classification of esophageal squamous cell carcinoma versus esophageal adenocarcinoma with quantitative explainability

  • Yixuan Wang,
  • Lili Yu,
  • Abiyasi Nanding,
  • Ke Jin,
  • Yin Zhang,
  • Wencheng Shao,
  • Shilong Liu

摘要

Background

Accurate histological subtyping of esophageal cancer is critical for treatment selection and prognostication. Whole-slide images (WSIs) provide rich morphological information, but they are extremely large and lack region-level annotations. We evaluated an interpretable deep learning pipeline for classifying esophageal squamous cell carcinoma (ESCC) and esophageal adenocarcinoma (EAC) using WSIs.

Methods

This retrospective study included 139 hematoxylin-and-eosin-stained WSIs from 139 patients (58 with EAC and 81 with ESCC). To differentiate EAC from ESCC, tissue tiles were generated and fixed-length embeddings were extracted using the original frozen Google ViT. During model fitting, each slide was represented by a randomly sampled bag of 128 tile embeddings and slide-level prediction was performed using an attention-gated multiple-instance learning model. We employed focal loss and weighted sampling strategies and evaluated internal performance using stratified fivefold cross-validation. The five original MIL checkpoints were applied to a balanced 50-case non-TCGA public cohort with an internal-only Platt mapping and post hoc full-cohort threshold optimization.

Results

The mean AUC was high in the training sets (0.98 ± 0.01) and internal validation folds (0.95 ± 0.05), with the best validation fold exhibiting an AUC of 0.99 and an accuracy of 96.4%. Neither normalized attention entropy (p = 0.364) nor maximum attention (p = 0.423) differed significantly between histologies. In post hoc full-cohort threshold optimization, internally calibrated external scores achieved apparent 80.0% accuracy and balanced accuracy, with adenocarcinoma sensitivity of 76.0% and ESCC sensitivity of 84.0% .

Conclusions

This model demonstrated potential as an interpretable slide-level screening approach for ESCC versus EAC classification, but the external operating threshold was domain dependent. Prospective multicenter validation using a prespecified cutoff is required.