Multi-instance Learning (MIL) has become a mainstream method for pathological image classification. Existing studies usually evaluate the interpretability of the model by visualizing image patches with high tumor probability through heat maps after model classification and verifying consistency with the real tumor area. However, such methods are post-analysis and fail to integrate interpretability into the model building process itself, resulting in differences between the model behavior and the actual diagnostic logic of pathologists. To address this problem, we propose a new MIL framework that simulates the pathologist’s diagnostic process and achieves intrinsic interpretability through the inherent design of the model architecture. We propose a key region selection strategy based on the Top-K attention to obtain information about the key region where the tumor is located; we also propose an innovative instance grouping dynamic mask strategy to achieve high-quality instance refinement for key regions. This coarse-to-fine hierarchical analysis not only makes the model decision process interpretable, but also generates more discriminative bag feature representation for classification tasks. We tested the model on two datasets and the results showed that our proposed model outperformed the mainstream methods. At the same time, the visualization analysis confirmed that the tumor region located by the model was highly consistent with the diagnostic basis of the pathologist, verifying the clinical interpretability of the method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

KSIR-MIL: Key Region Selection and Instance Refinement for Multi-instance Learning in Whole Slide Image Classification

  • Shaoguo Cui,
  • Jiangfeng Wu,
  • Binbin Sang,
  • Tiansong Li,
  • Yi Zhang,
  • Fumin Cheng,
  • Guofen Wang

摘要

Multi-instance Learning (MIL) has become a mainstream method for pathological image classification. Existing studies usually evaluate the interpretability of the model by visualizing image patches with high tumor probability through heat maps after model classification and verifying consistency with the real tumor area. However, such methods are post-analysis and fail to integrate interpretability into the model building process itself, resulting in differences between the model behavior and the actual diagnostic logic of pathologists. To address this problem, we propose a new MIL framework that simulates the pathologist’s diagnostic process and achieves intrinsic interpretability through the inherent design of the model architecture. We propose a key region selection strategy based on the Top-K attention to obtain information about the key region where the tumor is located; we also propose an innovative instance grouping dynamic mask strategy to achieve high-quality instance refinement for key regions. This coarse-to-fine hierarchical analysis not only makes the model decision process interpretable, but also generates more discriminative bag feature representation for classification tasks. We tested the model on two datasets and the results showed that our proposed model outperformed the mainstream methods. At the same time, the visualization analysis confirmed that the tumor region located by the model was highly consistent with the diagnostic basis of the pathologist, verifying the clinical interpretability of the method.