ViMILNet: Vision Transformer and MIL-Based Framework with MFMS Attention for Breast Cancer Diagnosis
摘要
Accurate classification of breast‑cancer histopathological images is indispensable for timely diagnosis and effective treatment, yet existing methods often falter when confronted with multi‑scale variability and fail to focus on diagnostically critical regions. To overcome these limitations, we present ViMILNet, a deep‑learning framework that couples an EfficientNet‑V2 backbone with Vision Transformers and augments them with a Multi‑Frequency Multi‑Scale (MFMS) attention module and a Multi‑Instance Learning (MIL) Transformer Aggregator. The network analyses high‑resolution slides at four magnifications (40×, 100×, 200×, and 400×), employing extensive data augmentation and patch‑wise instance modelling to improve generalisation and fine‑grained feature capture. MFMS fuses spatial and frequency cues across multiple scales, whereas the MIL component adaptively aggregates the most informative patches into a unified representation, yielding robust predictions. Experiments on the BreakHis dataset demonstrate that ViMILNet consistently surpasses conventional CNNs and recent Transformer‑based baselines, attaining accuracies above 98.3% and AUCs exceeding 99.3% at every magnification. This work contributes not only to the advancement of intelligent medical image analysis but also provides a clinically relevant, scalable solution that holds potential to support digital pathology workflows and improve healthcare delivery at scale.