<p>Accurate multi-label detection of thoracic abnormalities in chest X-rays is crucial for rapid and reliable clinical decision-making. However, high-performing large-scale vision transformers, such as Swin-Large, suffer from substantial computational cost and long inference times, limiting their feasibility in real-time and resource-constrained healthcare settings. Conversely, lightweight architectures offer fast inference but often lack the accuracy required for detecting multiple abnormalities simultaneously in complex medical data. To address this gap, we propose an integrated framework that combines Knowledge Distillation from a high-capacity Swin-Large teacher to four efficient student models (EfficientViT-B0, MobileViT-S, EdgeNeXt-Small, and GCViT-XXTiny), together with advanced ensemble fusion strategies and an XGBoost-based meta-learner. Knowledge distillation improved recall and F1-score in most students compared to their non-distilled counterparts, though large architectural capacity differences limited performance gains beyond the teacher. Nevertheless, the ensemble methods achieved substantial improvements: Einstein fusion delivered the highest recall (0.8744), representing over a 38% increase compared to the teacher; Minkowski fusion achieved the highest F1-score (0.6786); and Logit-Softmax fusion attained the highest ROC-AUC (0.8878), surpassing both the teacher and prior works. Computationally, distilled students required only 0.002–0.005 s per image, while ensemble inference (0.016 s) was over 46% faster than the teacher (0.030 s), preserving real-time feasibility. This balanced approach between high accuracy and low latency significantly narrows the gap between lightweight models and heavy transformers, making it a practical and effective solution for multi-label medical image analysis in time-critical and resource-limited clinical workflows.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Knowledge-distilled lightweight transformers for multi-label chest X-ray detection

  • Hamid Reza Ghaedsharaf,
  • Morteza Ebrahimi

摘要

Accurate multi-label detection of thoracic abnormalities in chest X-rays is crucial for rapid and reliable clinical decision-making. However, high-performing large-scale vision transformers, such as Swin-Large, suffer from substantial computational cost and long inference times, limiting their feasibility in real-time and resource-constrained healthcare settings. Conversely, lightweight architectures offer fast inference but often lack the accuracy required for detecting multiple abnormalities simultaneously in complex medical data. To address this gap, we propose an integrated framework that combines Knowledge Distillation from a high-capacity Swin-Large teacher to four efficient student models (EfficientViT-B0, MobileViT-S, EdgeNeXt-Small, and GCViT-XXTiny), together with advanced ensemble fusion strategies and an XGBoost-based meta-learner. Knowledge distillation improved recall and F1-score in most students compared to their non-distilled counterparts, though large architectural capacity differences limited performance gains beyond the teacher. Nevertheless, the ensemble methods achieved substantial improvements: Einstein fusion delivered the highest recall (0.8744), representing over a 38% increase compared to the teacher; Minkowski fusion achieved the highest F1-score (0.6786); and Logit-Softmax fusion attained the highest ROC-AUC (0.8878), surpassing both the teacher and prior works. Computationally, distilled students required only 0.002–0.005 s per image, while ensemble inference (0.016 s) was over 46% faster than the teacher (0.030 s), preserving real-time feasibility. This balanced approach between high accuracy and low latency significantly narrows the gap between lightweight models and heavy transformers, making it a practical and effective solution for multi-label medical image analysis in time-critical and resource-limited clinical workflows.