<p>Oral cancer is associated with substantial morbidity when detected at advanced stages, while image-based screening remains challenging because oral photographs show large variability in lesion appearance, anatomical site, illumination, framing, and acquisition conditions. This study proposes a Hierarchical Lesion-Aware Transformer (HLAT) for binary oral cancer image classification as a screening-oriented distinction between cancer and non-cancer images. The model integrates a Lesion-Preserving Stem, hierarchical Lesion-aware Recalibration Transformer blocks, Cross-Window Depthwise Convolution, Multi-Scale Gated ConvFFN modules, lesion-aware recalibration, and a fusion head designed to preserve local lesion morphology while capturing broader contextual information. Experiments were conducted on a curated public dataset of 6267 oral images from Kaggle, Zenodo, and Mendeley, comprising 3267 cancer and 3000 non-cancer images. After duplicate and invalid image removal, the dataset was evaluated using stratified tenfold cross-validation, an internal held-out test set, expanded baseline comparison, source-held-out validation, calibration analysis, ablation testing, and computational-efficiency assessment. HLAT achieved 99.14 ± 0.09% mean validation accuracy and 99.15% accuracy on the internal held-out test set, with sensitivity of 98.78%, specificity of 99.56%, F1-score of 99.18%, MCC of 0.983, Brier score of 0.012, and ECE of 0.009. In source-held-out evaluation, HLAT obtained a mean accuracy of 93.64 ± 0.63%, indicating reduced but still consistent performance under repository-level domain shift. The model required 24.7&#xa0;M parameters, 3.90 GFLOPs, and 0.310 ± 0.011&#xa0;ms forward-pass latency per image. These findings suggest strong internal performance and potential decision-support value for screening-oriented oral image analysis; however, external multicenter and prospective validation are required before clinical deployment or diagnostic use.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical lesion-aware transformer for oral cancer image classification

  • Chandan Karmakar,
  • Md Anwar Hossain,
  • Kallol Chakraborty Shekhor,
  • Md Sahid Hossain,
  • Md Abedur Rahman,
  • Md Sharifur Rahman,
  • Girigula Durga Bhavani,
  • Chala Wata

摘要

Oral cancer is associated with substantial morbidity when detected at advanced stages, while image-based screening remains challenging because oral photographs show large variability in lesion appearance, anatomical site, illumination, framing, and acquisition conditions. This study proposes a Hierarchical Lesion-Aware Transformer (HLAT) for binary oral cancer image classification as a screening-oriented distinction between cancer and non-cancer images. The model integrates a Lesion-Preserving Stem, hierarchical Lesion-aware Recalibration Transformer blocks, Cross-Window Depthwise Convolution, Multi-Scale Gated ConvFFN modules, lesion-aware recalibration, and a fusion head designed to preserve local lesion morphology while capturing broader contextual information. Experiments were conducted on a curated public dataset of 6267 oral images from Kaggle, Zenodo, and Mendeley, comprising 3267 cancer and 3000 non-cancer images. After duplicate and invalid image removal, the dataset was evaluated using stratified tenfold cross-validation, an internal held-out test set, expanded baseline comparison, source-held-out validation, calibration analysis, ablation testing, and computational-efficiency assessment. HLAT achieved 99.14 ± 0.09% mean validation accuracy and 99.15% accuracy on the internal held-out test set, with sensitivity of 98.78%, specificity of 99.56%, F1-score of 99.18%, MCC of 0.983, Brier score of 0.012, and ECE of 0.009. In source-held-out evaluation, HLAT obtained a mean accuracy of 93.64 ± 0.63%, indicating reduced but still consistent performance under repository-level domain shift. The model required 24.7 M parameters, 3.90 GFLOPs, and 0.310 ± 0.011 ms forward-pass latency per image. These findings suggest strong internal performance and potential decision-support value for screening-oriented oral image analysis; however, external multicenter and prospective validation are required before clinical deployment or diagnostic use.