<p>Lung CT analysis requires reliable feature representations, while fully supervised training is constrained by expert annotation cost. This study presents the Focal-SE Masked Autoencoder (Focal-SE MAE, where SE denotes Squeeze-and-Excitation), a self-supervised masked reconstruction framework for lung CT representation learning. The method retains the MAE reconstruction objective and modifies the encoder to jointly model multi-scale spatial context and channel responses. The Focal-SE MAE encoder was pre-trained without external pretrained weights. For downstream transfer, the encoder was separately pretrained using only the private patient-level training split, frozen, and evaluated as a feature extractor for four-class 64-slice CT sequence classification. On LUNA16, Focal-SE MAE achieved the best reconstruction among the evaluated MAE encoder variants using the common decoder over five independent runs, with a mean absolute error of <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(1.325\times 10^{-3}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>1.325</mn> <mo>×</mo> <msup> <mn>10</mn> <mrow> <mo>-</mo> <mn>3</mn> </mrow> </msup> </mrow> </math></EquationSource> </InlineEquation>, a peak signal-to-noise ratio of 46.54 dB, and a structural similarity index measure of 0.9939. In downstream four-class 64-slice CT sequence classification on a private CT dataset, it achieved an accuracy of 92.68% and a macro-averaged area under the receiver operating characteristic curve of 97.38%. These results support joint spatial and channel dependency modeling as a useful design choice for self-supervised lung CT representation learning within the evaluated reconstruction and CT sequence classification settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Focal-SE MAE: self-supervised feature learning for lung CT reconstruction and classification

  • Xilong Kang,
  • Satoru Ikebe,
  • Shingo Mabu

摘要

Lung CT analysis requires reliable feature representations, while fully supervised training is constrained by expert annotation cost. This study presents the Focal-SE Masked Autoencoder (Focal-SE MAE, where SE denotes Squeeze-and-Excitation), a self-supervised masked reconstruction framework for lung CT representation learning. The method retains the MAE reconstruction objective and modifies the encoder to jointly model multi-scale spatial context and channel responses. The Focal-SE MAE encoder was pre-trained without external pretrained weights. For downstream transfer, the encoder was separately pretrained using only the private patient-level training split, frozen, and evaluated as a feature extractor for four-class 64-slice CT sequence classification. On LUNA16, Focal-SE MAE achieved the best reconstruction among the evaluated MAE encoder variants using the common decoder over five independent runs, with a mean absolute error of \(1.325\times 10^{-3}\) 1.325 × 10 - 3 , a peak signal-to-noise ratio of 46.54 dB, and a structural similarity index measure of 0.9939. In downstream four-class 64-slice CT sequence classification on a private CT dataset, it achieved an accuracy of 92.68% and a macro-averaged area under the receiver operating characteristic curve of 97.38%. These results support joint spatial and channel dependency modeling as a useful design choice for self-supervised lung CT representation learning within the evaluated reconstruction and CT sequence classification settings.