Focal-SE MAE: self-supervised feature learning for lung CT reconstruction and classification
摘要
Lung CT analysis requires reliable feature representations, while fully supervised training is constrained by expert annotation cost. This study presents the Focal-SE Masked Autoencoder (Focal-SE MAE, where SE denotes Squeeze-and-Excitation), a self-supervised masked reconstruction framework for lung CT representation learning. The method retains the MAE reconstruction objective and modifies the encoder to jointly model multi-scale spatial context and channel responses. The Focal-SE MAE encoder was pre-trained without external pretrained weights. For downstream transfer, the encoder was separately pretrained using only the private patient-level training split, frozen, and evaluated as a feature extractor for four-class 64-slice CT sequence classification. On LUNA16, Focal-SE MAE achieved the best reconstruction among the evaluated MAE encoder variants using the common decoder over five independent runs, with a mean absolute error of