HD-CAE: hybrid encoding and deformable decoding autoencoder with cascade attention for visual anomaly detection
摘要
Autoencoder-based deep generative models leverage symmetrical encoding–decoding operations for visual anomaly detection and perform effectively across various scenarios. However, their ability to handle diverse static and dynamic irregularities is limited, often leading to the undesired generalization of abnormal visuals as normal. To address these shortcomings, we propose an asymmetrical encoder-decoder architecture featuring hybrid encoding, deformable decoding, and cascade attention for unsupervised anomaly detection. Our method emphasizes compressing comprehensive features and selectively reconstructing normal patterns. The encoder employs hybrid convolution operations to enhance spatial feature extraction, while the decoder integrates deformable operations to capture finer spatial details during reconstruction. The cascade attention mechanism refines encoding and decoding by focusing on the most relevant regions, ensuring more accurate anomaly detection. A multi-objective loss function constrains the reconstruction ability, mitigating overgeneralization and enhancing the model’s ability to differentiate between normal and abnormal visuals. This strategy enables high-fidelity reconstruction of normal visuals while disrupting abnormal pattern reconstruction. Extensive evaluations of benchmark datasets demonstrate the robustness and effectiveness of the proposed model. Our approach achieves competitive performance, with scores of 81% on Shanghai Tech, 99.1% on UCSD Ped2, and 90.1% on the CUHK Avenue dataset, highlighting its capability to detect visual irregularities across diverse environments.