MAS-ZSAS: A Zero-Shot Anomaly Segmentation Framework with Multi-attribute Guided Text Prompts
摘要
In zero-shot anomaly segmentation (ZSAS) tasks, CLIP aligns textual prompts with visual features, enabling the identification of anomalies in unseen categories, making it a pivotal tool for addressing complex challenges in anomaly detection. However, existing approaches often rely on binary status prompts (e.g., “good/damaged”), which are insufficient to describe diverse and fine-grained anomaly characteristics such as shape, texture, or scale—especially in complex industrial scenarios. In response, we introduce MAS-ZSAS, an innovative framework that constructs a multi-dimensional attribute space to generate semantically rich prompts for enhanced anomaly perception and segmentation. MAS-ZSAS comprises three key components: 1) Fore-MAS: Generates attribute-aware prompts and integrates global image features to improve generalization to unseen anomalies. 2) AAM: Learns intrinsic dependencies among attributes through attention-based association modeling. 3) After-MAS: Employs multi-head cross-attention to refine cross-modal alignment in localized regions. MAS-ZSAS achieves state-of-the-art results on the MVTec-AD and VisA benchmarks, with 97.9% AUROC and 92.6% AUPRO in anomaly segmentation.