<p>Zero-Shot Anomaly Detection (ZSAD) enables the identification of anomalies without requiring data from the target domain, meeting practical demands for rapid detection and response to unknown categories. Existing approaches rely on manually crafted, fixed textual prompts, which struggle to cope with the diversity and unpredictability of anomaly patterns. This paper proposes a novel method, V2TCASA, which dynamically generates semantically rich textual prompts based on visual features, overcoming the limitations of CLIP’s dependence on fixed manual prompts. The results demonstrate that V2TCASA achieves high-precision detection without requiring any prior knowledge of categories or anomaly states, while balancing detection accuracy and inference efficiency. The method achieves outstanding ZSAD accuracy across four benchmark anomaly detection datasets, attaining image-level and pixel-level AUROC scores of 91.4% and 90.3% on the MVTec-AD dataset, and 87.8% and 95.0% on the VisA dataset, demonstrating strong generalization and practical value.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

V2TCASA: Vision to text class-agnostic state-agnostic for industrial zero-shot anomaly detection

  • Cheng Jiang,
  • Lingxi Peng,
  • Haohuai Liu

摘要

Zero-Shot Anomaly Detection (ZSAD) enables the identification of anomalies without requiring data from the target domain, meeting practical demands for rapid detection and response to unknown categories. Existing approaches rely on manually crafted, fixed textual prompts, which struggle to cope with the diversity and unpredictability of anomaly patterns. This paper proposes a novel method, V2TCASA, which dynamically generates semantically rich textual prompts based on visual features, overcoming the limitations of CLIP’s dependence on fixed manual prompts. The results demonstrate that V2TCASA achieves high-precision detection without requiring any prior knowledge of categories or anomaly states, while balancing detection accuracy and inference efficiency. The method achieves outstanding ZSAD accuracy across four benchmark anomaly detection datasets, attaining image-level and pixel-level AUROC scores of 91.4% and 90.3% on the MVTec-AD dataset, and 87.8% and 95.0% on the VisA dataset, demonstrating strong generalization and practical value.