<p>Railway brake shoe defects pose substantial risks to railway operational safety because subtle micro-cracks, edge wear, and surface-adhered slag may be difficult to distinguish from rust, shadows, oil stains, and other industrial noise. To address this challenge, we propose Adapter-enhanced CLIP (ACLIP), a parameter-efficient vision-language anomaly detection framework for railway brake shoe inspection. ACLIP inserts lightweight residual adapters into the late transformer layers of a frozen CLIP ViT-L/14 backbone, preserving the general vision-language knowledge of CLIP while adapting it to fine-grained industrial defect patterns. A two-stage training strategy is introduced. Stage 1 disentangles normal and anomalous text anchors, and Stage 2 aligns local visual patches with the refined semantic anchors using image-level labels only. In this work, “annotation-free” refers strictly to the training stage: ACLIP does not use pixel-level masks for optimization. The generated heatmaps are therefore treated as defect-indicative visual explanations on the private brake shoe dataset, while quantitative localization evaluation is reported only on public datasets where pixel-level masks are available for testing. Experiments are conducted on a real-world railway brake shoe dataset under stratified five-fold cross-validation and on two public industrial anomaly detection benchmarks, MVTec AD and VisA. ACLIP achieves 99.28% classification accuracy and 99.98% image-level AUROC on the brake shoe dataset and maintains 85.00% image-level AUROC in the 2-shot setting. Public benchmark evaluation reports image-level AUROC, pixel-level AUROC, and AUPRO to assess both detection and heatmap-based localization behavior. The results indicate that ACLIP improves fine-grained defect discrimination and produces more focused anomaly evidence than representative CNN-based and CLIP-based baselines under the reported evaluation protocol.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adapter-enhanced CLIP for railway brake shoe anomaly detection

  • Salman Shehzad,
  • Chunsheng Yang,
  • Yuan-Gen Wang

摘要

Railway brake shoe defects pose substantial risks to railway operational safety because subtle micro-cracks, edge wear, and surface-adhered slag may be difficult to distinguish from rust, shadows, oil stains, and other industrial noise. To address this challenge, we propose Adapter-enhanced CLIP (ACLIP), a parameter-efficient vision-language anomaly detection framework for railway brake shoe inspection. ACLIP inserts lightweight residual adapters into the late transformer layers of a frozen CLIP ViT-L/14 backbone, preserving the general vision-language knowledge of CLIP while adapting it to fine-grained industrial defect patterns. A two-stage training strategy is introduced. Stage 1 disentangles normal and anomalous text anchors, and Stage 2 aligns local visual patches with the refined semantic anchors using image-level labels only. In this work, “annotation-free” refers strictly to the training stage: ACLIP does not use pixel-level masks for optimization. The generated heatmaps are therefore treated as defect-indicative visual explanations on the private brake shoe dataset, while quantitative localization evaluation is reported only on public datasets where pixel-level masks are available for testing. Experiments are conducted on a real-world railway brake shoe dataset under stratified five-fold cross-validation and on two public industrial anomaly detection benchmarks, MVTec AD and VisA. ACLIP achieves 99.28% classification accuracy and 99.98% image-level AUROC on the brake shoe dataset and maintains 85.00% image-level AUROC in the 2-shot setting. Public benchmark evaluation reports image-level AUROC, pixel-level AUROC, and AUPRO to assess both detection and heatmap-based localization behavior. The results indicate that ACLIP improves fine-grained defect discrimination and produces more focused anomaly evidence than representative CNN-based and CLIP-based baselines under the reported evaluation protocol.