<p>Image segmentation is critical in domains such as medical image analysis, autonomous driving, and remote sensing. While deep learning-based image segmentation systems have achieved remarkable performance, ensuring the reliability of such systems remains challenging due to the high cost of manual labeling effort, where labeling refers to annotating test inputs with pixel-level ground truth masks to determine whether model predictions are correct, which is prohibitively expensive and time-consuming. Test input prioritization offers a promising solution by identifying potentially misclassified inputs for early inspection. However, unlike in traditional deep learning, test prioritization in image segmentation presents unique challenges: 1) Image segmentation models output images rather than a single probability vector, making confidence-based uncertainty estimation unsuitable. 2) Evaluating prioritization effectiveness is more complex, as it relies on IoU, and whether a test is considered misclassified depends on an adjustable threshold, increasing assessment uncertainty. To address these challenges, we adapt complexity-based and learning-based prioritization approaches to image segmentation and conduct a comprehensive empirical study. We evaluate 26 methods across nine models and three datasets, analyzing factors such as IoU thresholds, embedding models, feature fusion techniques, and key parameters. Our study finds that: 1) Complexity-based methods outperform the baseline in 72.8% of cases, with an average APFD improvement of 0.035 (7%); 2) Learning-based methods consistently perform best across all settings, improving over complexity-based methods by 43.02%<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\sim \)</EquationSource> </InlineEquation>52.68% and over the random baseline by 52.98%; 3) DenseNet is the best-performing embedding model, achieving the best results in 56.0% of cases; 4) Concat is the most effective fusion technique, leading in 51.2% of cases; and 5) proximity-aware feature engineering further enhances performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Test input prioritization for image segmentation: an empirical study

  • Yinghua Li,
  • Xueqi Dang,
  • Wendkûuni C. Ouédraogo,
  • Jacques Klein,
  • Tegawendé F. Bissyandé

摘要

Image segmentation is critical in domains such as medical image analysis, autonomous driving, and remote sensing. While deep learning-based image segmentation systems have achieved remarkable performance, ensuring the reliability of such systems remains challenging due to the high cost of manual labeling effort, where labeling refers to annotating test inputs with pixel-level ground truth masks to determine whether model predictions are correct, which is prohibitively expensive and time-consuming. Test input prioritization offers a promising solution by identifying potentially misclassified inputs for early inspection. However, unlike in traditional deep learning, test prioritization in image segmentation presents unique challenges: 1) Image segmentation models output images rather than a single probability vector, making confidence-based uncertainty estimation unsuitable. 2) Evaluating prioritization effectiveness is more complex, as it relies on IoU, and whether a test is considered misclassified depends on an adjustable threshold, increasing assessment uncertainty. To address these challenges, we adapt complexity-based and learning-based prioritization approaches to image segmentation and conduct a comprehensive empirical study. We evaluate 26 methods across nine models and three datasets, analyzing factors such as IoU thresholds, embedding models, feature fusion techniques, and key parameters. Our study finds that: 1) Complexity-based methods outperform the baseline in 72.8% of cases, with an average APFD improvement of 0.035 (7%); 2) Learning-based methods consistently perform best across all settings, improving over complexity-based methods by 43.02% \(\sim \) 52.68% and over the random baseline by 52.98%; 3) DenseNet is the best-performing embedding model, achieving the best results in 56.0% of cases; 4) Concat is the most effective fusion technique, leading in 51.2% of cases; and 5) proximity-aware feature engineering further enhances performance.