In recent decades, many researchers have directed their focus towards confidence calibration of deep neural networks because trustworthiness is equally important as accuracy, especially in high-stake scenarios within real-world applications. Unlike methods that involve modifying the training process, post-hoc calibration methods establish excellent performance by calibrating the model on the calibration set without fine-tuning the trained classifiers. Consequently, they are widely used in real-world applications. However, most post-hoc calibration methods assume that the distribution of calibration data is consistent with that of the test data and previous researches always conduct experiments on a fixed calibration set, while the situation in real-world applications is often different. Furthermore, the influence of the factors on calibration set for post-hoc calibration remains unclear. Therefore, a systematic investigation targeting the guidance for the deployment of different post-hoc calibration paradigms is required. In this study, we conduct experiments to evaluate the performance of post-hoc calibration paradigms under different settings on the calibration set and the effectiveness of different augmentation strategies under subpopulation shift.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Rethinking the Reliability of Post-hoc Calibration Methods Under Subpopulation Shift

  • Yan Zhu,
  • Huan Ma,
  • Changqing Zhang,
  • Bingzhe Wu,
  • Huazhu Fu,
  • Joey Tianyi Zhou,
  • Qinghua Hu

摘要

In recent decades, many researchers have directed their focus towards confidence calibration of deep neural networks because trustworthiness is equally important as accuracy, especially in high-stake scenarios within real-world applications. Unlike methods that involve modifying the training process, post-hoc calibration methods establish excellent performance by calibrating the model on the calibration set without fine-tuning the trained classifiers. Consequently, they are widely used in real-world applications. However, most post-hoc calibration methods assume that the distribution of calibration data is consistent with that of the test data and previous researches always conduct experiments on a fixed calibration set, while the situation in real-world applications is often different. Furthermore, the influence of the factors on calibration set for post-hoc calibration remains unclear. Therefore, a systematic investigation targeting the guidance for the deployment of different post-hoc calibration paradigms is required. In this study, we conduct experiments to evaluate the performance of post-hoc calibration paradigms under different settings on the calibration set and the effectiveness of different augmentation strategies under subpopulation shift.