Robust Infidelity: When Faithfulness Measures on Masked Language Models Are Misleading
摘要
A common approach to quantifying neural text classifier interpretability is to calculate faithfulness metrics based on iteratively masking salient input tokens and measuring changes in model outputs. We propose that this property is better described as “sensitivity to iterative masking", and highlight pitfalls in using this measure for comparing text classifier interpretability. We show that iterative masking produces large variation in faithfulness scores between otherwise comparable Transformer encoder text classifiers. We further demonstrate that iteratively masked samples produce embeddings outside the distribution seen during training, resulting in unpredictable behaviour. We finally identify an underlying similarity between iterative masking and salience-based adversarial attacks that may negatively impact adversarial robustness. Our findings give insight into how masking affects neural text classifiers and provide guidance on how faithfulness measures should be interpreted.