<p>Large language models (LLMs) can produce harmful content when prompted with unsafe inputs, yet inference-time safety interventions often degrade general-purpose performance through over-refusal. We propose SURE (Semi-supervised Uncertainty-weighted Representation-based External detection), a framework that trains a lightweight classifier on the LLM’s hidden state representations to detect harmful inputs entirely outside the inference pipeline, preserving the model’s original capabilities. To overcome annotation scarcity, SURE queries the target LLM to assess unlabeled inputs and weights each pseudo-label by the classifier’s confidence, the LLM’s uncertainty, measured as the perplexity of its single-token safety self-assessment, and their agreement. Experiments across three open-source LLMs (8B–70B parameters) show that SURE achieves average harmonic mean scores of 88–90% with only 80 labeled samples, surpassing both inference-time and representation-based baselines. Performance plateaus at 40 labeled examples, and cross-category generalization confirms over 90% accuracy on held-out categories, indicating that the learned representations capture general safety features.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SURE: data-efficient safety guardrailing via internal representations and uncertainty-weighted pseudo-labels

  • Hanwen Li,
  • Jinhao Duan,
  • Chenxi Yuan,
  • James Diffenderfer,
  • Sandeep Madireddy,
  • Bhavya Kailkhura,
  • Kaidi Xu

摘要

Large language models (LLMs) can produce harmful content when prompted with unsafe inputs, yet inference-time safety interventions often degrade general-purpose performance through over-refusal. We propose SURE (Semi-supervised Uncertainty-weighted Representation-based External detection), a framework that trains a lightweight classifier on the LLM’s hidden state representations to detect harmful inputs entirely outside the inference pipeline, preserving the model’s original capabilities. To overcome annotation scarcity, SURE queries the target LLM to assess unlabeled inputs and weights each pseudo-label by the classifier’s confidence, the LLM’s uncertainty, measured as the perplexity of its single-token safety self-assessment, and their agreement. Experiments across three open-source LLMs (8B–70B parameters) show that SURE achieves average harmonic mean scores of 88–90% with only 80 labeled samples, surpassing both inference-time and representation-based baselines. Performance plateaus at 40 labeled examples, and cross-category generalization confirms over 90% accuracy on held-out categories, indicating that the learned representations capture general safety features.