Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. Therefore, ensuring the trustworthiness of LLMs emerges as an important topic. This chapter presents the TrustLLM framework (Sun et al., Trustllm: Trustworthiness in large language models. International Conference on Machine Learning (2024)), a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs. Specifically, we first introduce a set of principles for trustworthy LLMs that span eight dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. Based on the evaluation of 16 mainstream LLMs in TrustLLM (Sun et al., Trustllm: Trustworthiness in large language models. International Conference on Machine Learning (2024)), consisting of over 30 datasets, this chapter summarizes the main findings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Trustworthiness Evaluation of Large Language Models

  • Pin-Yu Chen,
  • Sijia Liu

摘要

Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. Therefore, ensuring the trustworthiness of LLMs emerges as an important topic. This chapter presents the TrustLLM framework (Sun et al., Trustllm: Trustworthiness in large language models. International Conference on Machine Learning (2024)), a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs. Specifically, we first introduce a set of principles for trustworthy LLMs that span eight dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. Based on the evaluation of 16 mainstream LLMs in TrustLLM (Sun et al., Trustllm: Trustworthiness in large language models. International Conference on Machine Learning (2024)), consisting of over 30 datasets, this chapter summarizes the main findings.