<p>AI ethics refers to the moral principles and guidelines governing the development and deployment of artificial intelligence systems, ensuring they align with human values and societal well-being. It encompasses the evaluation of AI outputs for fairness, safety, transparency, and respect for human rights. To advance systematic ethical evaluation, we introduce the EthicsLens dataset, comprising 38,808 responses generated by seven large language models. These responses were generated using diverse prompts designed to elicit appropriate and potentially sensitive responses. Each response was then annotated across sixteen ethical categories, including stereotyping, toxicity, misinformation, hate speech, harmful advice, privacy violations, political bias, false confidence, emotional or religious insensitivity, sexual content, manipulation, and impersonation. To classify ethical and unethical AI-generated content, the dataset is analysed using state-of-the-art classification methods, assessing its ability to support reliable ethical evaluation. Performance is reported both for binary ethical classification and multilabel violation identification. Results include accuracies of nearly 99% for binary classification tasks with SVM and CNN models, and macro-F1 scores of about 96% on multilabel tasks for Sentence-BERT transformer model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dataset-centric AI ethics classification

  • Aditya Kartik,
  • Surya Raj,
  • Akash Rattan,
  • Deepti Sahu

摘要

AI ethics refers to the moral principles and guidelines governing the development and deployment of artificial intelligence systems, ensuring they align with human values and societal well-being. It encompasses the evaluation of AI outputs for fairness, safety, transparency, and respect for human rights. To advance systematic ethical evaluation, we introduce the EthicsLens dataset, comprising 38,808 responses generated by seven large language models. These responses were generated using diverse prompts designed to elicit appropriate and potentially sensitive responses. Each response was then annotated across sixteen ethical categories, including stereotyping, toxicity, misinformation, hate speech, harmful advice, privacy violations, political bias, false confidence, emotional or religious insensitivity, sexual content, manipulation, and impersonation. To classify ethical and unethical AI-generated content, the dataset is analysed using state-of-the-art classification methods, assessing its ability to support reliable ethical evaluation. Performance is reported both for binary ethical classification and multilabel violation identification. Results include accuracies of nearly 99% for binary classification tasks with SVM and CNN models, and macro-F1 scores of about 96% on multilabel tasks for Sentence-BERT transformer model.