<p>Large Language Models (LLMs) have advanced various Natural Language Processing (NLP) tasks, such as text generation and translation, among others. However, these models often generate texts that can perpetuate biases. Existing approaches to mitigate these biases usually compromise knowledge retention. This study explores whether LLMs can produce safe, unbiased outputs without sacrificing knowledge or comprehension. We introduce the Safe and Responsible Large Language Model (<b>SR</b><InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10994_2025_6767_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="27" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{\text {LLM}}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mtext>LLM</mtext> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation>), which has been instruction fine-tuned atop of a safe fine-tuned auto-regressive decoder-only LLM to reduce biases in generated texts. We developed a specialized dataset with examples of unsafe and corresponding safe variations to train <b>SR</b><InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10994_2025_6767_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="27" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{\text {LLM}}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mtext>LLM</mtext> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> to identify and correct biased text. Experiments on our specialized dataset and out-of-distribution test sets reveal that <b>SR</b><InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10994_2025_6767_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="27" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{\text {LLM}}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mtext>LLM</mtext> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> effectively reduces biases while preserving knowledge integrity. This performance surpasses that of traditional fine-tuning of smaller language models and base LLMs that merely reply on prompting techniques. Our findings demonstrate that instruction fine-tuning on custom datasets tailored for tasks such as debiasing is a highly effective strategy for minimizing bias in LLM while preserving their inherent knowledge and capabilities. The code and dataset are accessible at <a href="https://github.com/shainarazavi/Safe-Responsible-LLM">https://github.com/shainarazavi/Safe-Responsible-LLM</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Developing safe and responsible large language model: can we balance bias reduction and language understanding?

  • Shaina Raza,
  • Oluwanifemi Bamgbose,
  • Shardul Ghuge,
  • Fatemeh Tavakoli,
  • Deepak John Reji,
  • Syed Raza Bashir

摘要

Large Language Models (LLMs) have advanced various Natural Language Processing (NLP) tasks, such as text generation and translation, among others. However, these models often generate texts that can perpetuate biases. Existing approaches to mitigate these biases usually compromise knowledge retention. This study explores whether LLMs can produce safe, unbiased outputs without sacrificing knowledge or comprehension. We introduce the Safe and Responsible Large Language Model (SR \(_{\text {LLM}}\) LLM ), which has been instruction fine-tuned atop of a safe fine-tuned auto-regressive decoder-only LLM to reduce biases in generated texts. We developed a specialized dataset with examples of unsafe and corresponding safe variations to train SR \(_{\text {LLM}}\) LLM to identify and correct biased text. Experiments on our specialized dataset and out-of-distribution test sets reveal that SR \(_{\text {LLM}}\) LLM effectively reduces biases while preserving knowledge integrity. This performance surpasses that of traditional fine-tuning of smaller language models and base LLMs that merely reply on prompting techniques. Our findings demonstrate that instruction fine-tuning on custom datasets tailored for tasks such as debiasing is a highly effective strategy for minimizing bias in LLM while preserving their inherent knowledge and capabilities. The code and dataset are accessible at https://github.com/shainarazavi/Safe-Responsible-LLM.