Using prompt-engineering and retrieval augmented generation, we can leverage pre-trained Large Language Models to answer domain-specific questions relying on information from textual sources. In this work, we discuss how to assess the trustworthiness of a module that performs such task: how to build a large, representative, and unbiased dataset of questions/answers by automatically generating variations and which metrics to compute. We apply the methodology to a use-case where a smart wheelchair provides answers about its functioning, presenting experimental results on a dataset of more than 1000 questions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing the Trustworthiness of Large Language Models on Domain-Specific Questions

  • Sandra Mitrović,
  • Matteo Mazzola,
  • Roberto Larcher,
  • Jérôme Guzzi

摘要

Using prompt-engineering and retrieval augmented generation, we can leverage pre-trained Large Language Models to answer domain-specific questions relying on information from textual sources. In this work, we discuss how to assess the trustworthiness of a module that performs such task: how to build a large, representative, and unbiased dataset of questions/answers by automatically generating variations and which metrics to compute. We apply the methodology to a use-case where a smart wheelchair provides answers about its functioning, presenting experimental results on a dataset of more than 1000 questions.