Assessing the Trustworthiness of Large Language Models on Domain-Specific Questions
摘要
Using prompt-engineering and retrieval augmented generation, we can leverage pre-trained Large Language Models to answer domain-specific questions relying on information from textual sources. In this work, we discuss how to assess the trustworthiness of a module that performs such task: how to build a large, representative, and unbiased dataset of questions/answers by automatically generating variations and which metrics to compute. We apply the methodology to a use-case where a smart wheelchair provides answers about its functioning, presenting experimental results on a dataset of more than 1000 questions.