Assessing Model Quality Using Large Language Models
摘要
Recently, researchers have explored whether large language models (LLMs) can be used as a substitute for domain experts to elicit information that should be represented in an enterprise model. This paper examines a slightly different application purpose, assessing an existing model’s quality using an LLM. We will analyze which aspects of model quality can be evaluated using an LLM in principle, referring to the established model quality framework SEQUAL. We will present a first test of assessing perceived semantic quality using ChatGPT. To examine the effect of different prompting strategies, we compared our results to the assessments of human domain experts. Our results suggest that LLMs are suitable for assessing the perceived semantic quality of models and provide a basis for considering further quality dimensions in future work.