Benchmarking LLMs for Business Architecture Modelling with Hierarchical Capability Maps
摘要
Business Capability Map is one of the core instruments of Business Architecture (BA) modeling and analysis and an essential tool for driving business/IT alignment. However, the process of crafting a structured and hierarchical overview of organisational capabilities as a business capability map is a manual, knowledge-intensive process that consumes a significant amount of effort and time. Large Language Models (LLMs) have demonstrated their ability to automate knowledge-intensive tasks such as business process modeling through their internalised knowledge. However, they tend to perform poorly when the task requires specific domain knowledge as opposed to handling general knowledge. This is a hurdle in adapting LLMs for BA modeling, as domain expertise is crucial for generating business architecture models. To further our understanding in this challenge, this paper presents a benchmark experiment that systematically and comprehensively evaluates the utility of LLMs in BA modeling. We propose BCM-Eval, a novel business capability map benchmark, and use it to evaluate key state-of-the-art LLMs in different prompt settings. We report on the potential and limitations of LLMs for business capability modeling, concluding that LLMs still have a limited grasp on industry expertise and do not precisely capture the semantics related to capability models. Our results also indicate the need for advanced prompting and domain knowledge augmentation techniques that can probe the knowledge of LLMs towards capability maps and other hierarchical BA models.