Live Prototyping for Evaluating Generative AI: Exploring the Potential and the Pitfalls
摘要
This paper explores the limitations of traditional prototyping methods for evaluating generative AI products and proposes live prototyping as a solution. Recognizing the critical role of Large Language Model (LLM) output quality in user experience, the authors argue that static or clickable prototypes fail to capture the dynamic and personalized nature of GenAI interactions. To address this, the authors developed live prototypes connected to real LLMs and employed them in user studies. The findings demonstrate that live prototypes enable users to explore their own use cases, leading to more accurate and relevant feedback. Additionally, live prototypes provide valuable insights into users’ perceptions of LLM output quality, facilitating improved model evaluation and stakeholder communication. This paper discusses the benefits of live prototypes for empathy building and aligning users’ expectations with product capabilities. The paper acknowledges the limitations of this approach, including the technical expertise required for development and maintenance. This work contributes to a growing understanding of effective evaluation methodologies for GenAI products in the field of human-computer interaction.