LLM-Augmentation for Idea Evaluation: Developing a Reference Model for Evaluation Pipelines
摘要
Automated approaches to idea evaluation increasingly leverage generative artificial intelligence to support decision-makers. However, contextualizing evaluations within specific domains remains challenging, particularly at varying levels of large language model (LLM) augmentation. Existing research employs embeddings to derive semantic insights, yet these representations often lack domain-specific contextualization. Recent advancements, such as chat-based LLMs, present new opportunities to incorporate context through prompting. To address these challenges, we propose a structured evaluation pipeline that integrates embeddings with feature engineering to enhance the contextualization of chat-based LLM evaluations. Using a real-world innovation challenge, we instantiate this pipeline and assess its predictive performance across different levels of augmentation. Our findings reveal that incorporating contextual information improves predictive accuracy but depends on fine-grained idea quality dimensions. By codifying our approach into a reference model, we provide a transferable framework that generalizes across various evaluation contexts employing LLMs.