Automated approaches to idea evaluation increasingly leverage generative artificial intelligence to support decision-makers. However, contextualizing evaluations within specific domains remains challenging, particularly at varying levels of large language model (LLM) augmentation. Existing research employs embeddings to derive semantic insights, yet these representations often lack domain-specific contextualization. Recent advancements, such as chat-based LLMs, present new opportunities to incorporate context through prompting. To address these challenges, we propose a structured evaluation pipeline that integrates embeddings with feature engineering to enhance the contextualization of chat-based LLM evaluations. Using a real-world innovation challenge, we instantiate this pipeline and assess its predictive performance across different levels of augmentation. Our findings reveal that incorporating contextual information improves predictive accuracy but depends on fine-grained idea quality dimensions. By codifying our approach into a reference model, we provide a transferable framework that generalizes across various evaluation contexts employing LLMs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLM-Augmentation for Idea Evaluation: Developing a Reference Model for Evaluation Pipelines

  • Philipp Gordetzki

摘要

Automated approaches to idea evaluation increasingly leverage generative artificial intelligence to support decision-makers. However, contextualizing evaluations within specific domains remains challenging, particularly at varying levels of large language model (LLM) augmentation. Existing research employs embeddings to derive semantic insights, yet these representations often lack domain-specific contextualization. Recent advancements, such as chat-based LLMs, present new opportunities to incorporate context through prompting. To address these challenges, we propose a structured evaluation pipeline that integrates embeddings with feature engineering to enhance the contextualization of chat-based LLM evaluations. Using a real-world innovation challenge, we instantiate this pipeline and assess its predictive performance across different levels of augmentation. Our findings reveal that incorporating contextual information improves predictive accuracy but depends on fine-grained idea quality dimensions. By codifying our approach into a reference model, we provide a transferable framework that generalizes across various evaluation contexts employing LLMs.