<p>The evaluation of Conversational Recommender Systems necessitates robust protocols to measure utility and user satisfaction. While human-in-the-loop testing remains the gold standard, scalability and reproducibility constraints have driven the field toward User Simulators. However, current simulation paradigms predominantly utilize rigid templates or closed-source Large Language Models that exhibit idealized behaviors. These approaches fail to capture user ambiguity, resulting in benchmarks that overestimate system proficiency by assuming crystallized user intent. To address this limitation, we introduce a family of open-weight user simulation models capable of generalizing across diverse e-commerce domains. Leveraging Teacher-Student distillation, we operationalize three distinct behavioral stereotypes: <i>Direct</i>, <i>Vague-Proactive</i>, and <i>Vague-Reactive</i>. Our evaluation of state-of-the-art Agentic Generative Conversational Recommender Systems reveals a critical <b>Robustness Gap</b>: while agents perform proficiently with decisive users, performance collapses when facing passivity and ambiguity. These findings underscore the necessity of our scalable framework for rigorously stress-testing the next generation of conversational agents against realistic, non-cooperative user behaviors.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Beyond the Lookup: Simulating Realistic User Uncertainty for the Evaluation of Conversational Agentic Recommenders

  • Alessandro Petruzzelli,
  • Alessandro Francesco Maria Martina,
  • Cataldo Musto,
  • Marco de Gemmis,
  • Pasquale Lops,
  • Giovanni Semeraro

摘要

The evaluation of Conversational Recommender Systems necessitates robust protocols to measure utility and user satisfaction. While human-in-the-loop testing remains the gold standard, scalability and reproducibility constraints have driven the field toward User Simulators. However, current simulation paradigms predominantly utilize rigid templates or closed-source Large Language Models that exhibit idealized behaviors. These approaches fail to capture user ambiguity, resulting in benchmarks that overestimate system proficiency by assuming crystallized user intent. To address this limitation, we introduce a family of open-weight user simulation models capable of generalizing across diverse e-commerce domains. Leveraging Teacher-Student distillation, we operationalize three distinct behavioral stereotypes: Direct, Vague-Proactive, and Vague-Reactive. Our evaluation of state-of-the-art Agentic Generative Conversational Recommender Systems reveals a critical Robustness Gap: while agents perform proficiently with decisive users, performance collapses when facing passivity and ambiguity. These findings underscore the necessity of our scalable framework for rigorously stress-testing the next generation of conversational agents against realistic, non-cooperative user behaviors.