Sequential recommendation problems have received increased research interest in recent years. Our knowledge about the effectiveness of sequential algorithms in practice is however limited. In this paper, we report on the outcomes of an A/B test on a video and movie streaming platform, where we benchmarked a sequential model against a non-sequential, personalized recommendation model, as well as a popularity-based baseline. Contrary to what we had expected from a preceding offline experiment, we observed that the popularity-based and the non-sequential models led to the highest click-through rates. However, in terms of the adoption of the recommendations, the sequential model was the most successful one in terms of viewing times. While our work points out the effectiveness of sequential models in practice, it also reminds us about important open challenges regarding (a) the sometimes limited predictive power of classic offline evaluations and (b) the dangers of optimizing recommendation models for click-through-rates.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Sequential Recommendations in the Wild: A Case Study on Offline Accuracy, Click Rates, and Consumption

  • Anastasiia Klimashevskaia,
  • Snorre Alvsvåg,
  • Christoph Trattner,
  • Alain D. Starke,
  • Astrid Tessem,
  • Dietmar Jannach

摘要

Sequential recommendation problems have received increased research interest in recent years. Our knowledge about the effectiveness of sequential algorithms in practice is however limited. In this paper, we report on the outcomes of an A/B test on a video and movie streaming platform, where we benchmarked a sequential model against a non-sequential, personalized recommendation model, as well as a popularity-based baseline. Contrary to what we had expected from a preceding offline experiment, we observed that the popularity-based and the non-sequential models led to the highest click-through rates. However, in terms of the adoption of the recommendations, the sequential model was the most successful one in terms of viewing times. While our work points out the effectiveness of sequential models in practice, it also reminds us about important open challenges regarding (a) the sometimes limited predictive power of classic offline evaluations and (b) the dangers of optimizing recommendation models for click-through-rates.