The Amazon-M2 dataset, introduced during the Amazon KDD Cup 2023, has become a landmark resource for multilingual, multi-locale product recommendation and title generation tasks. Its prominence continued into the KDD Cup 2024, further solidifying its importance. As interest in this dataset grows, this paper offers a comprehensive review of the machine learning and deep learning algorithms applied to the Amazon-M2 dataset, emphasizing their performance and practical applications. In addition, it provides an overview of the dataset’s structure and introduces its three primary tasks: session-based product recommendation, cross-locale product recommendation, and next-product title generation. The paper explores state-of-the-art techniques employed by top-ranking teams, such as Transformers, CNNs, and XGBoost, to solve the three tasks. Moreover, it also delves into key evaluation metrics, including MRR@K and BLEU scores, to provide deeper insights into the results. Finally, the paper recommends practices and highlights available source code to assist those seeking to build models using the Amazon-M2 dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Review of Machine Learning Algorithms for the Amazon-M2 Dataset

  • Chih Hsuan Hsieh

摘要

The Amazon-M2 dataset, introduced during the Amazon KDD Cup 2023, has become a landmark resource for multilingual, multi-locale product recommendation and title generation tasks. Its prominence continued into the KDD Cup 2024, further solidifying its importance. As interest in this dataset grows, this paper offers a comprehensive review of the machine learning and deep learning algorithms applied to the Amazon-M2 dataset, emphasizing their performance and practical applications. In addition, it provides an overview of the dataset’s structure and introduces its three primary tasks: session-based product recommendation, cross-locale product recommendation, and next-product title generation. The paper explores state-of-the-art techniques employed by top-ranking teams, such as Transformers, CNNs, and XGBoost, to solve the three tasks. Moreover, it also delves into key evaluation metrics, including MRR@K and BLEU scores, to provide deeper insights into the results. Finally, the paper recommends practices and highlights available source code to assist those seeking to build models using the Amazon-M2 dataset.