<p>In recent years, e-commerce has experienced growth in sales, brands and customers. Unfortunately, cybercriminals have taken advantage of this by creating fraudulent websites to scam customers. The large amount of new e-commerce websites outnumbers the manual reporting capabilities, exposing users to these attacks. In this work, we used machine learning techniques to identify possible fraudulent online stores. To achieve this, we created ELFW-2031 (E-commerce Legitimate Fraudulent Websites), an updated dataset of manually verified legitimate and fraudulent e-commerce websites and a comprehensive set of resources for researchers to compare their methods. We released this dataset for public use to overcome the lack of a comprehensive corpus of this type of websites. We also designed a novel set of 50 features using six different resources obtained from the website content and external services. We used these new features to train and test two models: (i) a model with all available resources focused on improving accuracy and (ii) a model focused on scalability independent of external services. The proposed models achieve F1 scores of <b>96.88%</b> and <b>96.53%</b> respectively using XGBoost. Finally, we evaluated the performance of the proposed features, showing that novel features from social media and the technology analysis were the most valuable ones.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fraud detection in e-commerce: a comparative analysis of features to enhance machine learning models

  • Manuel Sánchez-Paniagua,
  • Eduardo Fidalgo,
  • Enrique Alegre,
  • Francisco Jáñez-Martino

摘要

In recent years, e-commerce has experienced growth in sales, brands and customers. Unfortunately, cybercriminals have taken advantage of this by creating fraudulent websites to scam customers. The large amount of new e-commerce websites outnumbers the manual reporting capabilities, exposing users to these attacks. In this work, we used machine learning techniques to identify possible fraudulent online stores. To achieve this, we created ELFW-2031 (E-commerce Legitimate Fraudulent Websites), an updated dataset of manually verified legitimate and fraudulent e-commerce websites and a comprehensive set of resources for researchers to compare their methods. We released this dataset for public use to overcome the lack of a comprehensive corpus of this type of websites. We also designed a novel set of 50 features using six different resources obtained from the website content and external services. We used these new features to train and test two models: (i) a model with all available resources focused on improving accuracy and (ii) a model focused on scalability independent of external services. The proposed models achieve F1 scores of 96.88% and 96.53% respectively using XGBoost. Finally, we evaluated the performance of the proposed features, showing that novel features from social media and the technology analysis were the most valuable ones.