<p>Group spamming, where coordinated entities post deceptive reviews to manipulate product reputations, poses a significant threat to e-commerce integrity. Unlike traditional approaches that rely solely on behavioural patterns or linguistic features, this study introduces readability metrics as a novel and interpretable dimension for group spam detection. Spammers frequently craft highly readable and persuasive content to enhance credibility and mislead consumers. To capture this behaviour, we propose a two-phase framework that integrates sentiment-aware reviewer grouping with six quantitative readability measures (e.g., ARI, FRES, SMOG) and frequent pattern mining (FP-Growth) to isolate purely promotional or defamatory spammer groups. The model overcomes prior limitations in mixed-group identification. Evaluation on a balanced Amazon dataset using ten-fold cross-validation demonstrates that our Positive and Negative Group Spammer Readability Features (PGSRF and NGSRF) achieve 98.96% accuracy, 97% precision, and 95% recall with TabNet, and 99.48% accuracy, 96% precision, and 96% recall with Gradient Boosting Deep Networks (GBDN), outperforming behavioral and Linguistic baselines by 18–22%. Human annotation (k&#xa0;=&#xa0;0.71) further supports the method’s reliability. This work reveals that readability-based features are not only computationally efficient but also powerful indicators for spammer group detection, offering a novel and underutilised approach for strengthening trust in online review systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empowering e-commerce integrity: readability-driven detection of group spammers

  • Rupesh Kumar Dewang,
  • Anil Kumar Singh,
  • Arvind Mewada

摘要

Group spamming, where coordinated entities post deceptive reviews to manipulate product reputations, poses a significant threat to e-commerce integrity. Unlike traditional approaches that rely solely on behavioural patterns or linguistic features, this study introduces readability metrics as a novel and interpretable dimension for group spam detection. Spammers frequently craft highly readable and persuasive content to enhance credibility and mislead consumers. To capture this behaviour, we propose a two-phase framework that integrates sentiment-aware reviewer grouping with six quantitative readability measures (e.g., ARI, FRES, SMOG) and frequent pattern mining (FP-Growth) to isolate purely promotional or defamatory spammer groups. The model overcomes prior limitations in mixed-group identification. Evaluation on a balanced Amazon dataset using ten-fold cross-validation demonstrates that our Positive and Negative Group Spammer Readability Features (PGSRF and NGSRF) achieve 98.96% accuracy, 97% precision, and 95% recall with TabNet, and 99.48% accuracy, 96% precision, and 96% recall with Gradient Boosting Deep Networks (GBDN), outperforming behavioral and Linguistic baselines by 18–22%. Human annotation (k = 0.71) further supports the method’s reliability. This work reveals that readability-based features are not only computationally efficient but also powerful indicators for spammer group detection, offering a novel and underutilised approach for strengthening trust in online review systems.