<p>JavaScript is the most commonly used scripting language in today’s web development industry that powers a wide array of modern-day applications. The broad utilization of JavaScript significantly widens the attack surface on the web. A number of security risks result from its misuse, abuse, and vulnerable implementation in web environments, thus making possible such threats as cross-site scripting XSS, drive-by downloads, malicious advertising, and cryptojacking. For over two decades, researchers have been working hard to develop better methods and approaches for malicious JavaScript detection, but the battle between malicious JavaScript detection and attackers continues as both sides keep developing increasingly advanced methods and analytical systems. Motivated by the lack of a comprehensive, up-to-date survey spanning two decades of dataset-backed research in this area, this survey conducts a structured literature review covering January 2005 to November 2025 across ACM Digital Library (ACM DL), IEEE Xplore, SpringerLink, ScienceDirect, and Google Scholar and systematically selects 58 peer-reviewed studies reporting experiments on datasets and measurable results. We synthesize prior work across major detection approaches, code representations, and learning paradigms, and summarize the field’s progression from traditional feature-engineering methods toward deep learning and transformer-based techniques. This review outlines several important conclusions about detection robustness, availability of datasets, reproducibility, consistent evaluation, and potential adversarial attacks. Unlike previous works on the topic, it provides a concise taxonomy of feature engineering strategies that are connected to broader trends in detection performance and deployment practices. We also identify limitations in the field, including limited dataset availability, inconsistent evaluation settings, and insufficient robustness analysis under obfuscation and adversarial manipulation. Finally, this paper describes the current open problems of the field and proposes a future path forward toward building more robust, reproducible, and deployment-aware malicious JavaScript detection systems. Overall, this survey outlines the history of the field and highlights important deficiencies in the evidence available, as well as how future research can strengthen the reliability, comparability, and practical value of malicious JavaScript detection systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Malicious JavaScript detection using machine learning: a survey of taxonomy, trends, and open challenges

  • Qasim Bataineh,
  • Zeyad Hailat,
  • Qusay M. Al-Zubi,
  • Mst Shapna Akter,
  • Mohamed A. Zohdy

摘要

JavaScript is the most commonly used scripting language in today’s web development industry that powers a wide array of modern-day applications. The broad utilization of JavaScript significantly widens the attack surface on the web. A number of security risks result from its misuse, abuse, and vulnerable implementation in web environments, thus making possible such threats as cross-site scripting XSS, drive-by downloads, malicious advertising, and cryptojacking. For over two decades, researchers have been working hard to develop better methods and approaches for malicious JavaScript detection, but the battle between malicious JavaScript detection and attackers continues as both sides keep developing increasingly advanced methods and analytical systems. Motivated by the lack of a comprehensive, up-to-date survey spanning two decades of dataset-backed research in this area, this survey conducts a structured literature review covering January 2005 to November 2025 across ACM Digital Library (ACM DL), IEEE Xplore, SpringerLink, ScienceDirect, and Google Scholar and systematically selects 58 peer-reviewed studies reporting experiments on datasets and measurable results. We synthesize prior work across major detection approaches, code representations, and learning paradigms, and summarize the field’s progression from traditional feature-engineering methods toward deep learning and transformer-based techniques. This review outlines several important conclusions about detection robustness, availability of datasets, reproducibility, consistent evaluation, and potential adversarial attacks. Unlike previous works on the topic, it provides a concise taxonomy of feature engineering strategies that are connected to broader trends in detection performance and deployment practices. We also identify limitations in the field, including limited dataset availability, inconsistent evaluation settings, and insufficient robustness analysis under obfuscation and adversarial manipulation. Finally, this paper describes the current open problems of the field and proposes a future path forward toward building more robust, reproducible, and deployment-aware malicious JavaScript detection systems. Overall, this survey outlines the history of the field and highlights important deficiencies in the evidence available, as well as how future research can strengthen the reliability, comparability, and practical value of malicious JavaScript detection systems.