Navigating Data Imbalances in Cybersecurity: Identifying Malicious URLs with Multiple Labels and Extreme Data Imbalances with LGNet
摘要
Malicious URLs are a form of cyberattack, that manipulates individuals into disclosing sensitive information. Typically, Malicious URLs constitute a minor proportion of all searchable URLs and are marked by multi-labeling and a significant imbalance in sample data. These characteristics considerably diminish the effectiveness and precision of existing detection methodologies. This paper introduces LGNet, an innovative network framework designed to identify Malicious URLs in practical scenarios efficaciously. To tackle the multi-labeling issue, we have developed a label propagation algorithm that employs a confidence threshold limitation, ensuring high-confidence labeled URLs are acquired. We enhance the scalable tree system by using BILSTM and attention mechanism to address the substantial disparity between labeled and sample data in Malicious URL prediction. Our experimental results demonstrate that LGNet markedly surpasses existing state-of-the-art algorithms in detecting Malicious URLs.