Illustration or Illusion? Reassessing the Use of Machine Learning in Phishing Email Detection
摘要
Despite previous research illustrated the very high accuracy of machine learning (ML) algorithms in detecting phishing emails, billions of people continue to fall victim to email phishing, often due to detection systems failing to catch them. This study reassesses the use of ML algorithms in detecting previously unseen phishing emails and their ability to identify new phishing tactics. We apply three distinct ML algorithms across seven groups of datasets, and the findings of the thorough quantitative analysis indicate that the performance of ML algorithms substantially varies when evaluated on previously unseen or recent emails. We highlight the difficulty in generalizing learned patterns and we recommend potential solutions to address the identified challenges.