Spam tactics are always on the rise. Therefore, traditional machine learning solutions for spam detection are increasingly challenged by the need for adaptive and intelligent models that will be able to mimic elements of human reasoning. This study tries to evaluate the efficacy of five important Large Language Models (LLMs) Llama2, Mistral, Synthia, Zephyr, and CausalLM in classifying spam emails. These models were selected for their diverse architectures and were tested on a uniform dataset to assess their spam detection capabilities without prior specific training. The ground truth is known and used for assessment. The models operated under a zero-shot learning framework, and tested mainly with limited computing infrastructure. The results show that while Mistral, Synthia, and Zephyr achieved high accuracy rates of 85–90%, Llama2 and CausalLM under performed, highlighting issues related to task comprehension and response generation considering the smaller models—7b models—tested. Additionally, response time analysis revealed that although effective, the processing speed of these models could be a limitation for real-time applications and that more computational resources are essential for effective deployment. In conclusion, this study suggests that LLMs can be integrated as a core layer for more complicated spam detection processes and could potentially enhance the adaptability and accuracy of spam detection. However, more future research to optimize LLMs responsiveness and real-time learning capabilities in cyber security contexts will be necessary.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Study on LLMs for Spam Email Classification

  • Rami Mohd,
  • Zhen Ze Ong

摘要

Spam tactics are always on the rise. Therefore, traditional machine learning solutions for spam detection are increasingly challenged by the need for adaptive and intelligent models that will be able to mimic elements of human reasoning. This study tries to evaluate the efficacy of five important Large Language Models (LLMs) Llama2, Mistral, Synthia, Zephyr, and CausalLM in classifying spam emails. These models were selected for their diverse architectures and were tested on a uniform dataset to assess their spam detection capabilities without prior specific training. The ground truth is known and used for assessment. The models operated under a zero-shot learning framework, and tested mainly with limited computing infrastructure. The results show that while Mistral, Synthia, and Zephyr achieved high accuracy rates of 85–90%, Llama2 and CausalLM under performed, highlighting issues related to task comprehension and response generation considering the smaller models—7b models—tested. Additionally, response time analysis revealed that although effective, the processing speed of these models could be a limitation for real-time applications and that more computational resources are essential for effective deployment. In conclusion, this study suggests that LLMs can be integrated as a core layer for more complicated spam detection processes and could potentially enhance the adaptability and accuracy of spam detection. However, more future research to optimize LLMs responsiveness and real-time learning capabilities in cyber security contexts will be necessary.