In the current digital landscape, distinguishing between human-written texts and those generated by large language models (LLMs) is essential for information security and the prevention of academic fraud. Texts generated by LLMs often closely resemble high-quality human-written content, posing significant challenges for accurate identification. To tackle this, we introduce the Identify the Writer by Writing Style (IWWS) model, which integrates perplexity scores with text embeddings through feature fusion. Our innovative approach employs a similarity matrix and contrastive learning to improve the model’s ability to detect unique writing styles. Additionally, we present the HumanGenTextify dataset, which reflects real-world text generation scenarios and serves as a robust foundation for distinguishing between human and model-generated texts. Experimental results show that our IWWS model has superior performance over existing methods, achieving high accuracy in text source detection and offering insights into distinctive writing styles. In addition, our research paves the way for future advancements in automated LLMs-generated text detection and authenticity verification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Who is the Writer? Identifying the Generative Model by Writing Style

  • Jiawen Yan,
  • Baohua Zhang,
  • Wenyao Cui,
  • Huaping Zhang

摘要

In the current digital landscape, distinguishing between human-written texts and those generated by large language models (LLMs) is essential for information security and the prevention of academic fraud. Texts generated by LLMs often closely resemble high-quality human-written content, posing significant challenges for accurate identification. To tackle this, we introduce the Identify the Writer by Writing Style (IWWS) model, which integrates perplexity scores with text embeddings through feature fusion. Our innovative approach employs a similarity matrix and contrastive learning to improve the model’s ability to detect unique writing styles. Additionally, we present the HumanGenTextify dataset, which reflects real-world text generation scenarios and serves as a robust foundation for distinguishing between human and model-generated texts. Experimental results show that our IWWS model has superior performance over existing methods, achieving high accuracy in text source detection and offering insights into distinctive writing styles. In addition, our research paves the way for future advancements in automated LLMs-generated text detection and authenticity verification.