Automatic protocol reverse engineering is essential for various security applications. However, the growing complexity and diversity of today’s private protocols have led to the emergence of mixed-format traffic, making it increasingly difficult to obtain valid protocol reverse information. Existing research rarely focuses on this, and faces the challenges of difficulty in feature representation. This paper proposes EHFC, which focuses on clustering traffic of the same protocol format from mixed-format traffic, providing a high-quality traffic partitioning foundation for subsequent protocol reverse tasks. This method employs the pre-trained traffic model to enhance the feature representation and uses the HDBSCAN to achieve automated format clustering. We evaluate the proposed approach on ten widely used protocols. It obtains high homogeneity (0.92) and completeness (0.95) during traffic clustering in different formats, which is significantly better than the other methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EHFC: Enhanced Format Clustering via Pre-Trained Traffic Model

  • Zhen Wang,
  • Sen Zhao,
  • Laile Xi,
  • Haiqiang Fei,
  • Feng Cheng,
  • Hong Li,
  • Hongsong Zhu

摘要

Automatic protocol reverse engineering is essential for various security applications. However, the growing complexity and diversity of today’s private protocols have led to the emergence of mixed-format traffic, making it increasingly difficult to obtain valid protocol reverse information. Existing research rarely focuses on this, and faces the challenges of difficulty in feature representation. This paper proposes EHFC, which focuses on clustering traffic of the same protocol format from mixed-format traffic, providing a high-quality traffic partitioning foundation for subsequent protocol reverse tasks. This method employs the pre-trained traffic model to enhance the feature representation and uses the HDBSCAN to achieve automated format clustering. We evaluate the proposed approach on ten widely used protocols. It obtains high homogeneity (0.92) and completeness (0.95) during traffic clustering in different formats, which is significantly better than the other methods.