With the development of big data technology, the management of sensitive information online has become increasingly challenging, requiring innovative approaches for its effective control and classification. This paper addresses specific challenges in the classification of sensitive text in order to mitigate potential threats associated with the uncontrolled dissemination of sensitive information. Existing methods have shortcomings in extracting features from sensitive text, mainly due to insufficient feature extraction and interference caused by common features across domains. To address these challenges, we propose SecureSem, consisting of FeaEmbark and FeaPristine, which introduces innovative features to improve information exploitation within pre-trained language models such as BERT. The FeaEmbark module strengthens semantic features through feature standardisation operations, and the FeaPristine module purifies features through matrix mapping. Our experimental results, conducted on sensitive and general datasets, demonstrate the effectiveness of the SecureSem model. For example, our model achieves an improvement from 91.86% to 96.86% compared to BERT. This study improves sensitive text classification by introducing semantic features and optimisation operations, paving the way for further advances in the field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SecureSem: Sensitive Text Classification Based on Semantic Feature Optimization

  • Kangyuan Qin,
  • Ran Liu,
  • Min Yu,
  • Gang Li,
  • Mingqi Liu,
  • Jingyuan Li,
  • Weiqing Huang,
  • Jianguo Jiang

摘要

With the development of big data technology, the management of sensitive information online has become increasingly challenging, requiring innovative approaches for its effective control and classification. This paper addresses specific challenges in the classification of sensitive text in order to mitigate potential threats associated with the uncontrolled dissemination of sensitive information. Existing methods have shortcomings in extracting features from sensitive text, mainly due to insufficient feature extraction and interference caused by common features across domains. To address these challenges, we propose SecureSem, consisting of FeaEmbark and FeaPristine, which introduces innovative features to improve information exploitation within pre-trained language models such as BERT. The FeaEmbark module strengthens semantic features through feature standardisation operations, and the FeaPristine module purifies features through matrix mapping. Our experimental results, conducted on sensitive and general datasets, demonstrate the effectiveness of the SecureSem model. For example, our model achieves an improvement from 91.86% to 96.86% compared to BERT. This study improves sensitive text classification by introducing semantic features and optimisation operations, paving the way for further advances in the field.