SecureSem: Sensitive Text Classification Based on Semantic Feature Optimization
摘要
With the development of big data technology, the management of sensitive information online has become increasingly challenging, requiring innovative approaches for its effective control and classification. This paper addresses specific challenges in the classification of sensitive text in order to mitigate potential threats associated with the uncontrolled dissemination of sensitive information. Existing methods have shortcomings in extracting features from sensitive text, mainly due to insufficient feature extraction and interference caused by common features across domains. To address these challenges, we propose SecureSem, consisting of FeaEmbark and FeaPristine, which introduces innovative features to improve information exploitation within pre-trained language models such as BERT. The FeaEmbark module strengthens semantic features through feature standardisation operations, and the FeaPristine module purifies features through matrix mapping. Our experimental results, conducted on sensitive and general datasets, demonstrate the effectiveness of the SecureSem model. For example, our model achieves an improvement from 91.86% to 96.86% compared to BERT. This study improves sensitive text classification by introducing semantic features and optimisation operations, paving the way for further advances in the field.