In a world increasingly shaped by data-driven processes, the demand for robust data privacy has led to advancements in automated redaction technology designed to safeguard sensitive information within unstructured data. Existing tools often fall short, either by merely obscuring personally identifiable information (PII) or by entirely erasing it, leading to data loss and limiting document utility. Other tools rely on block encryption, which, while effective in security, causes context loss, impacting readability and comprehension. This paper proposes a novel framework for secure document redaction that integrates Natural Language Processing (NLP) with hybrid cryptography to overcome these limitations. The system leverages a pre-trained NLP model to identify and classify sensitive data using contextual analysis, enabling accurate redaction. By producing an irreversible key and utilizing stream cipher techniques with initialization vectors and XOR ciphers, the framework allows encryption at multiple granularities—bit, word, or sentence—adapting to enhance data security without compromising context. This approach signifies a scalable, secure advancement in data privacy, addressing critical shortcomings in existing tools and paving the way for more effective, automated redaction solutions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated Redaction Framework: Utilizing SpaCy and ChaCha20-Poly1305 for Secure Information Handling

  • A. Keerthana,
  • Sarang N. Asnani,
  • Rishi Dedhia,
  • Khyat A. Talsaniya,
  • Mosam Patel

摘要

In a world increasingly shaped by data-driven processes, the demand for robust data privacy has led to advancements in automated redaction technology designed to safeguard sensitive information within unstructured data. Existing tools often fall short, either by merely obscuring personally identifiable information (PII) or by entirely erasing it, leading to data loss and limiting document utility. Other tools rely on block encryption, which, while effective in security, causes context loss, impacting readability and comprehension. This paper proposes a novel framework for secure document redaction that integrates Natural Language Processing (NLP) with hybrid cryptography to overcome these limitations. The system leverages a pre-trained NLP model to identify and classify sensitive data using contextual analysis, enabling accurate redaction. By producing an irreversible key and utilizing stream cipher techniques with initialization vectors and XOR ciphers, the framework allows encryption at multiple granularities—bit, word, or sentence—adapting to enhance data security without compromising context. This approach signifies a scalable, secure advancement in data privacy, addressing critical shortcomings in existing tools and paving the way for more effective, automated redaction solutions.