Enhancing Sentence-Level Privacy Risk Classification in Electronic Health Records Using Clinical BERT: A Comparative Machine Learning Approach
摘要
This study extends an earlier novel approach to enhancing privacy risk classification in Electronic Health Records (EHRs) by employing sentence-level analysis within clinical notes. Transitioning from traditional document-level classification, which often hinders necessary data sharing due to its conservative nature, this research aims to strike a balance between maintaining data utility and ensuring privacy protection. Leveraging Clinical Bidirectional Encoder Representations from Transformers (BERT), a state-of-the-art Natural Language Processing (NLP) model tailored for clinical text, the study classifies sentences in EHRs as either containing or not containing protected health information (PHI). Utilizing the Harvard i2b2 dataset for model training and evaluation, the findings reveal that Clinical BERT surpasses earlier models in precision, recall, and f-score, highlighting its effectiveness in medical NLP tasks. This research emphasizes the significance of advanced NLP techniques, such as transformer-based models and transfer learning, in refining privacy risk classification within healthcare data. As the world transitions finding ways to ethically share data and resources between organizations, it is critical to uphold patient privacy.