The transformative influence of Electronic Health Records (EHRs) is manifested through their facilitation of streamlined management and dissemination of health data among healthcare practitioners. These fosters enhanced collaboration in patient care and contribute to superior healthcare outcomes. This phenomenon within healthcare informatics underscores the integration of EHRs with sophisticated technologies such as artificial intelligence and machine learning. The intricate functionalities of EHR are essential for precision, efficiency, and continuity in patient care. Deidentification is a crucial process for preserving patient privacy, facilitating safe data analysis, promoting public trust, and enabling global collaboration, especially in healthcare to accelerate research and response to public health crises. By anonymizing personal information, deidentification not only mitigates the risk of privacy breaches but also promotes data sharing and collaboration in medical research, advancing scientific discovery while respecting individual privacy rights. On the other hand, it poses challenges in balancing data utility and protection. In response to these challenges, the AI Cup 2023 - Privacy Protection and Medical Data Standardization Competition and the 2024 International Workshop on Deidentification of Electronic Medical Record Notes (IW-DMRN) were organized. We participate in AI Cup 2023 competition by employing a hybrid approach to deidentification across diverse datasets, aiming to enhance and refine techniques for safeguarding privacy in healthcare data utilization. The research contributes valuable insights for advancing comprehension and strategies in protecting sensitive information within varied datasets. The study utilized two separate corpora: the 2014 i2b2/UTHealth deidentification corpus and the 2016 CEGS N-GRID deidentification corpus, in addition to the OpenDeID v2 corpus provided in the competition. These datasets underwent processing and pre-processing stages, encompassing various steps. The deidentification method employed the OpenDeID pipeline which involved segmentation, tokenization, tagging, and labeling, following a structured BIESO tagging scheme. The experimental setup incorporated a BERT-base model, fine-tuned on discharge summaries. The output underwent evaluation by the competition organizer using their evaluation metrics encompassed micro and macro averaged precision, recall, and F1-scores. The model demonstrated superior performance in Run1, achieving macro-averaged F1-scores of 0.7596718, 0.6969127, and 0.7269402 for precision, recall, and F-measure, respectively. The model highlighted high accuracy in deidentifying electronic health record text notes. The evaluation results affirm the model’s efficiency and reliability in precise deidentification tasks. The results provide significant insights into the ethical and privacy-conscious utilization of healthcare data within dynamic multicenter healthcare frameworks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of OpenDeID Pipeline in the 2023 SREDH/AI-Cup Competition for Deidentification of Sensitive Health Information

  • Shalini Gupta,
  • Naga Lalitha Valli Alla,
  • Omkar Panchal,
  • Jan Witowski,
  • Jitendra Jonnagaddala

摘要

The transformative influence of Electronic Health Records (EHRs) is manifested through their facilitation of streamlined management and dissemination of health data among healthcare practitioners. These fosters enhanced collaboration in patient care and contribute to superior healthcare outcomes. This phenomenon within healthcare informatics underscores the integration of EHRs with sophisticated technologies such as artificial intelligence and machine learning. The intricate functionalities of EHR are essential for precision, efficiency, and continuity in patient care. Deidentification is a crucial process for preserving patient privacy, facilitating safe data analysis, promoting public trust, and enabling global collaboration, especially in healthcare to accelerate research and response to public health crises. By anonymizing personal information, deidentification not only mitigates the risk of privacy breaches but also promotes data sharing and collaboration in medical research, advancing scientific discovery while respecting individual privacy rights. On the other hand, it poses challenges in balancing data utility and protection. In response to these challenges, the AI Cup 2023 - Privacy Protection and Medical Data Standardization Competition and the 2024 International Workshop on Deidentification of Electronic Medical Record Notes (IW-DMRN) were organized. We participate in AI Cup 2023 competition by employing a hybrid approach to deidentification across diverse datasets, aiming to enhance and refine techniques for safeguarding privacy in healthcare data utilization. The research contributes valuable insights for advancing comprehension and strategies in protecting sensitive information within varied datasets. The study utilized two separate corpora: the 2014 i2b2/UTHealth deidentification corpus and the 2016 CEGS N-GRID deidentification corpus, in addition to the OpenDeID v2 corpus provided in the competition. These datasets underwent processing and pre-processing stages, encompassing various steps. The deidentification method employed the OpenDeID pipeline which involved segmentation, tokenization, tagging, and labeling, following a structured BIESO tagging scheme. The experimental setup incorporated a BERT-base model, fine-tuned on discharge summaries. The output underwent evaluation by the competition organizer using their evaluation metrics encompassed micro and macro averaged precision, recall, and F1-scores. The model demonstrated superior performance in Run1, achieving macro-averaged F1-scores of 0.7596718, 0.6969127, and 0.7269402 for precision, recall, and F-measure, respectively. The model highlighted high accuracy in deidentifying electronic health record text notes. The evaluation results affirm the model’s efficiency and reliability in precise deidentification tasks. The results provide significant insights into the ethical and privacy-conscious utilization of healthcare data within dynamic multicenter healthcare frameworks.