Comprehensive Evaluation of Pythia Model Efficiency in De-identification and Normalization for Enhanced Medical Data Management
摘要
This study presents our work developed for the AICUP2023-privacy protection and standardization of electronic medical record text notes. Our work focuses on exploring the efficiency of the Pythia model developed by the EleutherAI community applying it to the de-identification and normalization problems. The core objective of the research outcome is to achieve a dual purpose: on one hand, to protect patient privacy, and on the other hand, to enhance the automation and efficiency of medical data management, while striving to reduce manual de-identification errors caused by human factors. To this end, we will not only examine the standard configuration of the Pythia model but also delve into its performance under different parameter settings, to comprehensively evaluate and compare its effectiveness and results in handling sensitive medical data. Our experiment results highlight that merely increasing the model size might not be sufficient to yield the expected benefit for recognizing protected health information from text. However, for the task of temporal information normalization, the large language model demonstrates a significant performance improvement after fine-tuning.