Benchmark for Lymphoma Information Extraction and Automated Coding
摘要
We present a specialized dataset for automated lymphoma coding, comprising 162 case reports sourced from Chinese Medical Case Repository website. Each report is coded according to the ICD-O-3 and ICD-10 classification systems for both Hodgkin lymphoma and non-Hodgkin lymphoma cases. The dataset is strategically divided into three equal parts: a training set of 54 reports, and two test sets (A and B) containing 54 case reports each, providing a foundation for developing and evaluating large language models in oncology coding. Our resource is designed to facilitate advancements in automated tumor coding systems, with the potential to significantly enhance the efficiency and accuracy of lymphoma classification in clinical settings. The dataset, along with comprehensive evaluation metrics, is accessible at http://cips-chip.org.cn/2024/eval2 . It serves as a valuable tool for researchers and practitioners in medical informatics and oncology, particularly those focusing on Chinese language medical texts and international classification guidelines.