Background <p>Real-world data (RWD) in unstructured electronic health records (EHRs) is crucial for understanding complex diseases like cancer, but extracting structured information is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four small, locally deployable language models for information extraction from Italian EHRs.</p> Methods <p>We examine three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess the impact of prompt design on accuracy. We evaluate the performance of four open-source small language models (SLMs) for clinical information extraction from Italian electronic health records (EHRs) in the APOLLO 11 trial on non-small cell lung cancer (NSCLC). The extraction protocol involves four steps: problem definition, data preprocessing, Large Language Model (LLM)-based information extraction, and output evaluation.</p> Results <p>We show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM (Tumor, Node, Metastasis) staging, PD-L1 (Programmed death-ligand 1), and ECOG-PS (Eastern Cooperative Oncology Group-Performance Status) are more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improve performance and reduced hallucinations. Clinical expertise enhances consistency in the extraction, particularly among students using annotated examples.</p> Conclusions <p>The study confirms that privacy-preserving SLMs can be deployed locally for efficient and secure cancer data extraction. Findings highlight the need for hybrid systems combining SLMs with expert input and underline the importance of aligning clinical documentation practices with SLM capabilities. This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Clinician expertise and prompt engineering enhance cancer information extraction in electronic health records by small language models

  • Federica Corso,
  • Vittoria Peppoloni,
  • Laura Mazzeo,
  • Giuseppe Leone,
  • Luana Passos,
  • Vanja Mišković,
  • Justin Armanini,
  • Alberto Ferrarin,
  • Isabella Catharina Wiest,
  • Fabian Wolf,
  • Giulia Montelatici,
  • Rebecca Romanò,
  • Paolo Ambrosini,
  • Tommaso Capoccia,
  • Stefano Natangelo,
  • Simone Rota,
  • Paola Andena,
  • Marta De Ponti,
  • Alessandra Russo,
  • Giulia Stasi,
  • Leonardo Provenzano,
  • Andrea Spagnoletti,
  • Marco Meazza Prina,
  • Chiara Cavalli,
  • Claudia Giani,
  • Roberta Serino,
  • Michele Borracino,
  • Chiara Bonalume,
  • Rosa Maria Di Mauro,
  • Claudia Agosta,
  • Andra Diana Dumitrascu,
  • Giorgia Di Liberti,
  • Giulia Corrao,
  • Teresa Beninato,
  • Monica Ganzinelli,
  • Mario Occhipinti,
  • Marta Brambilla,
  • Claudia Proto,
  • Jakob Nikolas Kather,
  • Alessandra Laura Giulia Pedrocchi,
  • Filippo De Braud,
  • Giuseppe Lo Russo,
  • Paolo Baili,
  • Arsela Prelaj

摘要

Background

Real-world data (RWD) in unstructured electronic health records (EHRs) is crucial for understanding complex diseases like cancer, but extracting structured information is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four small, locally deployable language models for information extraction from Italian EHRs.

Methods

We examine three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess the impact of prompt design on accuracy. We evaluate the performance of four open-source small language models (SLMs) for clinical information extraction from Italian electronic health records (EHRs) in the APOLLO 11 trial on non-small cell lung cancer (NSCLC). The extraction protocol involves four steps: problem definition, data preprocessing, Large Language Model (LLM)-based information extraction, and output evaluation.

Results

We show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM (Tumor, Node, Metastasis) staging, PD-L1 (Programmed death-ligand 1), and ECOG-PS (Eastern Cooperative Oncology Group-Performance Status) are more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improve performance and reduced hallucinations. Clinical expertise enhances consistency in the extraction, particularly among students using annotated examples.

Conclusions

The study confirms that privacy-preserving SLMs can be deployed locally for efficient and secure cancer data extraction. Findings highlight the need for hybrid systems combining SLMs with expert input and underline the importance of aligning clinical documentation practices with SLM capabilities. This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.