Objective <p>To develop a privacy-preserving method for structuring free-text mammography reports using a locally fine-tuned, open-source large language model (LLM).</p> Materials and methods <p>In this multicenter study, 7161 unstructured mammography reports were collected from three institutions. The open-source Llama-3 model was fine-tuned via supervised learning using pseudo-labels from the commercial Qwen-Max model with low-rank adaptation. All labels were pseudo-labels generated by the commercial Qwen-Max model rather than human annotations. Structured outputs followed a BI-RADS-oriented nested JSON schema. Performance was evaluated across 23 features using Precision, Recall, and F1-score. Structural integrity was assessed using the JSON format accuracy (JFA) and field integrity accuracy (FIA) metrics. Statistical comparisons were performed using the paired Wilcoxon signed-rank test and Cohen’s d effect size.</p> Results <p>A total of 7161 reports were retrospectively obtained from three institutions and analyzed. The fine-tuned model achieved strong performance at epoch 10 (Precision 0.942, Recall 0.929, F1-score 0.932), with JFA and FIA reaching 0.964 and 1.000, respectively, showing significant gains over the base model (<i>p</i> &lt; 0.05, Cohen’s <i>d</i> &gt; 0.8). While slightly below Qwen-Max overall, the model exhibited moderate yet statistically significant differences (<i>p</i> &lt; 0.05; 0.5 &lt; Cohen’s <i>d</i> &lt; 0.8), particularly in the “special signs” category (F1 = 0.737 vs 0.947).</p> Conclusion <p>This method effectively converts mammography reports into structured data using a locally fine-tuned, open-source LLM. Although there is a slight performance trade-off, it improves privacy and can be deployed locally. Its accuracy, clinical relevance, and compliance make it a practical solution for medical institutions.</p> Key Points <p><Emphasis Type="BoldItalic">Question</Emphasis> <i>Free-text mammography reports lack standardization, making structured extraction difficult, while existing solutions often compromise privacy, adaptability, or require costly commercial tools</i>.</p> <p><Emphasis Type="BoldItalic">Findings</Emphasis> <i>Our fine-tuned LLaMA-3 model achieved high extraction accuracy (F1-score: 0.932) and complete structural integrity (FIA: 1.000) within a fully local deployment pipeline</i>.</p> <p><Emphasis Type="BoldItalic">Clinical relevance</Emphasis> <i>This method enables standardized and privacy-preserving mammography reporting without disruption to the clinical workflow, supporting safer AI integration, enhanced data quality, and compliance with regulations such as the General Data Protection Regulation (GDPR)</i>.</p> Graphical Abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BI-RADS-compliant structured mammography reporting using locally deployed large language models under privacy constraints

  • Wenjun Sheng,
  • Yudong Wang,
  • Lianxiang Xiao,
  • Bin Guo,
  • Yan Zhang,
  • Liang Qiao,
  • Qingqing Li,
  • Liu Yang,
  • Yi Zhang

摘要

Objective

To develop a privacy-preserving method for structuring free-text mammography reports using a locally fine-tuned, open-source large language model (LLM).

Materials and methods

In this multicenter study, 7161 unstructured mammography reports were collected from three institutions. The open-source Llama-3 model was fine-tuned via supervised learning using pseudo-labels from the commercial Qwen-Max model with low-rank adaptation. All labels were pseudo-labels generated by the commercial Qwen-Max model rather than human annotations. Structured outputs followed a BI-RADS-oriented nested JSON schema. Performance was evaluated across 23 features using Precision, Recall, and F1-score. Structural integrity was assessed using the JSON format accuracy (JFA) and field integrity accuracy (FIA) metrics. Statistical comparisons were performed using the paired Wilcoxon signed-rank test and Cohen’s d effect size.

Results

A total of 7161 reports were retrospectively obtained from three institutions and analyzed. The fine-tuned model achieved strong performance at epoch 10 (Precision 0.942, Recall 0.929, F1-score 0.932), with JFA and FIA reaching 0.964 and 1.000, respectively, showing significant gains over the base model (p < 0.05, Cohen’s d > 0.8). While slightly below Qwen-Max overall, the model exhibited moderate yet statistically significant differences (p < 0.05; 0.5 < Cohen’s d < 0.8), particularly in the “special signs” category (F1 = 0.737 vs 0.947).

Conclusion

This method effectively converts mammography reports into structured data using a locally fine-tuned, open-source LLM. Although there is a slight performance trade-off, it improves privacy and can be deployed locally. Its accuracy, clinical relevance, and compliance make it a practical solution for medical institutions.

Key Points

Question Free-text mammography reports lack standardization, making structured extraction difficult, while existing solutions often compromise privacy, adaptability, or require costly commercial tools.

Findings Our fine-tuned LLaMA-3 model achieved high extraction accuracy (F1-score: 0.932) and complete structural integrity (FIA: 1.000) within a fully local deployment pipeline.

Clinical relevance This method enables standardized and privacy-preserving mammography reporting without disruption to the clinical workflow, supporting safer AI integration, enhanced data quality, and compliance with regulations such as the General Data Protection Regulation (GDPR).

Graphical Abstract