Structured evidence-guided discrete diffusion for image-conditioned radiology report generation
摘要
Radiology report generation (RRG) converts chest X-rays into structured findings and impressions, but autoregressive systems can propagate early entity, negation, or anatomical errors, and language-model-based systems may rely on language priors when visual evidence is weak. We propose SEDRRG, structured evidence-guided discrete diffusion for radiology report generation, an image-conditioned framework that refines the whole report sequence through discrete token denoising. SEDRRG combines Swin-V2 hierarchical global and patch-level evidence with denoising-time global-local gating, and uses multi-factor structured supervision to align token recovery with report-derived medical token/phrase cues, image-text alignment, and section organization. Experiments on IU X-Ray and MIMIC-CXR show competitive benchmark-level performance against representative autoregressive, knowledge-enhanced, expert-token-based, and language-model-assisted baselines under reproduced and literature-reported comparison settings. A complementary CheXbert-based 14-observation clinical-efficacy evaluation on MIMIC-CXR yields precision of 0.322, recall of 0.375, and F1 of 0.346. Ablation and controlled analyses support the contributions of hierarchical evidence encoding, gated conditioning, and structured supervision. These results position SEDRRG as a benchmark-oriented algorithmic framework, not as evidence of clinical deployment readiness.