<p>Automated radiology report generation aims to alleviate documentation burdens but faces a critical trade-off between the computational overhead of Transformers and the inability of efficient State Space Models (SSMs) to capture two-dimensional spatial relationships. This study proposes MG-Hybrid to harmonize computational efficiency with anatomical fidelity while bridging the semantic gap between visual features and clinical text. We introduce a Mask-Guided Visual Disentanglement strategy that extracts structured organ-specific prototypes to preserve the spatial integrity of regions such as the heart and lungs, rather than treating images as flattened sequences. These features are processed by a Hybrid Anatomy Encoder which integrates a Mamba-based state-space module for organ-level representation refinement with a Transformer layer for global inter-organ reasoning. Additionally, a Visual-Textual Semantic Alignment mechanism is incorporated to anchor visual representations to clinical semantics explicitly. The framework was evaluated on the IU X-ray and MIMIC-CXR benchmarks. Extensive experiments demonstrate that the proposed approach achieves superior performance in clinical efficacy and precision-oriented metrics. The results indicate that the model generates reports that are more consistent with the reference reports and achieves improved clinical efficacy metrics while maintaining robust generation quality compared to existing baselines. By harmonizing the computational advantages of State Space Models (SSMs) with advanced spatial and semantic modeling, MG-Hybrid addresses the limitations of current architectures. This approach provides a viable pathway toward more accurate and computationally efficient automated diagnostic support.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MG-Hybrid: mask-guided hybrid state space models with semantic alignment for radiology report generation

  • Yilin Li,
  • Zijian Zhao,
  • Yizhuo Yuan,
  • Yan Yan

摘要

Automated radiology report generation aims to alleviate documentation burdens but faces a critical trade-off between the computational overhead of Transformers and the inability of efficient State Space Models (SSMs) to capture two-dimensional spatial relationships. This study proposes MG-Hybrid to harmonize computational efficiency with anatomical fidelity while bridging the semantic gap between visual features and clinical text. We introduce a Mask-Guided Visual Disentanglement strategy that extracts structured organ-specific prototypes to preserve the spatial integrity of regions such as the heart and lungs, rather than treating images as flattened sequences. These features are processed by a Hybrid Anatomy Encoder which integrates a Mamba-based state-space module for organ-level representation refinement with a Transformer layer for global inter-organ reasoning. Additionally, a Visual-Textual Semantic Alignment mechanism is incorporated to anchor visual representations to clinical semantics explicitly. The framework was evaluated on the IU X-ray and MIMIC-CXR benchmarks. Extensive experiments demonstrate that the proposed approach achieves superior performance in clinical efficacy and precision-oriented metrics. The results indicate that the model generates reports that are more consistent with the reference reports and achieves improved clinical efficacy metrics while maintaining robust generation quality compared to existing baselines. By harmonizing the computational advantages of State Space Models (SSMs) with advanced spatial and semantic modeling, MG-Hybrid addresses the limitations of current architectures. This approach provides a viable pathway toward more accurate and computationally efficient automated diagnostic support.