Automated Report Generation (ARG) enhances clinical decision- making through efficient and standardized report generation, playing an essential role in intelligent healthcare. Yet existing methods predominantly focus on single-frame analysis and fail to process multi-frame nasal endoscopy sequences. Conventional frameworks ignore spatiotemporal correlations between frames, resulting in fragmented reports that cannot model dynamic lesion progression. Two critical bottlenecks arise: (1) inter-frame interference, where anatomical similarities (e.g., inferior turbinate vs. nasal septum) induce semantic confusion (visual noise); (2) continuous-frame misclassification, where dynamic region transitions cause erroneous category assignments (classification noise). These issues degrade lesion localization accuracy and report consistency. To overcome these limitations, we propose the Visual Attribute-Guided Multi-dimensional Perception (VAMP) framework with two innovations: (1) a visual attribute-guided module that suppresses noise via spatiotemporal lesion tracking, and (2) a cross- dimensional perception aggregation module enabling anatomical pathological fusion through dual-stream spatial-channel coordination. Experiments on IRA-HUT-NA, IU-Xray, and MIMIC-CXR show VAMP achieves 25.4% improvement over the baseline in natural language generation and 16.0% enhancement in clinical validity assessment, significantly boosting diagnostic accuracy and report standardization.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VAMP: Visual Attribute-Guided Multi-dimensional Perception for Nasal Endoscopy Report Generation

  • Xinpan Yuan,
  • Jianuo Ju,
  • Liujie Hua,
  • Changhong Zhang,
  • Mingzhu Huang,
  • Shaomin Xie,
  • Wenguang Gan

摘要

Automated Report Generation (ARG) enhances clinical decision- making through efficient and standardized report generation, playing an essential role in intelligent healthcare. Yet existing methods predominantly focus on single-frame analysis and fail to process multi-frame nasal endoscopy sequences. Conventional frameworks ignore spatiotemporal correlations between frames, resulting in fragmented reports that cannot model dynamic lesion progression. Two critical bottlenecks arise: (1) inter-frame interference, where anatomical similarities (e.g., inferior turbinate vs. nasal septum) induce semantic confusion (visual noise); (2) continuous-frame misclassification, where dynamic region transitions cause erroneous category assignments (classification noise). These issues degrade lesion localization accuracy and report consistency. To overcome these limitations, we propose the Visual Attribute-Guided Multi-dimensional Perception (VAMP) framework with two innovations: (1) a visual attribute-guided module that suppresses noise via spatiotemporal lesion tracking, and (2) a cross- dimensional perception aggregation module enabling anatomical pathological fusion through dual-stream spatial-channel coordination. Experiments on IRA-HUT-NA, IU-Xray, and MIMIC-CXR show VAMP achieves 25.4% improvement over the baseline in natural language generation and 16.0% enhancement in clinical validity assessment, significantly boosting diagnostic accuracy and report standardization.