Recent advances in deep learning and artificial intelligence make automated image description creation a promising way to improve radiological interpretation, precision, and uniformity. This requires computational models to interpret images from X-rays and generate clinical assessment descriptions. An innovative dual-transformer system for IU Chest X-ray and NIH Chest X-ray datasets is presented in the study. Vision transformers (ViT) and the Generative Pre-trained Transformers 3 (GPT-3) language model are used to research how they affect X-ray image elements. The vision transformer encoder extracts image features, while the GPT decoder provides X-ray image text. The suggested approach scores highest in BLEU-1 to BLEU-4 for the IU Chest X-ray dataset with 0.824, 0.807, 0.786, and 0.759. It scores 0.983, 0.789, and 0.704 in CIDEr, METEOR, and ROUGE-L, respectively. For the NIH Chest X-ray dataset, the model yields BLEU-1 to BLEU-4 values of 0.804, 0.798, 0.782, and 0.766. Its METEOR is 0.756, CIDEr is 0.915, and ROUGE-L is 0.710. This showcases the capacity to provide descriptions that align with references annotated by humans.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Utilizing Transformer-Based Image Descriptors for Improving Chest X-Ray Analysis

  • Lakshita Agarwal,
  • Bindu Verma

摘要

Recent advances in deep learning and artificial intelligence make automated image description creation a promising way to improve radiological interpretation, precision, and uniformity. This requires computational models to interpret images from X-rays and generate clinical assessment descriptions. An innovative dual-transformer system for IU Chest X-ray and NIH Chest X-ray datasets is presented in the study. Vision transformers (ViT) and the Generative Pre-trained Transformers 3 (GPT-3) language model are used to research how they affect X-ray image elements. The vision transformer encoder extracts image features, while the GPT decoder provides X-ray image text. The suggested approach scores highest in BLEU-1 to BLEU-4 for the IU Chest X-ray dataset with 0.824, 0.807, 0.786, and 0.759. It scores 0.983, 0.789, and 0.704 in CIDEr, METEOR, and ROUGE-L, respectively. For the NIH Chest X-ray dataset, the model yields BLEU-1 to BLEU-4 values of 0.804, 0.798, 0.782, and 0.766. Its METEOR is 0.756, CIDEr is 0.915, and ROUGE-L is 0.710. This showcases the capacity to provide descriptions that align with references annotated by humans.