This study explores implementation and evaluation a multimodal (visual & text) language model for auto-captioning and visual question answering (VQA) tasks in histopathology images. The selected architecture integrates CLIP-prefix with GPT-2 models, combining the image understanding capabilities of vision models with the text generation abilities of language models. This approach demonstrates the potential of multimodal models to serve as supportive tools in medical image interpretation, assisting tasks such as auto-captioning and visual question answering (VQA), while maintaining manageable computational requirements. For auto-captioning, experiments were conducted using the ARCH dataset, which includes histopathology images and textual descriptions from scientific sources (books and journal papers). The best-performing model achieved an accuracy of \(82\%\) , a BLEU score of \(71\%\) , and a CIDER score of 4.9 metrics commonly used to assess the quality of generated textual descriptions. Qualitative analysis of test images suggests a strong alignment between predicted and actual descriptions. For VQA, a model was optimized using the PathVQA dataset, which contains diverse pathological images paired with questions and answers. The optimized model achieved an accuracy of \(84\%\) , a BLEU score of \(67\%\) , and a CIDER score of 1.9, reflecting strong alignment between generated answers and reference responses. Qualitative evaluations revealed a high consistency between predicted answers and corresponding visual content. These preliminary findings suggest that multimodal models can serve as intelligent assistants (“co-pathologists”) to support pathology workflows. By automating and improving the accuracy of image-based textual descriptions, these models could enhance diagnostic precision, reduce errors in pathology reports, and optimize clinical decision-making, especially in critical cases like cancer diagnosis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring a Multimodal Language Model for Auto-Captioning and Visual Question Answering in Histopathology Images

  • Johan Sánchez,
  • Angel Cruz-Roa

摘要

This study explores implementation and evaluation a multimodal (visual & text) language model for auto-captioning and visual question answering (VQA) tasks in histopathology images. The selected architecture integrates CLIP-prefix with GPT-2 models, combining the image understanding capabilities of vision models with the text generation abilities of language models. This approach demonstrates the potential of multimodal models to serve as supportive tools in medical image interpretation, assisting tasks such as auto-captioning and visual question answering (VQA), while maintaining manageable computational requirements. For auto-captioning, experiments were conducted using the ARCH dataset, which includes histopathology images and textual descriptions from scientific sources (books and journal papers). The best-performing model achieved an accuracy of \(82\%\) , a BLEU score of \(71\%\) , and a CIDER score of 4.9 metrics commonly used to assess the quality of generated textual descriptions. Qualitative analysis of test images suggests a strong alignment between predicted and actual descriptions. For VQA, a model was optimized using the PathVQA dataset, which contains diverse pathological images paired with questions and answers. The optimized model achieved an accuracy of \(84\%\) , a BLEU score of \(67\%\) , and a CIDER score of 1.9, reflecting strong alignment between generated answers and reference responses. Qualitative evaluations revealed a high consistency between predicted answers and corresponding visual content. These preliminary findings suggest that multimodal models can serve as intelligent assistants (“co-pathologists”) to support pathology workflows. By automating and improving the accuracy of image-based textual descriptions, these models could enhance diagnostic precision, reduce errors in pathology reports, and optimize clinical decision-making, especially in critical cases like cancer diagnosis.