In routine clinical practice, a vast amount of data is generated, including myriads of ultrasound recordings. However, their annotation and interpretation are labor-intensive; thus, a method that can incorporate this unlabeled data into deep learning pipelines would be highly beneficial. Video masked autoencoders (VideoMAE) are state-of-the-art pre-training techniques and have performed exceptionally well in various computer vision tasks. Accordingly, we hypothesized that a VideoMAE pre-trained on a large unlabeled dataset of ultrasound recordings could also perform well in a downstream task following supervised training on a smaller but labeled dataset. Nevertheless, we found that the conventional masking strategy of the VideoMAE pipeline may perform sub-optimally in the specific domain of ultrasound videos. Motivated by this, we proposed a novel region of interest (ROI)-aware masking method that considers the specific characteristics of this domain. We demonstrated that applying our method instead of the conventional masking strategy significantly improves the VideoMAE’s performance in clinically relevant downstream tasks, even when we reduced the labeled training dataset to one-tenth of its original sample size. The source code for this paper is available at https://github.com/szadam96/ROI-aware-masking .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Masked Autoencoders for Medical Ultrasound Videos Using ROI-Aware Masking

  • Ádám Szijártó,
  • Bálint Magyar,
  • Thomas Á. Szeier,
  • Máté Tolvaj,
  • Alexandra Fábián,
  • Bálint K. Lakatos,
  • Zsuzsanna Ladányi,
  • Zsolt Bagyura,
  • Béla Merkely,
  • Attila Kovács,
  • Márton Tokodi

摘要

In routine clinical practice, a vast amount of data is generated, including myriads of ultrasound recordings. However, their annotation and interpretation are labor-intensive; thus, a method that can incorporate this unlabeled data into deep learning pipelines would be highly beneficial. Video masked autoencoders (VideoMAE) are state-of-the-art pre-training techniques and have performed exceptionally well in various computer vision tasks. Accordingly, we hypothesized that a VideoMAE pre-trained on a large unlabeled dataset of ultrasound recordings could also perform well in a downstream task following supervised training on a smaller but labeled dataset. Nevertheless, we found that the conventional masking strategy of the VideoMAE pipeline may perform sub-optimally in the specific domain of ultrasound videos. Motivated by this, we proposed a novel region of interest (ROI)-aware masking method that considers the specific characteristics of this domain. We demonstrated that applying our method instead of the conventional masking strategy significantly improves the VideoMAE’s performance in clinically relevant downstream tasks, even when we reduced the labeled training dataset to one-tenth of its original sample size. The source code for this paper is available at https://github.com/szadam96/ROI-aware-masking .