Fetal Ultrasound Video Representation Learning Using Contrastive Rubik’s Cube Recovery
摘要
Contrastive learning (CL), which relies on the contrast between positive and negative pairs, has become the leading paradigm in self-supervised learning. In this paper, we propose a self-supervised learning framework, the feature-level Contrastive Rubik’s Cube Recovery (CRCR). CRCR creates contrastive sub-cube pairs from ultrasound video, which capture local spatio-temporal ultrasound features, unlike traditional CL methods which are spatial and work at the global frame level. This approach learns a representation with both intra- and inter-feature contrast to provide strong local feature discrimination. The proposed method is validated on two fetal ultrasound video tasks. Extensive experiments demonstrate that our approach is effective for learning representations that transfer to both in-domain (second-trimester) and cross-domain (first-trimester) clinical downstream classification tasks. In particular, CRCR outperforms four state-of-the-art contrastive learning-based methods on the in-domain task by 3.8%, 2.0%, 1.9% and 1.1%, with each improvement being statistically significant.