Semantic Video Diffusion Models for Long Echocardiogram Generation
摘要
Ultrasound imaging is the primary modality used in cardiology. However, the development and training of new algorithms are often hindered by the restrictions surrounding access to medical datasets and the shortage of sufficiently diverse annotated data. To address these challenges, a novel method for generating synthetic cardiac ultrasound sequences is introduced. This method combines two features: (a) the generation of long sequences using diffusion models, producing sequences with a number of frames exceeding the fixed number used during training, and (b) the guidance by anatomical semantic labels, enabling the creation of label-specific ultrasound data. Evaluations were conducted to assess both the temporal consistency of the synthetic sequences and their fidelity to the provided semantic labels.