Video sketching using multi-domain guidance and implicit encoding
摘要
Sketch data are a common element in visual communication. While synthesizing sketches from photographs has been extensively explored, creating sketches from video remains a complex challenge due to its inherent intricacy and the necessity for temporal consistency. This study delves into the generation of a sequence of vector sketches from a video clip. We have developed an optimization framework that utilizes the CLIP perceptual loss with guidance from multiple domains, including natural images and stylized line drawings. This approach aids in capturing the prominent visual content within a complex scene. We initialize the sketches by propagating control points from the keyframes through the video content deformation field. These initial points are implicitly encoded and serve as input to a transformer network that predicts the control point offsets for each frame. We also conduct an additional temporal refinement stage by using more precise initial points for optimization. Experimental results on the DAVIS video dataset demonstrate that our method successfully delivers high visual fidelity and temporal consistency.