Unsupervised natural scene image-to-image translation via independence-enhanced latent space
摘要
Natural scene image-to-image translation aims to keep the spatial content of one natural scene image and reproduce the style from another natural scene image. As content preservation and stylization are learned from unpaired and richly detailed images, it is a challenging task to achieve accurate and visually pleasing unsupervised image-to-image translation for natural scenes. To this end, we make an assumption of semantic feature recombination and propose an effective generative model via independence-enhanced latent space for natural scene image-to-image translation. In terms of content preservation, we design a self-attention guided skip connection (SASC), in which high-level semantic information is enhanced and then propagated to low-level layers to facilitate the restoration of local texture details. As for image stylization, we devise an imbalanced layer-instance normalization (ImbLIN) to strengthen the style-related feature learning. In addition, we introduce an orthogonal constraint (OC) to enforce the latent features to be independent of each other, which can enhance the ability of unsupervised disentanglement. Extensive experiments on three publicly available datasets show that our model can generate natural scene images with richer content details and more accurate styles, and achieve better qualitative and quantitative results than previous methods.