Despite the significant progress in conventional image style transfer, the results remain unsatisfactory. Witnessed by the powerful generative capabilities of large-scale text-to-image diffusion models, style transfer based on these models have attracted widespread attention from researchers. Given the challenge of acquiring large, style-consistent image datasets, training-free diffusion-based style transfer methods have gained significant research interest. These methods avoid complex processes such as retraining the model and fine-tuning, which enhances their generalizability. However, they inevitably lead to issues such as style degradation, content disruption and localized style inconsistency. To address these problems, we propose a training-free style transfer method with Style Enhancement and Localized Style Consistency (SELSC). Our method comprises two core components. First, we introduce style enhancement adapter that integrates the IP-adapter with the StyleID framework to enable efficient and lightweight style injection. Second, we propose localized style consistency attention mechanism that aligns style consistency with local semantic similarities across image regions. Extensive experiments demonstrate that our method achieves a better balance between style and content compared to traditional and diffusion-based style transfer baselines.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SELSC: A Style Transfer Method with Style Enhancement and Localized Style Consistency

  • Xinjie Ruan,
  • Yaoze Zhang,
  • Zekun Tian,
  • Yun Sheng,
  • Cong Liu

摘要

Despite the significant progress in conventional image style transfer, the results remain unsatisfactory. Witnessed by the powerful generative capabilities of large-scale text-to-image diffusion models, style transfer based on these models have attracted widespread attention from researchers. Given the challenge of acquiring large, style-consistent image datasets, training-free diffusion-based style transfer methods have gained significant research interest. These methods avoid complex processes such as retraining the model and fine-tuning, which enhances their generalizability. However, they inevitably lead to issues such as style degradation, content disruption and localized style inconsistency. To address these problems, we propose a training-free style transfer method with Style Enhancement and Localized Style Consistency (SELSC). Our method comprises two core components. First, we introduce style enhancement adapter that integrates the IP-adapter with the StyleID framework to enable efficient and lightweight style injection. Second, we propose localized style consistency attention mechanism that aligns style consistency with local semantic similarities across image regions. Extensive experiments demonstrate that our method achieves a better balance between style and content compared to traditional and diffusion-based style transfer baselines.