<p>Controllable image generation plays an important role in creative content synthesis; however, existing instruction-guided approaches often struggle to maintain stable multimodal semantic alignment and perceptual consistency under complex instructions. During iterative denoising, semantic drift and sampling-induced variability may lead to incomplete instruction execution and inconsistent visual quality. To address these challenges, this paper proposes Instruct-MR, a closed-loop diffusion-based framework that combines multimodal semantic alignment with preference-aware re-ranking. Specifically, a Qwen-VL multimodal large language model fine-tuned via Low-Rank Adaptation (LoRA) is employed to jointly encode textual instructions and visual inputs, thereby enhancing conditional semantic representations. In addition, a reconstruction-constrained optimization objective is introduced to facilitate conditional signal preservation during diffusion denoising while keeping most backbone diffusion parameters frozen. Furthermore, a preference-aware re-ranking module is incorporated at the inference stage to evaluate multiple generated candidates according to structural consistency, cross-modal semantic correspondence, and perceptual quality, thereby improving output reliability. Experimental results on several public benchmarks indicate that Instruct-MR consistently improves structural preservation and perceptual quality compared with representative baselines, including InstructPix2Pix, MGIE, SDEdit, and OmniGen2. These findings suggest that the integration of multimodal semantic alignment and preference-aware re-ranking contributes to more stable instruction-guided image generation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Diffusion based image generation with multi-modal semantic alignment and preference re-ranking, Instruct-MR

  • Jingyu Guo,
  • Yuan Wang,
  • Zhenhua Wu

摘要

Controllable image generation plays an important role in creative content synthesis; however, existing instruction-guided approaches often struggle to maintain stable multimodal semantic alignment and perceptual consistency under complex instructions. During iterative denoising, semantic drift and sampling-induced variability may lead to incomplete instruction execution and inconsistent visual quality. To address these challenges, this paper proposes Instruct-MR, a closed-loop diffusion-based framework that combines multimodal semantic alignment with preference-aware re-ranking. Specifically, a Qwen-VL multimodal large language model fine-tuned via Low-Rank Adaptation (LoRA) is employed to jointly encode textual instructions and visual inputs, thereby enhancing conditional semantic representations. In addition, a reconstruction-constrained optimization objective is introduced to facilitate conditional signal preservation during diffusion denoising while keeping most backbone diffusion parameters frozen. Furthermore, a preference-aware re-ranking module is incorporated at the inference stage to evaluate multiple generated candidates according to structural consistency, cross-modal semantic correspondence, and perceptual quality, thereby improving output reliability. Experimental results on several public benchmarks indicate that Instruct-MR consistently improves structural preservation and perceptual quality compared with representative baselines, including InstructPix2Pix, MGIE, SDEdit, and OmniGen2. These findings suggest that the integration of multimodal semantic alignment and preference-aware re-ranking contributes to more stable instruction-guided image generation.