<p>Aiming at the key problems of Low-Rank Adaptation (LoRA) methods in diffusion model fine-tuning, this study proposes a two-stage dynamic optimization method based on task relevance. Existing methods usually employ fixed probability density functions (e.g., uniform/normal distributions) for time-step sampling during LoRA training, relying on subjective prior knowledge, leading to less effective training results. Concurrently, static weight scaling factors during inference create conflicts between task adaptation capability and model stability.To resolve these issues, we innovatively introduce a Time-step Importance Probability Density Function (TIPDF) and dynamic weight scaling mechanism. In the training stage, the loss distribution of different denoising time-step is quantitatively analyzed by kernel density estimation, and a task-oriented TIPDF function is established to guide the model to focus on the feature learning of key time-steps. In the inference stage, a nonlinear Dynamic Weight Scaling Function (DWSF) is designed based on the TIPDF to realize the adaptive adjustment of LoRA weight scaling factor <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\alpha \)</EquationSource> </InlineEquation> at the time-step. Experiments show that in the In-Context LoRA generation task of the Flux1.dev model, in the training stage, the TIPDF sampling-based approach greatly improves training results. During the inference stage when task adaptation accuracy reaches 100%, our method achieves a 14.8% improvement in Feature Variance and an 8.2% increase in Structure Similarity Index Measure compared to conventional approaches, with a concurrent 16.3% reduction in Texture Deviation. Furthermore, additional experiments on cross-model and cross-task scenarios, as well as comparisons with multiple LoRA variants, demonstrate that the proposed TIPDF-DWSF framework maintains stable performance gains and strong generalization capability across different architectures and fine-tuning paradigms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TIPDF-DWSF: a task-oriented two-stage optimization framework for diffusion model LoRA fine-tuning

  • BaiHao Zhang,
  • Ting Lu,
  • Dan Qin

摘要

Aiming at the key problems of Low-Rank Adaptation (LoRA) methods in diffusion model fine-tuning, this study proposes a two-stage dynamic optimization method based on task relevance. Existing methods usually employ fixed probability density functions (e.g., uniform/normal distributions) for time-step sampling during LoRA training, relying on subjective prior knowledge, leading to less effective training results. Concurrently, static weight scaling factors during inference create conflicts between task adaptation capability and model stability.To resolve these issues, we innovatively introduce a Time-step Importance Probability Density Function (TIPDF) and dynamic weight scaling mechanism. In the training stage, the loss distribution of different denoising time-step is quantitatively analyzed by kernel density estimation, and a task-oriented TIPDF function is established to guide the model to focus on the feature learning of key time-steps. In the inference stage, a nonlinear Dynamic Weight Scaling Function (DWSF) is designed based on the TIPDF to realize the adaptive adjustment of LoRA weight scaling factor \(\alpha \) at the time-step. Experiments show that in the In-Context LoRA generation task of the Flux1.dev model, in the training stage, the TIPDF sampling-based approach greatly improves training results. During the inference stage when task adaptation accuracy reaches 100%, our method achieves a 14.8% improvement in Feature Variance and an 8.2% increase in Structure Similarity Index Measure compared to conventional approaches, with a concurrent 16.3% reduction in Texture Deviation. Furthermore, additional experiments on cross-model and cross-task scenarios, as well as comparisons with multiple LoRA variants, demonstrate that the proposed TIPDF-DWSF framework maintains stable performance gains and strong generalization capability across different architectures and fine-tuning paradigms.