A half-day tutorial on RL-based and non-RL alignment methods for visual diffusion models.
Aligning diffusion models with human preferences has become an increasingly important research area in visual generation. This half-day tutorial gives a structured introduction to the area, organized around the four main families of alignment methods: direct reward backpropagation, which differentiates a reward through the sampling process; DPO-style preference optimization, which learns from pairs of preferred and rejected samples; GRPO-based online RL, which samples groups of outputs and learns from their relative rewards; and forward-process RL, which reformulates the alignment objective on the forward diffusion process, close to the pre-training objective.
For each family, we cover the core formulation, representative methods, and their strengths and limitations, and we close with open challenges and future directions.
Liang Zheng
Jie Liu
Shuchen Xue
Slides will be posted here after the tutorial.