ECCV 2026 · Tutorial · Malmö, Sweden

Aligning Diffusion Models with Human Preferences

A half-day tutorial on RL-based and non-RL alignment methods for visual diffusion models.

Wednesday, 9 September 2026
Afternoon session
13:30 – 17:30
Central European Summer Time (CEST)
Quality View Oresundssalen - 2
Quality Hotel View, next to Malmömässan
Overview

Aligning diffusion models with human preferences has become an increasingly important research area in visual generation. This half-day tutorial gives a structured introduction to the area, organized around the four main families of alignment methods: direct reward backpropagation, which differentiates a reward through the sampling process; DPO-style preference optimization, which learns from pairs of preferred and rejected samples; GRPO-based online RL, which samples groups of outputs and learns from their relative rewards; and forward-process RL, which reformulates the alignment objective on the forward diffusion process, close to the pre-training objective.

For each family, we cover the core formulation, representative methods, and their strengths and limitations, and we close with open challenges and future directions.

What you will learn

  1. Problem setupWhat alignment means for visual diffusion models
  2. Direct reward backpropagationAligning diffusion models by backpropagating reward gradients through the sampling process
  3. DPO-style alignmentPreference optimization for diffusion models
  4. GRPO-based online RLLearning from the relative rewards of a group of samples
  5. Forward-process RLReformulating alignment on the forward diffusion process for more accurate likelihood estimation
  6. Looking aheadOpen challenges and future directions
People

Organizers & Speakers

Schedule

Wednesday, 9 September 2026 · 13:30 – 17:30 CEST

Room: Quality View Oresundssalen - 2  ·  Online: Join on Zoom

13:30 – 13:4515 min
Opening remarks: An introduction to diffusion model alignment
Liang Zheng
13:45 – 14:2540 min
Reward gradients tell you where to go: Aligning diffusion models with direct backpropagation
Zhanhao Liang
14:25 – 14:3510 min
Break
14:35 – 15:2045 min
Learning from preferences: DPO-style alignment for diffusion models
Zhanhao Liang
15:45 – 15:5510 min
Coffee break
15:55 – 16:4045 min
Sample, compare, improve: GRPO-based online RL for diffusion models
Jie Liu
16:40 – 16:455 min
Short break
16:45 – 17:3045 min
Rethinking the objective: Forward-process RL for diffusion models
Shuchen Xue
Materials

SlidesComing soon

Slides will be posted here after the tutorial.