This guide consolidates the public documentation for DanceOPD: the algorithm, configuration interface, smoke tests, public data preprocessing, SFT-teacher setup, and backend extension path.
- Release scope
- Method
- Configuration
- Smoke tests
- Public data preparation
- Teacher field preparation
- Run DanceOPD
- Custom backends
- Main experimental settings
- Interpretation notes
DanceOPD trains one student by querying frozen teacher velocity fields on states visited by the current student.
For a route (m), condition (c), student rollout state (z_t^\theta), and frozen teacher field (v_m):
[ \mathcal{L}{\text{DanceOPD}} = \mathbb{E}{m,c,t}\left[ \left|v_\theta(\operatorname{sg}(z_t^\theta), t, c)
- v_m(\operatorname{sg}(z_t^\theta), t, c)\right|_2^2 \right]. ]
student rollout: noise -> ... -> z_t -> ... -> clean latent
teacher query: v_T(z_t, t, c)
student query: v_S(z_t, t, c)
loss: || v_S - stopgrad(v_T) ||^2
| Component | Default |
|---|---|
| Route selection | hard route, one teacher per step |
| State source | current student trajectory |
| Query count | K=1 |
| Query bias | low-noise / semantic-side |
| Objective | direct velocity MSE |
| Student update | LoRA |
The provided configs intentionally contain no local paths. Fill the null
fields before training, or override them on the command line.
danceopd-train --config configs/sd35_danceopd.yaml \
--set model.pretrained_model='<MODEL_DIR>' \
--set data.prompts_csv='<PROMPTS_CSV>' \
--set training.output_dir='<OUTPUT_DIR>'| Field | Default |
|---|---|
| LoRA rank | 128 |
| rollout steps | 16 |
query states K |
1 |
| query bias | low_t |
| learning rate | 2e-4 |
| grad accumulation | 4 |
| save interval | 300 |
| max train steps | 3000 |
| precision | bf16 |
No random seed is set by the public configs.
| Template | Backend | Purpose |
|---|---|---|
configs/paper/zimage_edit_fusion.yaml |
Z-Image / DiffSynth | edit-fusion reproduction template |
configs/paper/sd35_edit_fusion.yaml |
SD3.5 / Diffusers | edit-fusion reproduction template |
Both templates are path-free. Fill data.prompts_csv, training.output_dir,
model paths, and teacher base_ckpt / lora_dir fields before launching.
pip install -e ".[all,smoke]"
bash scripts/bootstrap_smoke.shThis downloads a tiny official DiffSynth example subset and runs the full train/update/save loop with the zero-weight toy backend. It verifies code health without downloading large public model weights.
Expected outputs:
outputs/smoke_toy/step-1/toy_student.safetensors
outputs/smoke_toy/step-2/toy_student.safetensors
bash scripts/smoke_zimage.sh --dry-run
bash scripts/smoke_sd35.sh --dry-runThese validate dependency checks, config loading, prompt extraction, backend selection, and launch plumbing without loading all model weights.
RUN_TRAIN=1 BACKEND=zimage bash scripts/bootstrap_smoke.sh
RUN_TRAIN=1 BACKEND=sd35 bash scripts/bootstrap_smoke.shFirst runs may download large upstream model weights. This is intentionally not the default smoke path.
bash scripts/prepare_diffsynth_example.shThe helper downloads DiffSynth-Studio/diffsynth_example_dataset, extracts
prompts from the selected subset metadata, and writes:
data/diffsynth_example_dataset/danceopd_prompts.csv
To reuse an already downloaded dataset:
DIFFSYNTH_NO_DOWNLOAD=1 bash scripts/prepare_diffsynth_example.shThe public helper accepts either a Hugging Face dataset ID or a local
.csv / .json / .jsonl metadata file.
Install dataset support:
pip install -e ".[data]"Prepare edit-pair CSV for SFT teacher training:
python examples/prepare_omniedit.py \
--input TIGER-Lab/OmniEdit-Filtered-1.2M \
--output data/omniedit_sft_pairs.csv \
--format sft_pairs \
--max-rows 100000Prepare prompt-only CSV for DanceOPD:
python examples/prepare_omniedit.py \
--input TIGER-Lab/OmniEdit-Filtered-1.2M \
--output data/omniedit_prompts.csv \
--format promptsNormalized SFT-pair schema:
uid,source_image,edited_image,prompt,task,caption_dictcaption_dict is a JSON cell containing at least {"prompt": ...}.
DanceOPD distills teacher velocity fields. In the paper setting, those teachers are SFT LoRAs trained on edit data. The public repo does not ship those LoRAs or internal checkpoints, but it defines the interface expected by the DanceOPD stage.
| DanceOPD route | Example OmniEdit-style tasks | Purpose |
|---|---|---|
local_edit |
add, remove, replace, color, material, object | local edit preservation |
global_edit |
background, environment, weather, style, tone | global transformations |
t2i |
no LoRA | base text-to-image anchor |
The exact split can differ from ours. Each route only needs to represent a semantically meaningful frozen teacher field.
teachers:
- name: local_edit
base_ckpt: null
lora_dir: /path/to/local_edit_lora
weight: 1.0
- name: global_edit
base_ckpt: null
lora_dir: /path/to/global_edit_lora
weight: 1.0
- name: t2i
base_ckpt: null
lora_dir: null
weight: 1.0The same interface supports LoRA and non-LoRA teachers:
| Teacher case | base_ckpt |
lora_dir |
Behavior |
|---|---|---|---|
| Base route | null |
null |
use the pretrained base model as the frozen teacher |
| Full checkpoint teacher | /path/to/full_or_merged_ckpt |
null |
load a full, non-LoRA teacher checkpoint |
| LoRA teacher | null |
/path/to/peft_lora |
load the base model, then merge one PEFT LoRA |
| Full checkpoint + LoRA | /path/to/full_or_merged_ckpt |
/path/to/peft_lora |
load the full checkpoint first, then merge the LoRA |
Teacher LoRAs are merged into clean frozen teacher modules. They are not stacked on top of the student's trainable LoRA adapter.
For Z-Image, DiffSynth-Studio's Z-Image training scripts are a natural public starting point. For SD3.5, any PEFT-compatible SD3.5 LoRA SFT trainer can be used as long as the resulting LoRA can be loaded by PEFT.
Example Z-Image launch:
accelerate launch -m danceopd.cli.train \
--config configs/paper/zimage_edit_fusion.yaml \
--set data.prompts_csv=data/omniedit_prompts.csv \
--set teachers.0.lora_dir='<LOCAL_EDIT_LORA>' \
--set teachers.1.lora_dir='<GLOBAL_EDIT_LORA>' \
--set training.output_dir='<OUTPUT_DIR>'Example SD3.5 launch:
accelerate launch -m danceopd.cli.train \
--config configs/paper/sd35_edit_fusion.yaml \
--set model.pretrained_model='<SD35_MODEL_DIR>' \
--set data.prompts_csv=data/omniedit_prompts.csv \
--set teachers.0.lora_dir='<LOCAL_EDIT_LORA>' \
--set teachers.1.lora_dir='<GLOBAL_EDIT_LORA>' \
--set training.output_dir='<OUTPUT_DIR>'A backend only needs to implement DanceOPDBackend:
class MyBackend(DanceOPDBackend):
def prepare(self): ...
def compute_loss(self, prompt, route): ...
def backward(self, loss): ...
def optimizer_step(self): ...
def save(self, step): ...The core engine handles:
- CSV prompt iteration,
- teacher route sampling,
- gradient accumulation,
- logging,
- checkpoint cadence.
Use danceopd.backends.sd35_diffusers.SD35Backend as the clearest real-model
reference and danceopd.backends.toy.ToyBackend as the smallest complete
example.
| Component | Setting |
|---|---|
| objective | velocity MSE |
| routing | hard route, one teacher per step |
| query source | current student rollout |
| query states | K=1 |
| query bias | low-noise / semantic-side |
| rollout steps | 16 |
| LoRA rank | 128 |
| learning rate | 2e-4 |
| grad accumulation | 4 |
| precision | bf16 |
scripts/bootstrap_smoke.shverifies code health, not paper quality.- Real paper numbers require comparable teacher fields.
- Weights are intentionally user-provided.
- Teacher SFT code is not the main release path because SFT pipelines are often tightly coupled to base model, image-pair loader, VAE cache format, and local training environment.
- The DanceOPD stage is model-agnostic through the backend interface once compatible teacher fields are available, either as full checkpoints or PEFT LoRAs.
Back to README.md

