Clean Forcing: Drift-Resistant Autoregressive Video Diffusion with a Frozen Base

1Northeastern University   2Seoul National University   3NVIDIA
†Co-last authors

The last 20 s of an 80 s rollout of OmniDreams, an autoregressive driving world model built on Cosmos-Predict2.5, before (left) and after (right) applying the Clean Forcing corrector (strength 0.25α*(t)). The base degrades into blotches while the corrected rollout keeps the road scene. Same conditioning and seed.

Abstract

Autoregressive (AR) video diffusion models generate streaming video of arbitrary length by conditioning each new chunk of frames on those already generated. However, they suffer from exposure bias, where a model trained on clean context must continue from its own imperfect outputs, and errors compound until quality collapses within seconds (drift). Existing methods either require cluster-scale base retraining, apply training-free inference corrections with inconsistent results, or train frozen-base correctors for sampling and cache errors. We instead find that drift is largely deterministic and therefore learnable, as holding the noisy state fixed while varying only the history reveals a counterfactual velocity gap that is ~95% systematic across noise seeds.

Based on this observation, we introduce Clean Forcing, which trains a 5.9M-parameter LoRA corrector on the frozen base model to regress this gap, with a closed-loop stage that adds DAgger-style aggregation and a drift-contraction objective to expose the corrector to its own rollouts. Because the clean histories can be the base model's own single-shot generations, training needs no real videos, and the merged corrector adds no inference cost. In 50 s text-to-video generation, Clean Forcing reduces Δ-drift from +11.3 to +1.65 without real video and to −0.16 with 40 real clips, improving on the best published result of +3.57 with over 150× less data. Extensive experiments demonstrate that Clean Forcing outperforms prior methods in aesthetic quality and data efficiency with reduced drift for long-horizon video generation.

−0.16
Δ-drift over 50 s, from +11.27 (40 real clips)
+1.65
Δ-drift with zero real videos
5.9M
trainable parameters (0.4% of the base)
0
added inference cost once merged

Method

Clean Forcing overview (tap to enlarge)

(1) At the same noisy state zt, a drifted history hgen and a clean history hclean induce different velocities. About 95% of this gap is systematic across noise seeds, so it is learnable. (2) A LoRA student on the frozen base, conditioned on hgen, learns to match the frozen teacher's prediction under hclean at the same state. The clean histories can be the base model's own single-shot generations. (3) Closed-loop training adds DAgger aggregation over the corrector's own rollouts and a drift-contraction loss that backpropagates the next chunk's error through each committed prediction. At deployment the LoRA merges into the base weights.

Long-Horizon Results

Per-frame MUSIQ over 50 seconds

Per-frame quality over 50 s (mean MUSIQ ± SEM, 128 held-out MovieGen prompts). The adapted base and the training-free heuristics decay within about 15 s, while both Clean Forcing variants stay flat, on a frozen base trained on a single GPU in about a day.

Method comparisons over 50 s

Top row: adapted base, context noising, history guidance. Bottom row: Self Forcing, Clean Forcing (zero real videos), Clean Forcing (+40 real clips). All six are built on Wan2.1-T2V-1.3B. Same prompt and seed in every tile.

Transfer to Other Bases

Wan2.1-14B (10× larger), same recipe

Adapted 14B base (left) vs. + Clean Forcing (right).
Adapted 14B base (left) vs. + Clean Forcing (right).

Independently trained Causal-Forcing base, no per-model tuning

Causal-Forcing base (left) vs. + Clean Forcing (right).
Causal-Forcing base (left) vs. + Clean Forcing (right).

Other autoregressive world models

HY-WorldPlay (5B, camera-controlled) and OmniDreams (2B, action-conditioned driving), shown late in the rollout where the base drifts. These hosts use a reduced correction strength: 0.5α*(t) for HY-WorldPlay and 0.25α*(t) for OmniDreams.

HY-WorldPlay: walking toward a windmill.
HY-WorldPlay: cabin in snow.
HY-WorldPlay: castle.
OmniDreams: night drive.

BibTeX

@article{wang2026cleanforcing,
  title={Clean Forcing: Drift-Resistant Autoregressive Video Diffusion with a Frozen Base},
  author={Wang, Wenqing and Shin, Joonghyuk and Tremblay, Jonathan and Song, Chan Hee and Fu, Yun},
  journal={arXiv preprint},
  year={2026}
}