DynSC/Project

VIDEO GENERATION · SUBJECT CONSISTENCY

Towards Subject Consistency over Dynamic Subject Sets
in Video Generation

Tongcheng ZhangJun ZhuJianfei Chen

Tsinghua University

Abstract

We argue that as video generation extends to longer durations, subject consistency should be evaluated over dynamic subject sets. We therefore introduce DynSC-Eval, an evaluation framework that dynamically tracks eligible subjects throughout their visible lifespans and measures local continuity and global identity preservation using six complementary object-level metrics, with explicit detection of inconsistency events. To validate its effectiveness, we design synthetic experiments that actively inject inconsistency events, demonstrating both the sensitivity of DynSC-Eval and the limitations of existing metrics. Evaluations of diverse models on 5s, 15s, and 60s video generation further reveal substantial subject consistency differences that are obscured by conventional metrics. Beyond evaluation, we construct rewards from DynSC-Eval and apply DiffusionNFT post-training in an autonomous-driving testbed. On 5s generation, our approach reduces the six inconsistency metrics by an average of 13.82% for Wan-2.1-1.3B and 5.66% for SANA-2B, with improvements also observed on the I2V model ReSim. Qualitative comparisons further demonstrate the effectiveness of our method. We then extend generation to 10s and 30s through curriculum learning and show that consistency optimization remains effective while largely preserving other capabilities.

01Subject Consistency Comparison5s · 15s · 60s02Autonomous-driving testbedBefore / after consistency training