Attention Is All You Need
Ashish Vaswani, Noam Shazeer et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.
Autoregressive video generation has been limited by exposure bias and the inability to generate long, coherent videos beyond the training context window. Self Forcing partially addresses exposure bias by training on self-generated histories, but it fails to provide supervision for how earlier latents should be encoded into the key-value cache for future frames. This paper identifies and solves the historical context-gradient gap, a fundamental limitation in current autoregressive video models. By enabling future losses to supervise past context encoding, SGF allows models to learn more effective causal memory, which is crucial for maintaining subject identity and temporal consistency over long horizons.
The ability to extrapolate to several-minute-long videos from a 5-second training window is remarkable and suggests that SGF may unlock practical applications in long-form video generation, such as movie production, virtual environments, and video game cutscenes. This work is particularly relevant as the field moves toward native autoregressive video generation without reliance on external modules or post-processing.
The paper reports that SGF achieves stronger native long-video extrapolation than Self Forcing across multiple metrics: subject identity, background/layout consistency, and temporal stability. The experiments cover both frame-wise and chunk-wise generation under different initialization conditions. The most striking result is that with only a 5-second training window, SGF can extrapolate to videos lasting several minutes, demonstrating a dramatic improvement in temporal coherence. No specific numerical metrics (e.g., FID, FVD) are provided in the abstract, but the qualitative improvements are emphasized.
SGF addresses a core limitation in autoregressive video generation, enabling models to learn more effective causal memory without architectural changes. This could reduce the need for large training datasets and long context windows, making long-form video generation more accessible. The approach is general and could be applied to other autoregressive generative tasks, such as audio or text generation, where long-range dependencies are critical. By releasing code and models, the authors aim to accelerate research in this direction.
Ashish Vaswani, Noam Shazeer et al.
Jakubův, Jan, Chvalovský, Karel et al.
Pauli Virtanen, Ralf Gommers et al.
Tom B. Brown, Benjamin Mann et al.