CineScale is now integrated into Wan 🎉
TL;DR
Most video generators are trained at limited spatial resolutions due to the scarcity of high-resolution 4K video data and the prohibitive computational cost of large-scale training on such data. Most video diffusion models are trained on 720p videos and are therefore effectively limited to generating videos at similar resolutions during inference. To address this gap, we propose CineScale. CineScale, to the best of our knowledge, is the first tuning-free inference framework enabling pretrained video diffusion models to generate high-fidelity videos at resolutions far beyond those encountered during training, without any fine-tuning.
Method
CineScale enables pretrained video diffusion models to generate far beyond their native training resolution without retraining. It partitions query tokens into spatial tiles while retaining access to all global keys and values, then uses Adaptively Rectified RoPE to preserve precise local geometry and keep long-range positional offsets within the range supported by the pretrained model.
For complete method details and technical derivations, please refer to the paper.
Read the paperResults and Ablations
No Cascading
Cascading Steps
4K Video Gallery
Hover to Zoom In
Quantitative evaluation
VBench Comparison
Comparison across target resolutions, covering semantic consistency, temporal quality, and perceptual fidelity.
| Method | Subject Consistency↑ | Background Consistency↑ | Temporal Flickering↑ | Aesthetic Quality↑ | Imaging Quality↑ | Average↑ |
|---|---|---|---|---|---|---|
| Tuning-Free | ||||||
| Wan2.1-720p | 0.9570 | 0.9605 | 0.9845 | 0.5646 | 0.6828 | 0.8299 |
| Wan2.1-1K | 0.9540 | 0.9645 | 0.9898 | 0.4989 | 0.5826 | 0.7980 |
| Wan2.1-4K | 0.9470 | 0.9760 | 0.9950 | 0.2880 | 0.3740 | 0.7160 |
| CineScale-2K Ours | 0.9734 | 0.9777 | 0.9795 | 0.6488 | 0.7156 | 0.8581 |
| Tuning-Based | ||||||
| UltraWan-1K | 0.9586 | 0.9661 | 0.9853 | 0.5686 | 0.6966 | 0.8350 |
| UltraWan-4K | 0.9581 | 0.9611 | 0.9771 | 0.5769 | 0.7144 | 0.8375 |
| UltraGen-1080P | 0.9771 | 0.9777 | 0.9961 | 0.5819 | 0.7350 | 0.8536 |
| UltraGen-4K | 0.9854 | 0.9894 | 0.9933 | 0.5787 | 0.6832 | 0.8460 |
| LUVE-2K | 0.9583 | 0.9676 | 0.9818 | 0.5978 | 0.7115 | 0.8434 |
| LUVE-4K | 0.9536 | 0.9646 | 0.9809 | 0.5891 | 0.7133 | 0.8403 |
Citation
If you find CineScale useful in your research or projects, consider citing our paper:
@article{chen2026cinescale,
title={CineScale: Tuning-Free High-Resolution Video Generation},
author={Chen, Gordon and Qiu, Haonan and Yu, Ning and Huang, Ziqi and Debevec, Paul and Liu, Ziwei},
journal={arXiv preprint arXiv:2508.15774},
year={2026}
}