CineScale: Tuning-Free High-Resolution Video Generation

1Nanyang Technological University 2Netflix Eyeline Studios
*Equal contribution ✉Corresponding authors

Watch the video above in 4K

CineScale is now integrated into Wan 🎉

TL;DR

Most video generators are trained at limited spatial resolutions due to the scarcity of high-resolution 4K video data and the prohibitive computational cost of large-scale training on such data. Most video diffusion models are trained on 720p videos and are therefore effectively limited to generating videos at similar resolutions during inference. To address this gap, we propose CineScale. CineScale, to the best of our knowledge, is the first tuning-free inference framework enabling pretrained video diffusion models to generate high-fidelity videos at resolutions far beyond those encountered during training, without any fine-tuning.

Method

Local detail stays precise Global context stays visible Positions fit the pretrained range

CineScale enables pretrained video diffusion models to generate far beyond their native training resolution without retraining. It partitions query tokens into spatial tiles while retaining access to all global keys and values, then uses Adaptively Rectified RoPE to preserve precise local geometry and keep long-range positional offsets within the range supported by the pretrained model.

For complete method details and technical derivations, please refer to the paper.

Read the paper

Results and Ablations

Baseline comparison

No Cascading

RoPE
NTK-RoPE
AR-RoPE (Ours)
No-cascading comparison across the three positional encoding methods. Hover over a video to inspect fine details.
Step ablation

Cascading Steps

5 cascading steps
15 cascading steps
25 cascading steps
25-step detail 5× crop
RoPE
NTK-RoPE
AR-RoPE (Ours)
Cascading-step ablation at 5, 15, and 25 steps. The final column permanently enlarges the same foreground-gnome region for direct comparison.
4K video generation comparison across Kling 3.0, LTX 2.0, Wan 2.2, UltraWan, and CineScale
4K comparison across recent video generation systems.

Quantitative evaluation

VBench Comparison

Comparison across target resolutions, covering semantic consistency, temporal quality, and perceptual fidelity.

VBench comparison across different target resolutions
Method Subject Consistency↑ Background Consistency↑ Temporal Flickering↑ Aesthetic Quality↑ Imaging Quality↑ Average↑
Tuning-Free
Wan2.1-720p 0.95700.96050.98450.56460.68280.8299
Wan2.1-1K 0.95400.96450.98980.49890.58260.7980
Wan2.1-4K 0.94700.97600.99500.28800.37400.7160
CineScale-2K Ours 0.97340.97770.97950.64880.71560.8581
Tuning-Based
UltraWan-1K 0.95860.96610.98530.56860.69660.8350
UltraWan-4K 0.95810.96110.97710.57690.71440.8375
UltraGen-1080P 0.97710.97770.99610.58190.73500.8536
UltraGen-4K 0.98540.98940.99330.57870.68320.8460
LUVE-2K 0.95830.96760.98180.59780.71150.8434
LUVE-4K 0.95360.96460.98090.58910.71330.8403
Best, second-best, and third-best results are shown in bold, underline, and italics. 4K baseline scores follow UltraGen.

Citation

If you find CineScale useful in your research or projects, consider citing our paper:

@article{chen2026cinescale,
  title={CineScale: Tuning-Free High-Resolution Video Generation},
  author={Chen, Gordon and Qiu, Haonan and Yu, Ning and Huang, Ziqi and Debevec, Paul and Liu, Ziwei},
  journal={arXiv preprint arXiv:2508.15774},
  year={2026}
}