Research#Video generation#Benchmark
PKU, Tsinghua and Alibaba open-source SparkDiffusion: 265x single-GPU video
Open-source SparkDiffusion pushes 720P-14B video generation to 265x on one RTX 5090 via sparse attention, distillation and FP8; VBench-2.0 drops 0.5 points.

A joint team from Peking University, Tsinghua, Alibaba, UESTC and Harbin Institute of Technology open-sourced SparkDiffusion — described in the paper as the first unified open-source system combining sparse attention, few-step distillation and low-precision quantization for DiT video generation. Under best-case conditions (a single RTX 5090, the 720P-14B Wan2.1/Wan2.2 model pushed to 97% sparsity), generation accelerates 265x: 18 seconds for a 5-second clip, with VBench-2.0 dipping from 60.2 to 59.77. Paper, code and weights are all public.
The facts
- Method: RoLA sparse attention aligned to terminal supervision, plus CrossDistill trajectory-level hybrid distillation (3-step, CFG-free student), plus FP8 quantization.
- Speed numbers (paper-reported): 265x at 97% sparsity on a single RTX 5090 for 720P-14B (18 seconds per 5-second clip); 201x at 90% sparsity; 220x on H100 (8 seconds); 140x for 480P-1.3B.
- Quality cost: VBench-2.0 from 60.2 to 59.77 (-0.5); the paper concedes quality degrades slightly at 97% sparsity.
- Core finding: a “high-sparsity trap” — training loss keeps falling while generation quality collapses at extreme sparsity; the fix is terminal-aligned supervision via RoLA plus CrossDistill.
- Availability: arXiv 2609.23153, code at AlibabaResearch/SparkDiffusion, weights on Hugging Face; a fresh, non-peer-reviewed preprint.
Our take
265x is the ceiling number, but even the conservative 90%-sparsity setting (201x) pulls a 14B video model from minutes into tens of seconds on a consumer card — a real pipeline cost drop. Teams running text-to-video tools like Story-Flicks or their own video pipelines should rerun a comparison at their model and resolution. One caution: the benchmarks come from the authors’ own harness; validate quality on your own footage before wiring it in.