Research#Video generation#Benchmark

PKU, Tsinghua and Alibaba open-source SparkDiffusion: 265x single-GPU video

Open-source SparkDiffusion pushes 720P-14B video generation to 265x on one RTX 5090 via sparse attention, distillation and FP8; VBench-2.0 drops 0.5 points.

A professional video camera and microphone on a desk

A joint team from Peking University, Tsinghua, Alibaba, UESTC and Harbin Institute of Technology open-sourced SparkDiffusion — described in the paper as the first unified open-source system combining sparse attention, few-step distillation and low-precision quantization for DiT video generation. Under best-case conditions (a single RTX 5090, the 720P-14B Wan2.1/Wan2.2 model pushed to 97% sparsity), generation accelerates 265x: 18 seconds for a 5-second clip, with VBench-2.0 dipping from 60.2 to 59.77. Paper, code and weights are all public.

The facts

  • Method: RoLA sparse attention aligned to terminal supervision, plus CrossDistill trajectory-level hybrid distillation (3-step, CFG-free student), plus FP8 quantization.
  • Speed numbers (paper-reported): 265x at 97% sparsity on a single RTX 5090 for 720P-14B (18 seconds per 5-second clip); 201x at 90% sparsity; 220x on H100 (8 seconds); 140x for 480P-1.3B.
  • Quality cost: VBench-2.0 from 60.2 to 59.77 (-0.5); the paper concedes quality degrades slightly at 97% sparsity.
  • Core finding: a “high-sparsity trap” — training loss keeps falling while generation quality collapses at extreme sparsity; the fix is terminal-aligned supervision via RoLA plus CrossDistill.
  • Availability: arXiv 2609.23153, code at AlibabaResearch/SparkDiffusion, weights on Hugging Face; a fresh, non-peer-reviewed preprint.

Our take

265x is the ceiling number, but even the conservative 90%-sparsity setting (201x) pulls a 14B video model from minutes into tens of seconds on a consumer card — a real pipeline cost drop. Teams running text-to-video tools like Story-Flicks or their own video pipelines should rerun a comparison at their model and resolution. One caution: the benchmarks come from the authors’ own harness; validate quality on your own footage before wiring it in.