SparkDiffusion combines sparse attention, few step distillation and FP8 kernels; its team reports an 81 frame, 720p Wan 2.1 video generated in 18 seconds instead of 4,769 on one RTX 5090—about 265× faster in that test... Its “high sparsity trap” describes training loss improving while final video quality stalls or d...
Published byEdited with GPT-6 LunaImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: How does the open-source SparkDiffusion framework from Peking University, Tsinghua University and Alibaba accelerate Wan 2.1 and Wan 2.2 vid. Article summary: SparkDiffusion speeds up Wan video generation by combining three changes: it computes far less attention, distils generation into a few steps, and runs the resulting model with FP8 fused kernels. The researchers report u. Topic tags: general, academic, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, char
SparkDiffusion is an open-source acceleration framework for Wan video models from researchers at Peking University, Tsinghua University and Alibaba. It combines sparse attention, few-step distillation and FP8 fused kernels. The project reports up to 265× faster generation in a specific Wan 2.1 benchmark; that result should not be read as a speedup for every model, video or hardware setup. 6
8
Video diffusion models process long sequences of visual information, making attention a major part of the computation. Sparse attention reduces how many interactions the model calculates. But the SparkDiffusion researchers describe a failure mode at extreme sparsity: step-local training loss can continue to fall even while the final generated video stops improving or gets worse. 18
That gap between a training signal and the finished video is the “high-sparsity trap.” SparkDiffusion’s proposed response is staged: first warm up the sparse model to establish a coarse prior, then use distillation aimed at correcting the final generation trajectory. 8
SparkDiffusion targets different parts of the generation process rather than relying on one optimization alone:
The overall speedup comes from stacking these changes: less attention computation, fewer generation steps and faster execution of the remaining work. It is not attributable to 97% sparsity by itself. 8
For an 81-frame, 720p Wan 2.1 T2V 14B video, the reported generation time drops from 4,769 seconds to 18 seconds on a single RTX 5090—about 265× faster. A reported H100 comparison drops from 1,757 seconds to 8 seconds, or roughly 220×. These are results for specified benchmarks and hardware, not guarantees for other settings. 6
The framework also reports a roughly 140× speedup for a Wan 2.1 1.3B, 480p case. That result illustrates why the headline figure should be read alongside the model and resolution being tested. 5
7
The project describes SparkDiffusion as maintaining strong visual quality at high sparsity. But speed figures alone do not demonstrate quality parity across prompts, resolutions or model variants. The high-sparsity trap itself is a reminder that improving a step-level training metric does not necessarily mean the final video improves. 8
18
Treat the reported speed and quality claims as specific to the configurations evaluated. The available benchmark summaries do not provide enough detail to conclude that every generated video will match the dense Wan reference.
The provided checkpoint listings include Wan 2.1 T2V 14B at 480p and 720p, as well as Wan 2.2 T2V A14B at 480p. The listed sparsity ranges vary: for example, the Wan 2.1 480p checkpoint supports 90% sparsity, while listed 720p variants cover up to 95% or 97%; the Wan 2.2 480p listing covers 90% to 95%. So 97% sparsity is not a setting supported by every released checkpoint. 9
10
11
12
The project presents SparkDiffusion as a framework that can extend beyond the listed Wan releases. The sources provided here do not establish a firm timetable or verified results for future models or hardware. 8
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
SparkDiffusion combines sparse attention, few step distillation and FP8 kernels; its team reports an 81 frame, 720p Wan 2.1 video generated in 18 seconds instead of 4,769 on one RTX 5090—about 265× faster in that test...
SparkDiffusion combines sparse attention, few step distillation and FP8 kernels; its team reports an 81 frame, 720p Wan 2.1 video generated in 18 seconds instead of 4,769 on one RTX 5090—about 265× faster in that test... Its “high sparsity trap” describes training loss improving while final video quality stalls or degrades; SparkDiffusion’s staged approach aims to address that mismatch.
Released SparkWan checkpoints cover selected Wan 2.1 and Wan 2.2 configurations, with different supported sparsity levels.