Ted Hisokawa
Apr 22, 2025 02:14
Chipmunk leverages dynamic sparsity to speed up diffusion transformers, attaining important speed-ups in video and picture technology with out further coaching.
Chipmunk, a novel strategy to accelerating diffusion transformers, has been launched by Collectively.ai, promising substantial pace enhancements in video and picture technology. This technique makes use of dynamic column-sparse deltas with out requiring further coaching, in response to Collectively.ai.
Dynamic Sparsity for Sooner Processing
Chipmunk employs a method the place it caches consideration weights and MLP activations from earlier steps, dynamically computing sparse deltas towards these cached weights. This technique permits Chipmunk to attain as much as 3.7 occasions quicker video technology on platforms like HunyuanVideo in comparison with conventional strategies. The strategy exhibits a 2.16x pace enchancment in particular configurations and as much as 1.6 occasions quicker picture technology on FLUX.1-dev.
Addressing Diffusion Transformer Challenges
Diffusion Transformers (DiTs) are broadly used for video technology, however their excessive time and price necessities have restricted their accessibility. Chipmunk addresses these challenges by specializing in two key insights: the slow-changing nature of mannequin activations and their inherent sparsity. By reformulating these activations to compute cross-step deltas, the strategy enhances their sparsity and effectivity.
{Hardware}-Conscious Optimization
Chipmunk’s design features a hardware-aware sparsity sample that optimizes for dense shared reminiscence tiles utilizing non-contiguous columns in world reminiscence. This strategy, mixed with quick kernels, allows important computational effectivity and pace enhancements. The tactic takes benefit of GPUs’ choice for computing giant blocks, aligning with native tile sizes for optimum efficiency.
Kernel Optimizations
To additional improve efficiency, Chipmunk incorporates a number of kernel optimizations. These embody quick sparsity identification by customized CUDA kernels, environment friendly cache writeback utilizing the CUDA driver API, and warp-specialized persistent kernels. These improvements contribute to a extra environment friendly execution, decreasing computation time and useful resource utilization.
Open Supply and Neighborhood Engagement
Collectively.ai has embraced the open-source group by releasing Chipmunk’s assets on GitHub, inviting builders to discover and leverage these developments. This initiative is a part of a broader effort to speed up mannequin efficiency throughout numerous architectures, comparable to FLUX-1.dev and DeepSeek R1.
For extra detailed insights and technical documentation, readers can entry the complete weblog publish on Collectively.ai.
Picture supply: Shutterstock


