# WanIntellif **Repository Path**: Intellifusion_com/WanIntellif ## Basic Information - **Project Name**: WanIntellif - **Description**: wan2.1推理加速 - **Primary Language**: Python - **License**: Apache-2.0 - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 1 - **Forks**: 0 - **Created**: 2026-07-06 - **Last Updated**: 2026-09-29 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # WanIntellif This repository provides the official implementation of **WanIntellif**, a video generation acceleration framework that can speed up end-to-end diffusion generation by 100x to 270x on a single RTX 5090, while maintaining video quality. WanIntellif integrates and optimizes SageAttention, SLA, rCM, quantization, kernel fusion, type conversion elimination, and efficient VAE decoding to achieve end-to-end inference acceleration. WanIntellif is based on TurboDiffusion and extends the optimized Wan-2.1 inference path for RTX 5090 deployment. ## Contents - [Models](#models) - [Installation](#installation) - [Inference](#inference) - [Evaluation](#evaluation) - [Wan-2.1-T2V-14B-720P](#wan-21-t2v-14b-720p) - [Acknowledgements](#acknowledgements) **Note**: The current model is trained on **long English prompts**. If you use other types of prompts, please augment them to get better performance. The DiT checkpoint is available now. The distilled VAED decoder checkpoint is upcoming; before it is released, the public Wan-2.1 VAE checkpoint can be used to reproduce the optimized inference latency, while matching the final distilled-VAED output quality requires the upcoming checkpoint.
Acceleration decomposition
Original, DiT Time: 4767s
Original 14B sample 11
WanIntellif, DiT Time: 17.64s
WanIntellif 14B sample 11
An example of a 5-second video generated by Wan-2.1-T2V-14B-720P on a single RTX 5090.
## Models | Model Name | Checkpoint Link | Resolution | Status | | :-----------------------------------: | :----------------------------------------------------------: | :-------------: | :----: | | `TurboWan2.1-T2V-14B-720P` | [Huggingface Model](https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-14B-720P) | 720p | Available | | Distilled VAED decoder checkpoint | Upcoming | 720p | Upcoming | This branch publishes only the Wan-2.1-T2V-14B-720P optimized inference path. The available DiT checkpoint is the TurboDiffusion FP8 weight for Wan-2.1-T2V-14B-720P; WanIntellif uses it with the optimized inference stack in this repository. ## Installation The optimized Wan-2.1 14B 720p FP8 + VAED path has been validated on the following environment. | Component | Validated version | | :--- | :--- | | GPU | NVIDIA GeForce RTX 5090 (Blackwell, SM120) | | Python | 3.11.13 | | PyTorch | 2.8.0+cu129 | | CUDA Toolkit / NVCC | 12.9 / V12.9.86 | | Triton | 3.4.0 | | flash-attn | 2.8.3 | | GCC / G++ | 11.4.0 | | Ninja | 1.11.1 | Other recent CUDA/PyTorch combinations may work, but the latency numbers below should be reproduced with the validated stack first. Compile from source: ```bash cd turbodiffusion # Initialize the NVIDIA CUTLASS submodule required by the FP8/CUTLASS CUDA extension. git submodule update --init --recursive python3 -m pip install -U pip setuptools wheel packaging ninja python3 -m pip install torch==2.8.0 torchvision --index-url https://download.pytorch.org/whl/cu129 python3 -m pip install flash-attn==2.8.3 --no-build-isolation ``` SageSLA is required to reproduce the optimized RTX 5090 inference. Install [SpargeAttn](https://github.com/thu-ml/SpargeAttn) before building this repository: ```bash python3 -m pip install git+https://github.com/thu-ml/SpargeAttn.git@bfd980b --no-build-isolation ``` ### RTX 5090 / Blackwell Build Notes The FP8/CUTLASS backend is compiled as a CUDA extension and includes Blackwell code generation targets such as `sm_120a`. Make sure that: - `nvcc --version` reports CUDA Toolkit 12.9 or a toolkit that supports Blackwell code generation. - `CUDA_HOME` points to the same CUDA Toolkit used by PyTorch. - The PyTorch CUDA libraries are visible at runtime. - `ninja` is installed before running the editable install. Build and verify the optimized backends: ```bash export CUDA_HOME=${CUDA_HOME:-/usr/local/cuda} export TORCH_CUDA_ARCH_LIST=${TORCH_CUDA_ARCH_LIST:-12.0} export LD_LIBRARY_PATH="$(python3 - <<'PY' import pathlib, torch print(pathlib.Path(torch.__file__).parent / "lib") PY ):${LD_LIBRARY_PATH:-}" python3 -m pip install -e . --no-build-isolation python3 - <<'PY' import turbo_diffusion_ops import spas_sage_attn._qattn import spas_sage_attn._fused print("optimized backends OK") PY ``` ## Inference This branch publishes the optimized Wan-2.1-T2V-14B-720P FP8 checkpoint. Use `--quant_linear_fp8` when running inference. The public Wan-2.1 VAE checkpoint can run the optimized `vaed_op_repl` backend and reproduce the optimized inference latency. The distilled VAED checkpoint is upcoming and is required to reproduce the same final VAED output quality. 1. Download the public Wan-2.1 VAE and umT5 text encoder checkpoints: ```bash mkdir checkpoints cd checkpoints wget https://huggingface.co/Wan-AI/Wan2.1-T2V-14B/resolve/main/Wan2.1_VAE.pth wget https://huggingface.co/Wan-AI/Wan2.1-T2V-14B/resolve/main/models_t5_umt5-xxl-enc-bf16.pth ``` 2. Download the optimized Wan-2.1-T2V-14B-720P FP8 checkpoint. This is the TurboDiffusion weight used by WanIntellif: ```bash wget https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-14B-720P/resolve/main/TurboWan2.1-T2V-14B-720P-FP8.pth ``` 3. Run the Wan-2.1-T2V-14B-720P inference script: ```bash export DIT_PATH=/path/to/TurboWan2.1-T2V-14B-720P-FP8.pth export VAE_PATH=/path/to/Wan2.1_VAE.pth export TEXT_ENCODER_PATH=/path/to/models_t5_umt5-xxl-enc-bf16.pth mkdir -p output/e2e_logs for run in 1 2 3 4 5; do SAVE_PATH="output/generated_wan21_14b_720p_fp8_vaed_run${run}.mp4" \ bash scripts/inference_wan2.1_t2v.sh \ 2>&1 | tee "output/e2e_logs/wan21_14b_720p_fp8_vaed_run${run}.log" done ``` The log should show `vae_backend=vaed_op_repl`. Read the `Timing summary` line for `dit_sampling_seconds`, `vae_decode_seconds`, and `e2e_dit_vae_seconds`. On the validated RTX 5090 environment, typical steady-state values are about 17.6s DiT sampling and 4.6s VAE decoding with the public Wan-2.1 VAE checkpoint. ## Evaluation We evaluate video generation on **a single RTX 5090 GPU**. The E2E Time refers to DiT sampling latency plus VAE decoding latency, excluding text encoding, Python process startup, and video file writing. ### Wan-2.1-T2V-14B-720P
Original, E2E Time: 4774.2s
Original 14B 720p sample 0
FastVideo, E2E Time: 79.8s
FastVideo 14B 720p sample 0
TurboDiffusion, E2E Time: 31.2s
TurboDiffusion 14B 720p sample 0
WanIntellif, E2E Time: 22.4s
WanIntellif 14B 720p sample 0
Original, E2E Time: 4774.2s
Original 14B 720p sample 3
FastVideo, E2E Time: 79.8s
FastVideo 14B 720p sample 3
TurboDiffusion, E2E Time: 31.2s
TurboDiffusion 14B 720p sample 3
WanIntellif, E2E Time: 22.4s
WanIntellif 14B 720p sample 3
Original, E2E Time: 4774.2s
Original 14B 720p sample 6
FastVideo, E2E Time: 79.8s
FastVideo 14B 720p sample 6
TurboDiffusion, E2E Time: 31.2s
TurboDiffusion 14B 720p sample 6
WanIntellif, E2E Time: 22.4s
WanIntellif 14B 720p sample 6
## Acknowledgements WanIntellif is based on [TurboDiffusion](https://github.com/thu-ml/TurboDiffusion) and uses its weights. If you use this repository, please also cite TurboDiffusion: ```bibtex @article{zhang2025turbodiffusion, title={TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times}, author={Zhang, Jintao and Zheng, Kaiwen and Jiang, Kai and Wang, Haoxu and Stoica, Ion and Gonzalez, Joseph E and Chen, Jianfei and Zhu, Jun}, journal={arXiv preprint arXiv:2512.16093}, year={2025} } ```