# WanIntellif
**Repository Path**: Intellifusion_com/WanIntellif
## Basic Information
- **Project Name**: WanIntellif
- **Description**: wan2.1推理加速
- **Primary Language**: Python
- **License**: Apache-2.0
- **Default Branch**: main
- **Homepage**: None
- **GVP Project**: No
## Statistics
- **Stars**: 1
- **Forks**: 0
- **Created**: 2026-07-06
- **Last Updated**: 2026-09-29
## Categories & Tags
**Categories**: Uncategorized
**Tags**: None
## README
# WanIntellif
This repository provides the official implementation of **WanIntellif**, a video generation acceleration framework that can speed up end-to-end diffusion generation by 100x to 270x on a single RTX 5090, while maintaining video quality.
WanIntellif integrates and optimizes SageAttention, SLA, rCM, quantization, kernel fusion, type conversion elimination, and efficient VAE decoding to achieve end-to-end inference acceleration.
WanIntellif is based on TurboDiffusion and extends the optimized Wan-2.1 inference path for RTX 5090 deployment.
## Contents
- [Models](#models)
- [Installation](#installation)
- [Inference](#inference)
- [Evaluation](#evaluation)
- [Wan-2.1-T2V-14B-720P](#wan-21-t2v-14b-720p)
- [Acknowledgements](#acknowledgements)
**Note**: The current model is trained on **long English prompts**. If you use other types of prompts, please augment them to get better performance.
The DiT checkpoint is available now. The distilled VAED decoder checkpoint is upcoming; before it is released, the public Wan-2.1 VAE checkpoint can be used to reproduce the optimized inference latency, while matching the final distilled-VAED output quality requires the upcoming checkpoint.
|
Original, DiT Time: 4767s
|
WanIntellif, DiT Time: 17.64s
|
An example of a
5-second video generated by Wan-2.1-T2V-14B-720P on a single
RTX 5090.
## Models
| Model Name | Checkpoint Link | Resolution | Status |
| :-----------------------------------: | :----------------------------------------------------------: | :-------------: | :----: |
| `TurboWan2.1-T2V-14B-720P` | [Huggingface Model](https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-14B-720P) | 720p | Available |
| Distilled VAED decoder checkpoint | Upcoming | 720p | Upcoming |
This branch publishes only the Wan-2.1-T2V-14B-720P optimized inference path.
The available DiT checkpoint is the TurboDiffusion FP8 weight for Wan-2.1-T2V-14B-720P; WanIntellif uses it with the optimized inference stack in this repository.
## Installation
The optimized Wan-2.1 14B 720p FP8 + VAED path has been validated on the following environment.
| Component | Validated version |
| :--- | :--- |
| GPU | NVIDIA GeForce RTX 5090 (Blackwell, SM120) |
| Python | 3.11.13 |
| PyTorch | 2.8.0+cu129 |
| CUDA Toolkit / NVCC | 12.9 / V12.9.86 |
| Triton | 3.4.0 |
| flash-attn | 2.8.3 |
| GCC / G++ | 11.4.0 |
| Ninja | 1.11.1 |
Other recent CUDA/PyTorch combinations may work, but the latency numbers below should be reproduced with the validated stack first.
Compile from source:
```bash
cd turbodiffusion
# Initialize the NVIDIA CUTLASS submodule required by the FP8/CUTLASS CUDA extension.
git submodule update --init --recursive
python3 -m pip install -U pip setuptools wheel packaging ninja
python3 -m pip install torch==2.8.0 torchvision --index-url https://download.pytorch.org/whl/cu129
python3 -m pip install flash-attn==2.8.3 --no-build-isolation
```
SageSLA is required to reproduce the optimized RTX 5090 inference. Install [SpargeAttn](https://github.com/thu-ml/SpargeAttn) before building this repository:
```bash
python3 -m pip install git+https://github.com/thu-ml/SpargeAttn.git@bfd980b --no-build-isolation
```
### RTX 5090 / Blackwell Build Notes
The FP8/CUTLASS backend is compiled as a CUDA extension and includes Blackwell code generation targets such as `sm_120a`. Make sure that:
- `nvcc --version` reports CUDA Toolkit 12.9 or a toolkit that supports Blackwell code generation.
- `CUDA_HOME` points to the same CUDA Toolkit used by PyTorch.
- The PyTorch CUDA libraries are visible at runtime.
- `ninja` is installed before running the editable install.
Build and verify the optimized backends:
```bash
export CUDA_HOME=${CUDA_HOME:-/usr/local/cuda}
export TORCH_CUDA_ARCH_LIST=${TORCH_CUDA_ARCH_LIST:-12.0}
export LD_LIBRARY_PATH="$(python3 - <<'PY'
import pathlib, torch
print(pathlib.Path(torch.__file__).parent / "lib")
PY
):${LD_LIBRARY_PATH:-}"
python3 -m pip install -e . --no-build-isolation
python3 - <<'PY'
import turbo_diffusion_ops
import spas_sage_attn._qattn
import spas_sage_attn._fused
print("optimized backends OK")
PY
```
## Inference
This branch publishes the optimized Wan-2.1-T2V-14B-720P FP8 checkpoint. Use `--quant_linear_fp8` when running inference.
The public Wan-2.1 VAE checkpoint can run the optimized `vaed_op_repl` backend and reproduce the optimized inference latency. The distilled VAED checkpoint is upcoming and is required to reproduce the same final VAED output quality.
1. Download the public Wan-2.1 VAE and umT5 text encoder checkpoints:
```bash
mkdir checkpoints
cd checkpoints
wget https://huggingface.co/Wan-AI/Wan2.1-T2V-14B/resolve/main/Wan2.1_VAE.pth
wget https://huggingface.co/Wan-AI/Wan2.1-T2V-14B/resolve/main/models_t5_umt5-xxl-enc-bf16.pth
```
2. Download the optimized Wan-2.1-T2V-14B-720P FP8 checkpoint. This is the TurboDiffusion weight used by WanIntellif:
```bash
wget https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-14B-720P/resolve/main/TurboWan2.1-T2V-14B-720P-FP8.pth
```
3. Run the Wan-2.1-T2V-14B-720P inference script:
```bash
export DIT_PATH=/path/to/TurboWan2.1-T2V-14B-720P-FP8.pth
export VAE_PATH=/path/to/Wan2.1_VAE.pth
export TEXT_ENCODER_PATH=/path/to/models_t5_umt5-xxl-enc-bf16.pth
mkdir -p output/e2e_logs
for run in 1 2 3 4 5; do
SAVE_PATH="output/generated_wan21_14b_720p_fp8_vaed_run${run}.mp4" \
bash scripts/inference_wan2.1_t2v.sh \
2>&1 | tee "output/e2e_logs/wan21_14b_720p_fp8_vaed_run${run}.log"
done
```
The log should show `vae_backend=vaed_op_repl`. Read the `Timing summary` line for `dit_sampling_seconds`, `vae_decode_seconds`, and `e2e_dit_vae_seconds`. On the validated RTX 5090 environment, typical steady-state values are about 17.6s DiT sampling and 4.6s VAE decoding with the public Wan-2.1 VAE checkpoint.
## Evaluation
We evaluate video generation on **a single RTX 5090 GPU**. The E2E Time refers to DiT sampling latency plus VAE decoding latency, excluding text encoding, Python process startup, and video file writing.
### Wan-2.1-T2V-14B-720P
|
Original, E2E Time: 4774.2s
|
FastVideo, E2E Time: 79.8s
|
TurboDiffusion, E2E Time: 31.2s
|
WanIntellif, E2E Time: 22.4s
|
|
Original, E2E Time: 4774.2s
|
FastVideo, E2E Time: 79.8s
|
TurboDiffusion, E2E Time: 31.2s
|
WanIntellif, E2E Time: 22.4s
|
|
Original, E2E Time: 4774.2s
|
FastVideo, E2E Time: 79.8s
|
TurboDiffusion, E2E Time: 31.2s
|
WanIntellif, E2E Time: 22.4s
|
## Acknowledgements
WanIntellif is based on [TurboDiffusion](https://github.com/thu-ml/TurboDiffusion) and uses its weights. If you use this repository, please also cite TurboDiffusion:
```bibtex
@article{zhang2025turbodiffusion,
title={TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times},
author={Zhang, Jintao and Zheng, Kaiwen and Jiang, Kai and Wang, Haoxu and Stoica, Ion and Gonzalez, Joseph E and Chen, Jianfei and Zhu, Jun},
journal={arXiv preprint arXiv:2512.16093},
year={2025}
}
```