I am a first-year PhD student in the CS Dept. at Tsinghua University, focusing on efficient training and inference of large models.
- 🏠 My Homepage.
I am a first-year PhD student in the CS Dept. at Tsinghua University, focusing on efficient training and inference of large models.
TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Vidu S1: A Real-Time Interactive Video Generation Model
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
A Survey of Efficient Attention Methods: Hardware-efficient, Sparse, Compact, and Linear Attention
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention