Back to News
Infini-AI-Lab (CMU)•May 20, 2026
MonarchRT
Paid
Video Generation
Efficient Attention
Diffusion Models
Real-Time AI
CMU
Open Source
Research
Triton Kernels
### TL;DR
MonarchRT is a research method from Carnegie Mellon University that enables true real-time video generation by parameterizing attention maps in video diffusion transformers as sparse Monarch matrices. It achieves up to 95% effective attention sparsity with no quality loss, unlocking 16 FPS real-time video generation on a single RTX 5090.
Key Insights & Metrics
Pricing
Free (Open Source, Apache-2.0)
Cost structure
Version
latest
Current release version
Hardware
NVIDIA GPU (RTX 5090 for real-time 16 FPS); CUDA with Triton support
Compute requirements
Category
Paid
Licensing model
Region
United States
Primary region
Key Features
- Monarch matrix attention factorization — sparsely parameterizes 3D video attention maps using structured Monarch matrices that respect spatiotemporal block alignment, achieving 95% effective sparsity with no quality loss
- 1.4–11.8× speedup over FlashAttention (-2, -3, -4) kernels via custom Triton kernel implementation — first method to achieve true real-time 16 FPS video generation with Self-Forcing on a single RTX 5090
- Tiled Monarch parameterization solves monotonic compute-accuracy trade-off and 1-iteration + finetuning approach reduces iterative refinement overhead for real-time few-step diffusion models
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!