Back to News
Infini-AI-Lab (CMU)May 20, 2026

MonarchRT

Paid
Video Generation
Efficient Attention
Diffusion Models
Real-Time AI
CMU
Open Source
Research
Triton Kernels

Explore MonarchRT

Visit the official website to learn more and get started

### TL;DR

MonarchRT is a research method from Carnegie Mellon University that enables true real-time video generation by parameterizing attention maps in video diffusion transformers as sparse Monarch matrices. It achieves up to 95% effective attention sparsity with no quality loss, unlocking 16 FPS real-time video generation on a single RTX 5090.

Key Insights & Metrics

Pricing
Free (Open Source, Apache-2.0)
Cost structure
Version
latest
Current release version
Hardware
NVIDIA GPU (RTX 5090 for real-time 16 FPS); CUDA with Triton support
Compute requirements
Category
Paid
Licensing model
Region
United States
Primary region

Key Features

  • Monarch matrix attention factorization — sparsely parameterizes 3D video attention maps using structured Monarch matrices that respect spatiotemporal block alignment, achieving 95% effective sparsity with no quality loss
  • 1.4–11.8× speedup over FlashAttention (-2, -3, -4) kernels via custom Triton kernel implementation — first method to achieve true real-time 16 FPS video generation with Self-Forcing on a single RTX 5090
  • Tiled Monarch parameterization solves monotonic compute-accuracy trade-off and 1-iteration + finetuning approach reduces iterative refinement overhead for real-time few-step diffusion models

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode