Back to News
Microsoft•August 9, 2026
VibeVoice-ASR
Open Source
Audio
### TL;DR
VibeVoice-ASR is a unified speech-to-text model developed by Microsoft Research, designed to process up to 60 minutes of continuous audio in a single pass. It generates structured transcriptions that include speaker identification, timestamps, and content, with support for customized hotwords to enhance accuracy in domain-specific contexts.
Key Insights & Metrics
Pricing
Free
Cost structure
Version
N/A
Current release version
Hardware
N/A
Compute requirements
Category
Open Source
Licensing model
Region
United States
Primary region
Key Features
- 60-minute single-pass processing
- Customized hotwords for domain-specific accuracy
- Structured transcriptions with speaker identification and timestamps
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!