LongSpeech v1.0
LongSpeech is a large-scale long-form audio understanding dataset with 100,000+ audio segments (~10 min each), designed to benchmark and train Audio LLMs on long-form speech. Presented at ICASSP 2026, it covers 8 tasks: ASR, translation, summarization, speaker counting, QA, and emotion analysis.
100,000+ long-form audio segments (~10 minutes each) for benchmarking Audio LLMs
Covers 8 tasks: ASR, translation, summarization, speaker counting, QA, emotion analysis
Publicly released on Hugging Face with accompanying ICASSP 2026 paper