Back to News
GoogleJanuary 1, 2026

WaxalNLP

Open Source
Audio

Explore WaxalNLP

Visit the official website to learn more and get started

### TL;DR

The WaxalNLP dataset is a comprehensive collection of audio data for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) tasks in 14 African languages. It aims to enhance the accuracy and fluency of speech and language technologies for underserved African languages and serves as a resource for digital preservation.

Key Insights & Metrics

Pricing
Free
Cost structure
Version
1.0.0
Current release version
Hardware
CPU only
Compute requirements
Category
Open Source
Licensing model
Region
United States
Primary region

Key Features

  • Contains approximately 1,250 hours of transcribed natural speech for ASR
  • Includes about 240 hours of scripted natural speech for TTS
  • Covers 14 African languages spoken by over 100 million people across 40 Sub-Saharan countries
  • Acquired through partnerships with Makerere University, The University of Ghana, Digital Umuganda, and Media Trust
  • Funded by Google and the Gates Foundation to be openly accessible

Discussion

1
Upvotes
0
Downvotes
1 review

Sign in to leave a review

Reviews

Upvote

asif@marktechpost.com • Jan 26, 2026

Quick vote

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode