BembaSpeech: A Speech Recognition Corpus for the Bemba Language

Sikasote, Claytone; Anastasopoulos, Antonios

Computer Science > Computation and Language

arXiv:2102.04889 (cs)

[Submitted on 9 Feb 2021]

Title:BembaSpeech: A Speech Recognition Corpus for the Bemba Language

Authors:Claytone Sikasote, Antonios Anastasopoulos

Download PDF

Abstract: We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting over 24 hours of read speech in the Bemba language, a written but low-resourced language spoken by over 30% of the population in Zambia. To assess its usefulness for training and testing ASR systems for Bemba, we train an end-to-end Bemba ASR system by fine-tuning a pre-trained DeepSpeech English model on the training portion of the BembaSpeech corpus. Our best model achieves a word error rate (WER) of 54.78%. The results show that the corpus can be used for building ASR systems for Bemba. The corpus and models are publicly released at this https URL.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2102.04889 [cs.CL]
	(or arXiv:2102.04889v1 [cs.CL] for this version)

Submission history

From: Antonios Anastasopoulos [view email]
[v1] Tue, 9 Feb 2021 15:42:00 UTC (51 KB)

Full-text links:

Download:

(license)

Current browse context:

cs.CL

< prev | next >

new | recent | 2102

Change to browse by:

References & Citations

export bibtex citation

Bookmark

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer (What is the Explorer?)

Code

Code Associated with this Article

arXiv Links to Code (What is Links to Code?)

Recommenders and Search Tools

Connected Papers (What is Connected Papers?)

CORE Recommender (What is CORE?)

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs and how to get involved.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)