Skip to main content

Common Voice Kazakh

· 5 min read

Бұл мақаланың қазақша нұсқасы мұнда (Kazakh version here).

Update (2026)

Good news since this was written: Kazakh has launched on Common Voice and is open for contributions. Following the effort the 2019 post below calls for — translating the interface and collecting sentences, with the help of native Kazakh speakers — the language went live. As of the 26.0 release, contributors have accumulated a few hours of validated audio — a modest but real start. You can add your voice now at commonvoice.mozilla.org/kk.

The 2019 post below has aged into a piece of that history.

Wouldn’t it be great, if Google Assistant or Yandex Alisa spoke Kazakh?

One necessary component of such speech-enabled digital assistants is a so-called automatic speech recognition (ASR) or speech-to-text (STT) system.

Large amounts of audio data (thousands of hours, see below), from many different people, along with transcriptions, are needed to train a good speech-to-text system using machine learning methods. So far, due to lack of appropriately licensed, freely available audio data in them, building a high-accuracy free/libre/open-source software (FLOSS) speech recognition system is out of reach for most languages.

Update (2026)

Thanks to pre-trained models like Whisper and fine-tuning, the amount of training data needed to build an accurate ASR system for a language the original Whisper does not support is much less nowadays — in the order of tens of hours, not thousands. See subsequent posts on this blog for examples of that.

Fortunately, there is a way to change this for the better. In 2018, Mozilla, the company behind the Firefox web-browser and many other programs, launched the Common Voice project.

What is Mozilla Common Voice?

Here is a quote from the project's About page:

Common Voice is a publicly available voice dataset, powered by the voices of volunteer contributors around the world. People who want to build voice applications can use the dataset to train machine learning models.

How you can help

We at Taruen want to launch Common Voice in Kazakh, and took the first steps towards that goal. But for that to happen, we need help from native speakers of Kazakh. If you are one, here is how you can help:

  1. Review whether the Common Voice website has been correctly translated into Kazakh on Pontoon, Mozilla’s localization tool. About one-third of the translations were authored by us — non-native Kazakh speakers — and thus might be incorrect.

  2. Review Kazakh sentences we’ve submitted to the Common Voice Sentence Collector tool. At least 5000 reviewed sentences are needed to “launch” a language on Common Voice.

These sentences were taken from Volumes 65 and 68 of the mighty 100-volume set with works of Kazakh folklore, called “Бабалар сөзі” and published by M. Auezov Institute of Literature and Art. By Article 8 of Kazakhstani Copyright Law, works of folklore are exempt from copyright and are thus in the public domain and suitable for submitting to Common Voice. Moreover, proverbs from volumes 65 and 68 match other criteria of the Common Voice project as well — they don’t contain digits, foreign letters and other symbols not allowed in the dataset, they are mostly short, arguably pedagogical/entertaining and thus fun to read.

Besides usual checks for correct spelling and grammaticality required by Common Voice, when reviewing sentences, we also ask you to make sure that none of the sentences are remotely offensive to men, women, parents, children, religious people, non-religious people, Southern Kazakhstanis, Northern Kazakhstanis, Western Kazakhstanis, Eastern Kazakhstanis, Kazakhs of China, Kazakhs of … You get the idea. A glimpse over the sentences did not suggest that they would contain anything like that, but again, that’s something for native speakers to judge.

We hope that together we can assemble enough data so that anyone who wishes to do so can train a speech-to-text system for Kazakh. On the FAQ page of the project, Mozillians mention 10000 hours as an approximate number of validated hours needed to train a production speech-to-text system from scratch, so that’s something to strive for. It does sound like a lot, but when divided among even a conservative number of Kazakh speakers, it requires each person to record sentences for about 4 seconds. That is a much less scary number.

(As the Update above notes, fine-tuning a pre-trained model like Whisper needs far less data — but more data never hurts, and a large, openly-licensed Kazakh voice dataset is valuable regardless of how models are trained.)