Fine-tuning Whisper for Laz: An End-to-End Journey
Taruen is a language technology studio that puts equal emphasis on supporting all languages, regardless of speaker numbers. Turkey is home to many languages besides Turkish, and one of them is Laz — a Kartvelian language spoken primarily along the southeastern Black Sea coast, listed by UNESCO as definitely endangered.
While based in Istanbul, we had the privilege of meeting people from the Laz Institute, an organization founded in 2013 dedicated to preserving, developing, and revitalizing the Laz language. They do extensive work: producing textbooks, organizing language courses (Laz is now an elective in some Turkish schools and universities), publishing books through the Lazika Yayın Kollektifi, and crucially for our purposes, contributing voices to Mozilla Common Voice.
The Institute also operates the LazuriTV YouTube channel, which hosts roughly 105 hours of spoken Laz content — about 4× the data available on Common Voice (28 hours). However, most of these videos lack transcriptions or subtitles.
Our previous involvement with speech recognition was a 2021 paper on low-resource ASR for Turkic languages. Out of three motivations — sharpening our own skills, learning the latest developments in the speech-to-text field, and wanting to help the Laz Institute eventually transcribe their YouTube archive to feed back into Common Voice — we decided to fine-tune a Whisper-small ASR model on the Common Voice Laz data.
This post documents how we did it. The resulting model is available on the Hugging Face Hub at Taruen/whisper-small-laz. (We've also built a browser-based Laz transliterator that converts Georgian Mkhedruli script into customizable Latin orthographies, with options for exactly the ejective and affricate consonants the model has to get right below.)
