Taruen
A custom language technology and data studio based in Istanbul.
Open for 20 Hours / Week Contracts & Technical Consulting
Alongside developing independent language technology products on taruen.com and educational tools on selimcan.org, I am actively taking on remote contract work (up to 20 hours/week). If you need an ASR/TTS model fine-tuned, PyTorch models optimized into high-throughput ONNX runtimes, hardened Azure ML endpoints, or large-scale Elasticsearch indices populated, let's talk.
Services & Technical Capabilities
Hands-on technical consulting and engineering practice focused on speech processing, inference acceleration, cloud MLOps, and symbolic computational linguistics:
Fine-tune Automatic Speech Recognition models (e.g. OpenAI Whisper, Conformer) and Text-to-Speech (TTS) models for specialized domains, distinct accents, or low-resource languages. Experience includes reaching 28.5% WER on held-out Laz test sets, building Coqui STT models, and curating speech synthesis datasets (Chuvash TTS).
Convert PyTorch and TensorFlow models to ONNX format. Apply INT8/FP16 quantization, graph pruning, and operator fusion with Hugging Face Optimum. Benchmark and optimize models for specific target production GPUs to multiply requests/second and reduce cloud inference cost.
Deploy production models to Microsoft Azure Machine Learning (Azure ML) endpoints ensuring high availability and zero downtime. Load-test endpoints under heavy simulated concurrency using Locust to right-size compute and validate latency SLAs. Audit and harden Docker images in Azure Container Registry (ACR) to remediate known CVE vulnerabilities.
Architect Elasticsearch mappings, tokenizers, and custom analyzers for complex similarity queries (phonetic matching, orthographic homoglyphs, and multilingual cross-search). Build high-throughput ingestion and data clearance pipelines capable of populating tens of millions of records (such as 20M+ USPTO public trademark archives).
Engineer deterministic, rule-based NLP tools when training data is sparse or unavailable. Build finite-state morphological transducers (HFST, Apertium), syntactic and dependency parsers using Constraint Grammar (VISL CG-3) under Universal Dependencies, and rule-based machine translation engines.
Studio Products & Edge Services
We build privacy-preserving language tools running directly on user hardware, alongside open-source machine learning models:
edge.taruen.com
A browser-based language technology playground built in ClojureScript (re-frame) and Transformers.js. All inference runs 100% on-device inside Web Workers; no audio or text is ever sent to a server.
- Laz Transliterator: A client-side utility for bidirectional script conversion of Laz text between the Georgian Mkhedruli script and Latin orthographies.
- Edge Speech-to-Text: In-browser on-device transcription running OpenAI's Whisper-small in a Web Worker.
whisper-small-laz (Hugging Face)
OpenAI's Whisper-small fine-tuned for Laz, an endangered South Caucasian language. Trained on Common Voice data from the Mozilla Data Collective, achieving 28.5% word error rate on the held-out test set for a language Whisper did not originally support. Read the end-to-end technical write-up on the Taruen blog →
Open Datasets
Taruen is dedicated to the digital preservation and technological advancement of regional and low-resource languages. All 12 of our machine-learning-ready datasets are published and verified on the Mozilla Data Collective:
- Kazakh Proverbs Text CorpusProverbs
- Uzbek Proverbs Text CorpusProverbs
- Kumyk Proverbs & SayingsProverbs
- Chechen Proverbs & SayingsProverbs
- Bulgarian Proverbs & SayingsProverbs
- Chuvash TTSSpeech / Audio
- World Factbook (JSON)JSON Data
About the Founder
Taruen is the independent practice of Ilnar Salimzianov, an NLP / ML engineer and computational linguist based in Istanbul. Holding an M.Sc. in Computational Linguistics from the University of Stuttgart and a Specialist's degree in German Philology (focus Linguistics) from Kazan State University, Ilnar has architected scalable language technology, high-throughput search infrastructure, and ML-ready datasets for organizations including the Mozilla Data Collective, US-based legal tech startups, and ProWritingAid.
Taruen operates on a digital craftsman ethos: building bespoke, robust data infrastructure where no technical problem is solved twice, eliminating bloated dependencies, and maximizing compute efficiency.
Read more about Ilnar's academic publications and open-source work on ifs.name → · GitHub · GitLab
Work With Taruen
Looking for an experienced NLP / ML engineer for a ~20 hours/week contract or technical advisory?
contact@taruen.com