# Kaldi Projects at AI Tinkerers

> Canonical HTML: https://aitinkerers.org/technologies/kaldi
> Markdown URL: https://aitinkerers.org/technologies/kaldi.md
> Technology record last updated: 2026-03-04T16:47:18Z
> Generated: 2026-09-22T17:53:18Z

The industry-standard C++ toolkit for speech recognition, providing finite-state transducer based modeling and deep learning integration.

Kaldi is the definitive open-source framework for speech processing (ASR). Built on OpenFST, it offers a modular C++ codebase that supports linear algebra, acoustic modeling, and extensive feature extraction. Researchers use it to build robust systems like the LibriSpeech and Switchboard recipes, leveraging its flexible integration with CUDA for GPU-accelerated neural network training. It remains the primary engine for speech scientists requiring precise control over the decoding graph and lattice generation.

- Official technology site: https://kaldi-asr.org/
- Public AI Tinkerers demos and talks: 2
- Result page: 1 of 1

## Recent Public Talks and Demos

### [UzbekVoice](https://tashkent.aitinkerers.org/talks/rsvp_7WmwgOqtMgg)

UzbekVoiceAI is an innovative company specializing in speech processing and artificial intelligence solutions for the Uzbek-speaking audience. Our mission is to deliver high-quality speech recognition (STT) and text-to-speech (TTS) technologies for the Uzbek language, enhancing user interaction with digital devices and services. We develop solutions for speech-to-text conversion, text-to-speech synthesis, and large language models (LLMs) tailored to the unique features of the Uzbek language.

- Event context: AI Tinkerers - Tashkent Inaugural Meetup (October) — 2024-10-31 — Tashkent
- Public talk page: https://tashkent.aitinkerers.org/talks/rsvp_7WmwgOqtMgg

### [YT shorts finder](https://austin.aitinkerers.org/talks/rsvp_TpbSnDWuX0s)

A way to extract meaning from youtube shorts and search over your personalized database of them, using a mix of vision models, speech transcription models, and general purpose LLMs

- Event context: Community AI Demos - September Edition — 2024-09-12 — Austin
- Public talk page: https://austin.aitinkerers.org/talks/rsvp_TpbSnDWuX0s

## Related Technologies

- [Amazon Transcribe](https://aitinkerers.org/technologies/amazon-transcribe) ([Markdown](https://aitinkerers.org/technologies/amazon-transcribe.md)) — 2 public demos
- [BERT](https://aitinkerers.org/technologies/bert) ([Markdown](https://aitinkerers.org/technologies/bert.md)) — 179 public demos
- [CMU Sphinx](https://aitinkerers.org/technologies/cmu-sphinx) ([Markdown](https://aitinkerers.org/technologies/cmu-sphinx.md)) — 2 public demos
- [Google Cloud Speech-to-Text](https://aitinkerers.org/technologies/google-cloud-speech-to-text) ([Markdown](https://aitinkerers.org/technologies/google-cloud-speech-to-text.md)) — 2 public demos
- [GPT-3](https://aitinkerers.org/technologies/gpt-3) ([Markdown](https://aitinkerers.org/technologies/gpt-3.md)) — 191 public demos
- [GPT-4](https://aitinkerers.org/technologies/gpt-4) ([Markdown](https://aitinkerers.org/technologies/gpt-4.md)) — 529 public demos
- [IBM Watson Speech to Text](https://aitinkerers.org/technologies/ibm-watson-speech-to-text) ([Markdown](https://aitinkerers.org/technologies/ibm-watson-speech-to-text.md)) — 2 public demos
- [Amazon Polly](https://aitinkerers.org/technologies/amazon-polly) ([Markdown](https://aitinkerers.org/technologies/amazon-polly.md)) — 1 public demo
- [Azure Speech to Text](https://aitinkerers.org/technologies/azure-speech-to-text) ([Markdown](https://aitinkerers.org/technologies/azure-speech-to-text.md)) — 1 public demo
- [BLOOM](https://aitinkerers.org/technologies/bloom) ([Markdown](https://aitinkerers.org/technologies/bloom.md)) — 115 public demos
- [Database](https://aitinkerers.org/technologies/database) ([Markdown](https://aitinkerers.org/technologies/database.md)) — 8 public demos
- [DeepSpeech](https://aitinkerers.org/technologies/deepspeech) ([Markdown](https://aitinkerers.org/technologies/deepspeech.md)) — 1 public demo
- [eSpeak](https://aitinkerers.org/technologies/espeak) ([Markdown](https://aitinkerers.org/technologies/espeak.md)) — 1 public demo
- [Festival](https://aitinkerers.org/technologies/festival) ([Markdown](https://aitinkerers.org/technologies/festival.md)) — 1 public demo
- [Google Cloud Text-to-Speech](https://aitinkerers.org/technologies/google-cloud-text-to-speech) ([Markdown](https://aitinkerers.org/technologies/google-cloud-text-to-speech.md)) — 1 public demo
- [IBM Watson Text to Speech](https://aitinkerers.org/technologies/ibm-watson-text-to-speech) ([Markdown](https://aitinkerers.org/technologies/ibm-watson-text-to-speech.md)) — 1 public demo
- [Keras](https://aitinkerers.org/technologies/keras) ([Markdown](https://aitinkerers.org/technologies/keras.md)) — 74 public demos
- [Llama-2](https://aitinkerers.org/technologies/llama-2) ([Markdown](https://aitinkerers.org/technologies/llama-2.md)) — 227 public demos
