# Speech Processing Projects at AI Tinkerers

> Canonical HTML: https://aitinkerers.org/technologies/speech-processing
> Markdown URL: https://aitinkerers.org/technologies/speech-processing.md
> Technology record last updated: 2026-03-04T16:47:03Z
> Generated: 2026-09-21T15:37:21Z

Speech Processing is the digital signal process that enables machines to accurately convert spoken audio to text (ASR) and synthesize human-like voice from text (TTS).

This technology is a critical fusion of digital signal processing and machine learning, allowing systems to acquire, manipulate, and interpret human speech signals. Core tasks include Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). ASR, powered by deep neural networks, drives high-impact applications: think virtual assistants like Siri and Alexa, or the call routing system AT&amp;T deployed in 1992. TTS provides the natural voice output. The goal is seamless human-computer interaction, enabling hands-free dictation, which is demonstrably faster—up to 3x the speed of typing—and fundamentally reshaping customer experience and enterprise workflow.

- Official technology site: https://en.wikipedia.org/wiki/Speech_processing
- Public AI Tinkerers demos and talks: 1
- Result page: 1 of 1

## Recent Public Talks and Demos

### [UzbekVoice](https://tashkent.aitinkerers.org/talks/rsvp_7WmwgOqtMgg)

UzbekVoiceAI is an innovative company specializing in speech processing and artificial intelligence solutions for the Uzbek-speaking audience. Our mission is to deliver high-quality speech recognition (STT) and text-to-speech (TTS) technologies for the Uzbek language, enhancing user interaction with digital devices and services. We develop solutions for speech-to-text conversion, text-to-speech synthesis, and large language models (LLMs) tailored to the unique features of the Uzbek language.

- Event context: AI Tinkerers - Tashkent Inaugural Meetup (October) — 2024-10-31 — Tashkent
- Public talk page: https://tashkent.aitinkerers.org/talks/rsvp_7WmwgOqtMgg

## Related Technologies

- [Amazon Polly](https://aitinkerers.org/technologies/amazon-polly) ([Markdown](https://aitinkerers.org/technologies/amazon-polly.md)) — 1 public demo
- [Amazon Transcribe](https://aitinkerers.org/technologies/amazon-transcribe) ([Markdown](https://aitinkerers.org/technologies/amazon-transcribe.md)) — 2 public demos
- [Azure Speech to Text](https://aitinkerers.org/technologies/azure-speech-to-text) ([Markdown](https://aitinkerers.org/technologies/azure-speech-to-text.md)) — 1 public demo
- [BERT](https://aitinkerers.org/technologies/bert) ([Markdown](https://aitinkerers.org/technologies/bert.md)) — 179 public demos
- [CMU Sphinx](https://aitinkerers.org/technologies/cmu-sphinx) ([Markdown](https://aitinkerers.org/technologies/cmu-sphinx.md)) — 2 public demos
- [eSpeak](https://aitinkerers.org/technologies/espeak) ([Markdown](https://aitinkerers.org/technologies/espeak.md)) — 1 public demo
- [Festival](https://aitinkerers.org/technologies/festival) ([Markdown](https://aitinkerers.org/technologies/festival.md)) — 1 public demo
- [Google Cloud Speech-to-Text](https://aitinkerers.org/technologies/google-cloud-speech-to-text) ([Markdown](https://aitinkerers.org/technologies/google-cloud-speech-to-text.md)) — 2 public demos
- [Google Cloud Text-to-Speech](https://aitinkerers.org/technologies/google-cloud-text-to-speech) ([Markdown](https://aitinkerers.org/technologies/google-cloud-text-to-speech.md)) — 1 public demo
- [GPT-3](https://aitinkerers.org/technologies/gpt-3) ([Markdown](https://aitinkerers.org/technologies/gpt-3.md)) — 191 public demos
- [GPT-4](https://aitinkerers.org/technologies/gpt-4) ([Markdown](https://aitinkerers.org/technologies/gpt-4.md)) — 529 public demos
- [IBM Watson Speech to Text](https://aitinkerers.org/technologies/ibm-watson-speech-to-text) ([Markdown](https://aitinkerers.org/technologies/ibm-watson-speech-to-text.md)) — 2 public demos
- [IBM Watson Text to Speech](https://aitinkerers.org/technologies/ibm-watson-text-to-speech) ([Markdown](https://aitinkerers.org/technologies/ibm-watson-text-to-speech.md)) — 1 public demo
- [Kaldi](https://aitinkerers.org/technologies/kaldi) ([Markdown](https://aitinkerers.org/technologies/kaldi.md)) — 2 public demos
- [Keras](https://aitinkerers.org/technologies/keras) ([Markdown](https://aitinkerers.org/technologies/keras.md)) — 74 public demos
- [Microsoft Azure Text-to-Speech](https://aitinkerers.org/technologies/microsoft-azure-text-to-speech) ([Markdown](https://aitinkerers.org/technologies/microsoft-azure-text-to-speech.md)) — 1 public demo
- [Mozilla DeepSpeech](https://aitinkerers.org/technologies/mozilla-deepspeech) ([Markdown](https://aitinkerers.org/technologies/mozilla-deepspeech.md)) — 1 public demo
- [ONNX](https://aitinkerers.org/technologies/onnx) ([Markdown](https://aitinkerers.org/technologies/onnx.md)) — 83 public demos
