# OpenAI Whisper Projects at AI Tinkerers

> Canonical HTML: https://aitinkerers.org/technologies/openai-whisper
> Markdown URL: https://aitinkerers.org/technologies/openai-whisper.md
> Technology record last updated: 2026-02-22T17:59:33Z
> Generated: 2026-09-23T06:35:22Z

Whisper is OpenAI's robust, open-source Automatic Speech Recognition (ASR) system, trained on 680,000 hours of diverse audio.

This is Whisper: a high-performance, general-purpose ASR model from OpenAI. It was trained on a massive 680,000 hours of multilingual, multitask data, resulting in exceptional robustness against accents, background noise, and technical language. The model is a Transformer sequence-to-sequence architecture, engineered for multiple tasks: multilingual transcription, speech-to-English translation, and language identification. Developers leverage the open-source code and various model sizes (tiny, base, small, medium, large) to balance transcription speed with near human-level accuracy for diverse applications.

- Official technology site: https://github.com/openai/whisper
- Public AI Tinkerers demos and talks: 10
- Result page: 1 of 1

## Recent Public Talks and Demos

### [I gave my reachy mini a brain](https://nyc.aitinkerers.org/talks/rsvp_DmKV1qTMbSw)

Reachy Mini robot with an OpenClaw brain

- Event context: 🦞Demo Night: OpenClaw ft Convex — 2026-02-17 — New York City
- Public talk page: https://nyc.aitinkerers.org/talks/rsvp_DmKV1qTMbSw

### [AI powered agent to improve public speaking sklils](https://tiruchirappalli.aitinkerers.org/talks/rsvp_HxfxvA64rE8)

Demo will be divided into the following 1. I will showcase how the tool works. 2. I will detail down the technologies used 3. Ask on the improvements/best models to use/optimise the output

- Event context: AI Tinkerers Trichy: January Meetup &amp; Live Demos — 2026-01-31 — Tiruchirappalli
- Public talk page: https://tiruchirappalli.aitinkerers.org/talks/rsvp_HxfxvA64rE8

### [AI and music - entertaining people's ears with AI](https://toronto.aitinkerers.org/talks/rsvp_V0tbwYU7b48)

I will be presenting my work on AI music. First, I will demonstrate that low parameter count and efficiency could be achieved for AI music generation. This comes from my recent publication on music diffusion that is conditioned on the vocal. Second, I will showcase my project Xing Xing. It is a karaoke singing app that separates tracks and has live transcription. It uses vaiious off-the-market AI tools and models. It demonstrates scalability and different ways that musical AI could be applied

- Event context: AI Tinkerers Toronto - December Meetup sponsored by Auth0 and TribalScale! — 2025-12-03 — Toronto
- Public talk page: https://toronto.aitinkerers.org/talks/rsvp_V0tbwYU7b48

### [Natural Language to Robot Motions](https://seattle.aitinkerers.org/talks/rsvp_xJ2TlLuu5u0)

Zettaware is an agentic robotics platform for robot motion planning.

- Event context: Frontier Builds Demo Night: Experiments at the Edge of AI — 2025-11-13 — Seattle
- Public talk page: https://seattle.aitinkerers.org/talks/rsvp_xJ2TlLuu5u0

### [Easy indexing of NASCAR archived footage](https://seattle.aitinkerers.org/talks/rsvp_-X1ysKvGvHA)

I'm excited to show off what I've been doing but actually, I'm looking for feedback and advice. The last time I did this I worked on Bing Image and Video search and it was a custom pipeline and mostly we just wanted to find video clips, not really search within them for specific scenes. For NASCAR I need a low-cost rapid way to index race footage for a project I'm working on. So far, I have used Amazon Rekognition and Twelve Labs. I will demo both but if someone in the group has experience using these or other off the shelf indexing tools and integrating into Media Asset Management solutions I'd love to connect.

- Event context: October Meetup - Science Fair at Foundations — 2025-10-23 — Seattle
- Public talk page: https://seattle.aitinkerers.org/talks/rsvp_-X1ysKvGvHA

### [Creating an AI Helper to add speech input to any windows app!](https://raleigh.aitinkerers.org/talks/rsvp_G4t1qPgdXao)

I will show a novel way to use low level Windows API to integrate local (or cloud-based) speech-to-text model for adding speech input to applications that don’t natively support it on Windows. Existing applications we use today were built before modern AI and LLMs were available. Adding Voice input to an existing application without making any changes to that application is an enabler for all of us. You can use it entirely locally on your PC if you like, so it is private and secure. I will also demo how it can be used with Multi-Modal LLMs to assist you with more advanced tasks. This can be extended to use local and/or open source language models like Qwen and GPT-OSS. This is open source and available for anyone to use, tinker with and apply to other unique problems.

- Event context: AI Tinkerers - Raleigh Inaugural Meetup (September 2025) — 2025-09-30 — Raleigh
- Public talk page: https://raleigh.aitinkerers.org/talks/rsvp_G4t1qPgdXao

### [Aura: A Locally Hosted AI Gaming Companion](https://dc.aitinkerers.org/talks/rsvp_PPJAfIPNvoQ)

Aura is a locally hosted AI companion that observes live gameplay, interprets on-screen activity, and interacts with players through voice-based commentary. It combines screen capture, computer vision, speech recognition, and voice synthesis into a modular system designed for real-time operation without relying on cloud services. Aura adapts to different games, learns from user feedback, and allows dynamic personality and voice customization

- Event context: AI Tinkerers - DC Metro Meetup (July 10th 2025) — 2025-07-10 — DC
- Public talk page: https://dc.aitinkerers.org/talks/rsvp_PPJAfIPNvoQ

### [Sakhr](https://dubai.aitinkerers.org/talks/rsvp_7XBGehpeVwc)

A student assistant app that can generate notes, reminders, and quizzes from lecture transcription

- Event context: AI Tinkerers - Dubai Meetup #5 (February) — 2025-02-05 — Dubai
- Public talk page: https://dubai.aitinkerers.org/talks/rsvp_7XBGehpeVwc

### [AI Call Analyst](https://medellin.aitinkerers.org/talks/rsvp__9HY7DDnB4Y)

El desarrollo en el que estoy trabajando es un sistema que analiza audios de llamadas entre Reclutador y candidato (en inglés) y evalua la llamada del reclutador según critetios establecidos del negocio y también utiliza un modelo de Azure AI Service para evualuar su pronunciación. Todo esto se combina para darle una puntuación final de 0 a 100 al reclutador y ver en qué está fallando, al final se descarga un documento en word con todos los insights de la llamada, esto se hace utilizando los LLM de OpenAI, whisper de OpenAI y un modelo de Azure y se crea una interfaz con Streamlit

- Event context: AI Tinkerers Medellín #9 - 29 de Enero 2025 — 2025-01-29 — Medellín
- Public talk page: https://medellin.aitinkerers.org/talks/rsvp__9HY7DDnB4Y

### [UzbekVoice](https://tashkent.aitinkerers.org/talks/rsvp_7WmwgOqtMgg)

UzbekVoiceAI is an innovative company specializing in speech processing and artificial intelligence solutions for the Uzbek-speaking audience. Our mission is to deliver high-quality speech recognition (STT) and text-to-speech (TTS) technologies for the Uzbek language, enhancing user interaction with digital devices and services. We develop solutions for speech-to-text conversion, text-to-speech synthesis, and large language models (LLMs) tailored to the unique features of the Uzbek language.

- Event context: AI Tinkerers - Tashkent Inaugural Meetup (October) — 2024-10-31 — Tashkent
- Public talk page: https://tashkent.aitinkerers.org/talks/rsvp_7WmwgOqtMgg

## Related Technologies

- [Llama 3](https://aitinkerers.org/technologies/llama-3) ([Markdown](https://aitinkerers.org/technologies/llama-3.md)) — 38 public demos
- [PyTorch](https://aitinkerers.org/technologies/pytorch) ([Markdown](https://aitinkerers.org/technologies/pytorch.md)) — 273 public demos
- [React](https://aitinkerers.org/technologies/react) ([Markdown](https://aitinkerers.org/technologies/react.md)) — 220 public demos
- [Amazon Polly](https://aitinkerers.org/technologies/amazon-polly) ([Markdown](https://aitinkerers.org/technologies/amazon-polly.md)) — 1 public demo
- [Amazon Rekognition](https://aitinkerers.org/technologies/amazon-rekognition) ([Markdown](https://aitinkerers.org/technologies/amazon-rekognition.md)) — 1 public demo
- [Amazon S3](https://aitinkerers.org/technologies/amazon-s3) ([Markdown](https://aitinkerers.org/technologies/amazon-s3.md)) — 8 public demos
- [Amazon Transcribe](https://aitinkerers.org/technologies/amazon-transcribe) ([Markdown](https://aitinkerers.org/technologies/amazon-transcribe.md)) — 2 public demos
- [Azure AI Service](https://aitinkerers.org/technologies/azure-ai-service) ([Markdown](https://aitinkerers.org/technologies/azure-ai-service.md)) — 1 public demo
- [Azure Speech Service](https://aitinkerers.org/technologies/azure-speech-service) ([Markdown](https://aitinkerers.org/technologies/azure-speech-service.md)) — 2 public demos
- [Azure Speech to Text](https://aitinkerers.org/technologies/azure-speech-to-text) ([Markdown](https://aitinkerers.org/technologies/azure-speech-to-text.md)) — 1 public demo
- [BERT](https://aitinkerers.org/technologies/bert) ([Markdown](https://aitinkerers.org/technologies/bert.md)) — 179 public demos
- [Claude](https://aitinkerers.org/technologies/claude) ([Markdown](https://aitinkerers.org/technologies/claude.md)) — 174 public demos
- [Claude-3](https://aitinkerers.org/technologies/claude-3) ([Markdown](https://aitinkerers.org/technologies/claude-3.md)) — 110 public demos
- [CMU Sphinx](https://aitinkerers.org/technologies/cmu-sphinx) ([Markdown](https://aitinkerers.org/technologies/cmu-sphinx.md)) — 2 public demos
- [Coqui TTS](https://aitinkerers.org/technologies/coqui-tts) ([Markdown](https://aitinkerers.org/technologies/coqui-tts.md)) — 1 public demo
- [Demucs](https://aitinkerers.org/technologies/demucs) ([Markdown](https://aitinkerers.org/technologies/demucs.md)) — 1 public demo
- [Docker](https://aitinkerers.org/technologies/docker) ([Markdown](https://aitinkerers.org/technologies/docker.md)) — 147 public demos
- [eSpeak](https://aitinkerers.org/technologies/espeak) ([Markdown](https://aitinkerers.org/technologies/espeak.md)) — 1 public demo
