# Multimodal Models Projects at AI Tinkerers

> Canonical HTML: https://aitinkerers.org/technologies/multimodal-models
> Markdown URL: https://aitinkerers.org/technologies/multimodal-models.md
> Technology record last updated: 2026-03-17T16:11:11Z
> Generated: 2026-09-21T13:43:17Z

AI systems that process and integrate multiple data modalities—like text, image, and audio—to achieve human-like, context-aware understanding.

Multimodal models fuse disparate data types (text, video, audio) into a single, unified representation, enabling advanced reasoning and generation across modalities. Key players like Google's Gemini 2.5 Pro handle massive 2-million-token contexts, processing entire codebases or two hours of video footage at once. This capability drives real-world applications: a GPT-4o-powered agent can analyze a customer's voice tone and a screenshot simultaneously, and a vision-language model can generate a detailed image description from a simple text prompt. The technology moves AI beyond single-input limitations, delivering a more holistic and versatile intelligence.

- Official technology site: https://cloud.google.com/multimodal-ai
- Public AI Tinkerers demos and talks: 6
- Result page: 1 of 1

## Recent Public Talks and Demos

### [Building an AI WhatsApp guide for the Valencia Fallas festival](https://valencia.aitinkerers.org/talks/rsvp_73Oz6MkCmvU)

I built an AI-powered WhatsApp assistant that acts as a digital guide for the Valencia Fallas festival. Visitors can ask about the main Fallas monuments and receive explanations about their meaning, satire, and artistic concept. For the most important Fallas, the assistant also delivers pre-recorded audio explanations in Spanish and Valencian, allowing visitors to experience them as if they were using an audio guide. The project demonstrates how conversational AI can turn a messaging app into an accessible cultural guide for large public events without requiring users to install a dedicated app.

- Event context: AI Tinkerers Valencia March Meetup — 2026-03-17 — Valencia
- Public talk page: https://valencia.aitinkerers.org/talks/rsvp_73Oz6MkCmvU

### [Omni ingestion RAG](https://medellin.aitinkerers.org/talks/rsvp_FQ91nU7B_8Q)

Ingesta multimodal en aplicación RAG (Retrieval Augmented Generation) empleando unnestructure y modelos multi modales para procesar imágenes, tablas, y texto.

- Event context: AI Tinkerers Medellín #8 - 5 de Diciembre — 2024-12-05 — Medellín
- Public talk page: https://medellin.aitinkerers.org/talks/rsvp_FQ91nU7B_8Q

### [Heal.dev demo](https://paris.aitinkerers.org/talks/rsvp_9fLF913HZS0)

I built Heal.dev https://www.heal.dev/, an AI agent that takes care of website UI testing automatically

- Event context: AI Tinkerers - Paris Meetup on October 15th — 2024-10-15 — Paris
- Public talk page: https://paris.aitinkerers.org/talks/rsvp_9fLF913HZS0

### [Multimodal Groq Demo](https://denver-boulder.aitinkerers.org/talks/rsvp_TiwUC6zS46U)

Groq is the leading low-latency AI inference provider, and we've been cooking up some magic! Demo gods permitting, we'd love to show a sneak peek of some voice and multimodal models we have running on our hardware.

- Event context: AI Tinkerers Denver - June Meetup — 2024-06-11 — Denver
- Public talk page: https://denver-boulder.aitinkerers.org/talks/rsvp_TiwUC6zS46U

### [Gamified Reality](https://sf.aitinkerers.org/talks/rsvp_1ok-_RTnhL4)

This demo shows how you can use the latest small multimodal models to get real-time inferences and then connect those to actions. In this case, I use it to create a gamified reality app where you earn points as you do various real world activities and those are understood by the AI and matched to achievements.

- Event context: AI Tinkerers - San Francisco - April 2024 Meetup — 2024-04-30 — San Francisco
- Public talk page: https://sf.aitinkerers.org/talks/rsvp_1ok-_RTnhL4

### [Make Video Just as Easy as Text: Introduction to Twelve Labs Video Foundation Model](https://sf.aitinkerers.org/talks/rsvp_mWAtc7EvDtg)

Twelve Labs has been developing a general purpose video foundation model to enable multimodal, contextual video understanding so video can be as easy as text.

- Event context: 🤖🔄🧠 AI Tinkerers SF - August Meetup — 2023-08-10 — San Francisco
- Public talk page: https://sf.aitinkerers.org/talks/rsvp_mWAtc7EvDtg

## Related Technologies

- [RAG](https://aitinkerers.org/technologies/rag) ([Markdown](https://aitinkerers.org/technologies/rag.md)) — 147 public demos
- [AI agent](https://aitinkerers.org/technologies/ai-agent) ([Markdown](https://aitinkerers.org/technologies/ai-agent.md)) — 8 public demos
- [ChromeDriver](https://aitinkerers.org/technologies/chromedriver) ([Markdown](https://aitinkerers.org/technologies/chromedriver.md)) — 1 public demo
- [Computer Vision](https://aitinkerers.org/technologies/computer-vision) ([Markdown](https://aitinkerers.org/technologies/computer-vision.md)) — 22 public demos
- [Cypress](https://aitinkerers.org/technologies/cypress) ([Markdown](https://aitinkerers.org/technologies/cypress.md)) — 2 public demos
- [Foundation models](https://aitinkerers.org/technologies/foundation-models) ([Markdown](https://aitinkerers.org/technologies/foundation-models.md)) — 5 public demos
- [GeckoDriver](https://aitinkerers.org/technologies/geckodriver) ([Markdown](https://aitinkerers.org/technologies/geckodriver.md)) — 1 public demo
- [GPT-V](https://aitinkerers.org/technologies/gpt-v) ([Markdown](https://aitinkerers.org/technologies/gpt-v.md)) — 1 public demo
- [Groq](https://aitinkerers.org/technologies/groq) ([Markdown](https://aitinkerers.org/technologies/groq.md)) — 23 public demos
- [Heal](https://aitinkerers.org/technologies/heal) ([Markdown](https://aitinkerers.org/technologies/heal.md)) — 1 public demo
- [Image](https://aitinkerers.org/technologies/image) ([Markdown](https://aitinkerers.org/technologies/image.md)) — 4 public demos
- [Image Processing](https://aitinkerers.org/technologies/image-processing) ([Markdown](https://aitinkerers.org/technologies/image-processing.md)) — 3 public demos
- [Images](https://aitinkerers.org/technologies/images) ([Markdown](https://aitinkerers.org/technologies/images.md)) — 2 public demos
- [KServe](https://aitinkerers.org/technologies/kserve) ([Markdown](https://aitinkerers.org/technologies/kserve.md)) — 1 public demo
- [LLM](https://aitinkerers.org/technologies/llm) ([Markdown](https://aitinkerers.org/technologies/llm.md)) — 123 public demos
- [LLMs](https://aitinkerers.org/technologies/llms) ([Markdown](https://aitinkerers.org/technologies/llms.md)) — 83 public demos
- [NVIDIA Tesla T4](https://aitinkerers.org/technologies/nvidia-tesla-t4) ([Markdown](https://aitinkerers.org/technologies/nvidia-tesla-t4.md)) — 1 public demo
- [Object Detection](https://aitinkerers.org/technologies/object-detection) ([Markdown](https://aitinkerers.org/technologies/object-detection.md)) — 2 public demos
