# RAG Projects at AI Tinkerers

> Canonical HTML: https://aitinkerers.org/technologies/rag
> Markdown URL: https://aitinkerers.org/technologies/rag.md
> Technology record last updated: 2026-09-18T15:13:52Z
> Generated: 2026-09-22T17:54:02Z

RAG (Retrieval-Augmented Generation) is the GenAI framework that grounds LLMs (like GPT-4) on external, verified data, drastically reducing model hallucinations and providing verifiable sources.

RAG is a critical GenAI architecture: it solves the LLM 'hallucination' problem by inserting a retrieval step before generation. A user query is vectorized, then used to query an external knowledge base (e.g., a Pinecone vector database) for relevant document chunks (typically 512-token segments). These retrieved facts augment the original prompt, providing the LLM (e.g., Gemini or Llama 3) the specific, current, or proprietary context required. This process ensures the final response is accurate and grounded in domain-specific data, avoiding the high cost and latency of full model retraining.

- Official technology site: https://en.wikipedia.org/wiki/Retrieval-augmented_generation
- Public AI Tinkerers demos and talks: 147
- Result page: 1 of 7

## Recent Public Talks and Demos

### [Bias-Free RAG: Simulating Political Candidates with AI](https://saopaulo.aitinkerers.org/talks/rsvp_uTVDpnRMXT4)

Café com o Candidato is an open-source site where users pick a 2026 Brazilian presidential candidate and chat with an AI simulation powered by RAG over public data, replying with the candidate's speech patterns and mannerisms — always with a clear disclaimer that it's a simulation with no affiliation. In the live demo I'll show the chat working end-to-end and, behind it, the RAG pipeline: how a user's question pulls the most relevant chunks from public sources via pgvector before reaching the model.

- Event context: AI Tinkerers SP e Oracle - Meetup Agosto — 2026-08-27 — São Paulo
- Public talk page: https://saopaulo.aitinkerers.org/talks/rsvp_uTVDpnRMXT4

### [PalliAssist: Transforming Palliative Care with Compassionate AI](https://mombasa.aitinkerers.org/talks/rsvp_lhA4geGRg08)

PalliAssist is an AI-powered palliative care companion that helps patients, caregivers, and healthcare providers manage symptoms, medications, appointments, and access trusted care guidance through compassionate, personalized support. During this demo, we'll showcase the working web application, walk through the complete patient and caregiver workflow, demonstrate how Gemma powers real-time AI conversations, explain our system architecture and RAG pipeline, highlight key sections of our codebase on GitHub, and show how the platform delivers intelligent, accessible, and privacy-conscious palliative care in low-resource settings.

- Event context: AI Tinkerers – Mombasa Chapter Launch · 22 August 2026 — 2026-08-22 — Mombasa
- Public talk page: https://mombasa.aitinkerers.org/talks/rsvp_lhA4geGRg08

### [Automated context compaction strategies](https://san-diego.aitinkerers.org/talks/rsvp_7lJPz-RERho)

This talk demonstrates how leveraging concepts from database design to manage context and semantically indexed time-weighted memories can improve the quality of interactions with LLMs while reducing token usage. It introduces a new concept called Context Structured Merge that only retrieves the most relevant historical context on demand, reducing token usage over the lifetime of a model interaction.

- Event context: Self-hosting Models and Managing Token Spend — 2026-08-21 — San Diego
- Public talk page: https://san-diego.aitinkerers.org/talks/rsvp_7lJPz-RERho

### [How we managed to recover a stolen bike thanks to my AI-powered platform](https://valencia.aitinkerers.org/talks/rsvp_ObUQQ3GZ0Ds)

I'll demo all the AI integrations I did for our subscription bike company.

- Event context: AI Tinkerers Valencia July Demo Night — 2026-07-28 — Valencia
- Public talk page: https://valencia.aitinkerers.org/talks/rsvp_ObUQQ3GZ0Ds

### [The Completion Utility Stack: Launching Thousands of AI-Native Businesses for the Agentic Economy](https://orange-county.aitinkerers.org/talks/rsvp_g0ZRnVDKJI8)

I built NetShow IQ1, the full agentic operating stack from NetShow.AI for creating AI-first, AI-native businesses where digital crews move users from intent to completed outcome. IQ1 is designed around the Completion Utility: the idea that, just as electricity, water, and gas became foundational utilities for modern life, reliable task completion becomes a new utility for the agentic economy. In the live demonstration, I’ll show how IQ1 turns a request like “I need this handled” into an orchestrated workflow across virtual agents, tools, memory, MCPs, skills, approvals, and reporting. I’ll walk through the working system, architecture, agent harness, workflow routing, tool calls, logs, and how a business-specific digital crew can be composed for real consumer and business services. NetShow.AI is building economic infrastructure for thousands of AI-first businesses and services across major categories of life, work, commerce, local services, operations, support, home, and environment.

- Event context: AI Tinkerers Orange County: Tuesday, July 21, 2026 at Centercode — 2026-07-22 — Orange County
- Public talk page: https://orange-county.aitinkerers.org/talks/rsvp_g0ZRnVDKJI8

### [Oracle AI para Saúde: Agentes Inteligentes, Memória Persistente e Dados Multimodelo](https://saopaulo.aitinkerers.org/talks/rsvp_-dbi8r0iY-c)

Veja como a Oracle está impulsionando a próxima geração de soluções de IA para saúde com agentes omnichannel que mantêm contexto entre WhatsApp, voz e e-mail, além de recursos avançados como Long-Term Memory e o Autonomous Multi-Model Database para criar aplicações inteligentes, seguras e escaláveis.

- Event context: AI Tinkerers SP - Meetup de Junho — 2026-06-25 — São Paulo
- Public talk page: https://saopaulo.aitinkerers.org/talks/rsvp_-dbi8r0iY-c

### [Apertus: SwissAI’s fully-transparent multilingual LLM](https://lausanne.aitinkerers.org/talks/rsvp_yoFNhymPd9I)

Apertus is an Apache 2.0 family of open language models. This demo focuses on the Mini instruct variants, distilled into 0.5B, 1.5B, and 4B parameter checkpoints and uses Transformers.js and WebGPU for fast, client-side inference.

- Event context: AI Tinkerers Lausanne June 2026 Meetup — 2026-06-25 — Lausanne
- Public talk page: https://lausanne.aitinkerers.org/talks/rsvp_yoFNhymPd9I

### [Iterating on AI features](https://valencia.aitinkerers.org/talks/rsvp_WrlMEwVSia0)

Showing PostHog's AI features.

- Event context: AI Tinkerers Valencia June Meetup ft. PostHog — 2026-06-16 — Valencia
- Public talk page: https://valencia.aitinkerers.org/talks/rsvp_WrlMEwVSia0

### [Utilizing lightweight AI models for the Modern Storefront ecommerce](https://amman.aitinkerers.org/talks/rsvp_mWfpqXVV_RA)

Ecommerce storefront app with enhanced features using small LLMs

- Event context: AI Tinkerers Amman: Agentic Workflows &amp; Self-Hosted LLMs — 2026-05-30 — Amman
- Public talk page: https://amman.aitinkerers.org/talks/rsvp_mWfpqXVV_RA

### [Patent Mining for Engineers: Building an Agentic RAG System for Inventive Problem Solving using TRIZ &amp; AI](https://poland.aitinkerers.org/talks/rsvp_oYsjqZvaY7E)

A pipeline that parses patent PDFs, extracts Technical Contradictions, classifies solutions into TRIZ Inventive Principles, and indexes everything into a vector database. This collection is then feeding an AI Agent that helps engineers solve real inventive problems. The demo starts with a raw patent PDF, submits it live to the processing endpoint, and walks through what gets extracted and indexed. Then, given a real mechanical engineering problem, the agent frames it as a TRIZ contradiction, retrieves relevant patents, and proposes concrete solution ideas, powered by domain knowledge, not just plain LLM generation.

- Event context: AI Tinkerers Poland #3 - Meetup in Wrocław — 2026-05-06 — Poland
- Public talk page: https://poland.aitinkerers.org/talks/rsvp_oYsjqZvaY7E

### [Agents Building Agents: Reflective Optimization Loops](https://toronto.aitinkerers.org/talks/rsvp_4gS7qartFFk)

I'll show how an AI agent can build and optimize another AI agent, using reflective optimization to find issues, optimize evals, iterate on architecture, find the optimal prompt/model, and more. We've built a system with multiple levels of reflective optimization for agent development. - GEPA: reflective prompt optimization - Synthetic eval generation: going from a 1-off bug to an proper eval you can use in reflective optimization - Expanding reflective optimization beyond prompts: model selection, tool use, subagents -- reflective optimization can drive all levels of agent optimization.

- Event context: AI Tinkerers Toronto - April 2026 - hosted by Shopify — 2026-04-29 — Toronto
- Public talk page: https://toronto.aitinkerers.org/talks/rsvp_4gS7qartFFk

### [Scaling RAG: Hybrid Search and Hierarchical Chunking for 780k Pages](https://poland.aitinkerers.org/talks/rsvp_BCaEvuBCHLM)

I built a custom desktop-server search engine designed to help me instantly find and manage documents within my 40GB PDF library. Technical Overview: - The Interface: A Windows application where I can search and browse through the results easily. - The Search Brain: A backend powered by FastAPI that uses "hybrid search" - Data Processing: Python and Bash scripts that handle the heavy lifting, such as pulling Markdown and generating page thumbnails from every file. - Annotation AI: vLLM based LLM server that extract metadata. - The Future: I am currently adding a RAG (Retrieval-Augmented Generation) feature so I can ask the AI questions directly about the content of my documents.

- Event context: AI Tinkerers Poland - Meetup in Gdańsk #1 — 2026-04-23 — Poland
- Public talk page: https://poland.aitinkerers.org/talks/rsvp_BCaEvuBCHLM

### [Compose and Dragons: Tiny Language Models in Action](https://paris.aitinkerers.org/talks/rsvp_lNiq-CojExE)

Let's debunk some beliefs about (very) small LLMs, those that make less than 4b of parameters. We often hear: They are useless and do not know how to do anything, they know nothing, they are bad at calling (so no MCP) This is partly wrong, and we can fix the rest and build generative AI systems with these very small models. Among other things, we will see how to create NPCs with a personality, a master dungeon that will manage your movements, fights... and allow you to talk to this or that NPC ...

- Event context: AI Tinkerers Paris: Docker Agentic Workflows (Devoxx Kickoff) — 2026-04-21 — Paris
- Public talk page: https://paris.aitinkerers.org/talks/rsvp_lNiq-CojExE

### [Stop the Confident BS: Reflective Retrieval Agents and Human-in-the-Loop Interrupts](https://cologne.aitinkerers.org/talks/rsvp_LyKSpKSeRVQ)

We built a reflective agent prototype that evaluates its own retrieved context and halts for human clarification before it hallucinates. In the live demo, we'll first break a standard one-prompt RAG setup to show how it confidently gives answers when faced with poorly defined context. Then, we will query our prototype, showcasing the live execution. You will see the agent evaluate its context, hit an uncertainty threshold, trigger a Human-in-the-Loop (HITL) interrupt to ask for missing parameters, and finally generate a factually grounded answer.

- Event context: AI Tinkerers Cologne 4: Live Technical Demos — 2026-04-16 — Cologne
- Public talk page: https://cologne.aitinkerers.org/talks/rsvp_LyKSpKSeRVQ

### [Your AI is Two Versions Behind](https://zurich.aitinkerers.org/talks/rsvp_k0YRpX9Yc4E)

A simple, repeatable system for filling in what AI models don’t know. Every model has a training cutoff, and everything after that date is a blind spot. I built structured markdown folders that capture what changed in a language or framework since the cutoff, and feed them directly into AI coding tools as project context. My first implementation covers Go 1.25 and 1.26, but the approach works for anything. I’ll demo the folder structure, show the before/after difference in AI output, and walk through gobot, a zero-dependency Go app built entirely with AI that had the missing knowledge loaded.

- Event context: AI Tinkerers Zurich April 9th — 2026-04-09 — Zürich
- Public talk page: https://zurich.aitinkerers.org/talks/rsvp_k0YRpX9Yc4E

### [Miró: Synthetic Audience Analysis using LLM Agents and Graph Database](https://manizales.aitinkerers.org/talks/rsvp_QJ7Kot1_hP8)

In this demo, I will do a code deep-dive into "Miró", an engine I built to forecast the social and critical reception of upcoming books using synthetic readers. Instead of a product pitch, I will focus entirely on the technical architecture and the integration layer between LLMs and Graph Database. I'll walk through the code live, showing: Agent Generation Pipeline: How I parse static PDFs containing psychological profiles and translate them into "Synthetic Reader" nodes with their respective master prompts. The Predictive Engine: The orchestration code that drives how agents "read" the book, interact with each other, and how these interactions continuously update the graph database state. Graph-based Memory Management: A look into how I solved the challenge of persistent agent memory by dynamically creating complex relationships (such as CHATTED_ABOUT) and appending conversation histories as extendable edge properties using RAG. Analysis Dashboard: A quick look at how the Python backend consumes this dynamic graph network to feed an interactive react frontend, rendering resonance, friction, and abandonment connections.

- Event context: 🚀 ¡14vo Encuentro de AI Tinkerers Manizales! 🤖 — 2026-03-25 — Manizales
- Public talk page: https://manizales.aitinkerers.org/talks/rsvp_QJ7Kot1_hP8

### [\[UofT\] Beyond Baseline RAG: Building a Reliable and Transparent Policy Chatbot for Government Guidance](https://toronto.aitinkerers.org/talks/rsvp_oPuu6wo8x8Q)

Government procurement policies are often complex, lengthy, and distributed across multiple directives and guidance documents. Employees seeking clarification must manually search through these materials, which can be time-consuming and may lead to inconsistent interpretations. In this talk, we present the design of a retrieval-augmented policy chatbot that assists users in navigating procurement policies by answering questions directly from source documents while providing transparent citations. Our system uses a Retrieval-Augmented Generation (RAG) architecture to ground responses in official policy text. Documents are segmented and embedded into a retrieval index, allowing the system to surface relevant policy excerpts in response to natural language queries. The language model then generates answers strictly from the retrieved evidence and provides citations so users can verify the source material. Beyond a baseline RAG implementation, the project explores several mechanisms to improve reliability and transparency. These include structure-aware document chunking, detection of contradictory or overlapping policy statements, and a self-verification loop that checks generated answers for unsupported claims. The system also incorporates user feedback signals and an evaluation pipeline that measures retrieval accuracy, evidence grounding, and response quality. Together, these components aim to demonstrate how AI assistants can support policy interpretation while maintaining transparency, accountability, and user trust.

- Event context: AI Tinkerers Toronto - March - hosted by Mozilla! — 2026-03-25 — Toronto
- Public talk page: https://toronto.aitinkerers.org/talks/rsvp_oPuu6wo8x8Q

### [\[UofT\] Give Your Local File System Memory - Intelligent Document Reference](https://toronto.aitinkerers.org/talks/rsvp_xdm4yT8hgkU)

Our application is an intelligent document search and question-answering system designed to help users quickly find information within their personal files. Instead of relying on file names or exact keyword matches, the system analyzes the actual content of documents and allows users to search using natural language queries. The application automatically indexes files from the user’s file system, extracts their content, and organizes the information in a way that makes it easy to retrieve later. When a user asks a question or searches for a topic, the system identifies the most relevant files and sections of text, then returns either the file paths or a summarized answer supported with citations to the original documents. The system supports multiple file formats, including documents, spreadsheets, images, and text files, enabling users to search across different types of data in one place. By combining semantic search with AI-powered reasoning, the application helps users navigate large collections of files more efficiently and quickly locate the information they need.

- Event context: AI Tinkerers Toronto - March - hosted by Mozilla! — 2026-03-25 — Toronto
- Public talk page: https://toronto.aitinkerers.org/talks/rsvp_xdm4yT8hgkU

### [Coding with AI: What Works &amp; What Doesn't](https://manchester-nh.aitinkerers.org/talks/rsvp_WlqmhQoEJnw)

We all know AI dev tools are powerful, but they can also be incredibly frustrating if you use them wrong. In this session, I'll share my personal "dos and don'ts" for navigating the current landscape of AI assistants. We'll cover the right (and wrong) ways to use different tools, and how to avoid the common traps that actually slow you down. Finally, I’ll show you what happens when you build your own rules: a live demo of a self-learning agent I built with Claude Code and mcp.json that records sessions, writes its own memories, and gets smarter as I code.

- Event context: AI Tinkerers Manchester (Bedford), NH - March 2026 Meetup — 2026-03-18 — Manchester NH
- Public talk page: https://manchester-nh.aitinkerers.org/talks/rsvp_WlqmhQoEJnw

### [VLLM and Qdrant - GPU goes Brrrr!](https://manchester-nh.aitinkerers.org/talks/rsvp_RGPw96tcjiA)

This demo goes over the fundamentals of VLLM and the QDrant vector database. We'll spin up some Docker containers with the LLM, Database and Embedding model, and then run some interesting benchmarks. I'll demonstrate just how much more powerful VLLM can be on hardware when compared to sequential model runners.

- Event context: AI Tinkerers Manchester (Bedford), NH - March 2026 Meetup — 2026-03-18 — Manchester NH
- Public talk page: https://manchester-nh.aitinkerers.org/talks/rsvp_RGPw96tcjiA

### [Building an AI WhatsApp guide for the Valencia Fallas festival](https://valencia.aitinkerers.org/talks/rsvp_73Oz6MkCmvU)

I built an AI-powered WhatsApp assistant that acts as a digital guide for the Valencia Fallas festival. Visitors can ask about the main Fallas monuments and receive explanations about their meaning, satire, and artistic concept. For the most important Fallas, the assistant also delivers pre-recorded audio explanations in Spanish and Valencian, allowing visitors to experience them as if they were using an audio guide. The project demonstrates how conversational AI can turn a messaging app into an accessible cultural guide for large public events without requiring users to install a dedicated app.

- Event context: AI Tinkerers Valencia March Meetup — 2026-03-17 — Valencia
- Public talk page: https://valencia.aitinkerers.org/talks/rsvp_73Oz6MkCmvU

### [From AI Agent Demo to Enterprise Reality Usecases](https://ho-chi-minh-city.aitinkerers.org/talks/rsvp_tu-93CcrdmY)

With experience building internal products and an AI-first mission, GreenNode – a member of VNG Group – has pioneered the development of AI agent use cases both inside the company and for external customers. In this talk, the speaker will systematically share the journey of deploying real-world agent use cases in-house (Project Manager Agent, product-focused RAG chatbot, and internal knowledge-base RAG chatbot) and how these agents are being adopted across the enterprise.

- Event context: AI Tinkerers Ho Chi Minh City: From Prompt to Agent — 2026-03-07 — Ho Chi Minh City
- Public talk page: https://ho-chi-minh-city.aitinkerers.org/talks/rsvp_tu-93CcrdmY

### [ai-flow.eu Evaluation Suite: Systematic Testing for LLM Apps](https://cologne.aitinkerers.org/talks/rsvp_FslCeIYc_OY)

I will demo the ai-flow.eu evaluation suite and walk through how it is built end to end. We will start with a minimal evaluation project and define a dataset of prompts and expected behavior. Then we will run automated evaluations across multiple models and configurations and capture structured results. I will show how we: - version datasets and test cases - run deterministic checks plus LLM judged scoring - track regressions between runs - generate a simple report that is useful for engineers The focus is on code, architecture, and the practical workflow to make LLM changes measurable.

- Event context: AI Tinkerers Cologne 3: Demos, Code, and Architecture — 2026-03-05 — Cologne
- Public talk page: https://cologne.aitinkerers.org/talks/rsvp_FslCeIYc_OY

### [Let's talk about Embeddings](https://cologne.aitinkerers.org/talks/rsvp_NX6I1fvuINY)

I will talk about why embeddings are such a great thing. They can do so many tasks that we set out a huge LLM to do, but in a much more efficient and cost saving way. There are tons of use cases for embeddings, and in this talk, I just want to give a simple insight into some use cases of embeddings, beside RAG. I want to cover (not sure if this is the final list yet): - RAG - Image Search - Image Classifier - Advanced Image Classifier with an added MLP Head - Text Matching across languages - Getting Clear Text Input for Customer Intention Analysis (Main Focus) - And a short example of how you can use that clear text input to improve what you are offering as a company. (Main Focus) As the 5 Minute Time slot is very narrow, I will likely focus on the Clear Text Input Analysis part, as I think that is quite a nice use case for embedding based, customer facing search. While I will not show a lot of code in this presentation, coding this yourself is so easy, that anyone could do it without seeing any code. It's more about the idea and concept for this usecase.

- Event context: AI Tinkerers Cologne 3: Demos, Code, and Architecture — 2026-03-05 — Cologne
- Public talk page: https://cologne.aitinkerers.org/talks/rsvp_NX6I1fvuINY

## Related Technologies

- [GPT-4](https://aitinkerers.org/technologies/gpt-4) ([Markdown](https://aitinkerers.org/technologies/gpt-4.md)) — 529 public demos
- [GPT-3](https://aitinkerers.org/technologies/gpt-3) ([Markdown](https://aitinkerers.org/technologies/gpt-3.md)) — 191 public demos
- [BERT](https://aitinkerers.org/technologies/bert) ([Markdown](https://aitinkerers.org/technologies/bert.md)) — 179 public demos
- [RoBERTa](https://aitinkerers.org/technologies/roberta) ([Markdown](https://aitinkerers.org/technologies/roberta.md)) — 118 public demos
- [BLOOM](https://aitinkerers.org/technologies/bloom) ([Markdown](https://aitinkerers.org/technologies/bloom.md)) — 115 public demos
- [Llama-2](https://aitinkerers.org/technologies/llama-2) ([Markdown](https://aitinkerers.org/technologies/llama-2.md)) — 227 public demos
- [PaLM 2](https://aitinkerers.org/technologies/palm-2) ([Markdown](https://aitinkerers.org/technologies/palm-2.md)) — 116 public demos
- [OpenAI API](https://aitinkerers.org/technologies/openai-api) ([Markdown](https://aitinkerers.org/technologies/openai-api.md)) — 520 public demos
- [Python](https://aitinkerers.org/technologies/python) ([Markdown](https://aitinkerers.org/technologies/python.md)) — 662 public demos
- [Embeddings](https://aitinkerers.org/technologies/embeddings) ([Markdown](https://aitinkerers.org/technologies/embeddings.md)) — 22 public demos
- [Generative AI](https://aitinkerers.org/technologies/generative-ai) ([Markdown](https://aitinkerers.org/technologies/generative-ai.md)) — 45 public demos
- [LLM](https://aitinkerers.org/technologies/llm) ([Markdown](https://aitinkerers.org/technologies/llm.md)) — 123 public demos
- [GraphRAG](https://aitinkerers.org/technologies/graphrag) ([Markdown](https://aitinkerers.org/technologies/graphrag.md)) — 13 public demos
- [FastAPI](https://aitinkerers.org/technologies/fastapi) ([Markdown](https://aitinkerers.org/technologies/fastapi.md)) — 181 public demos
- [LangChain](https://aitinkerers.org/technologies/langchain) ([Markdown](https://aitinkerers.org/technologies/langchain.md)) — 445 public demos
- [React](https://aitinkerers.org/technologies/react) ([Markdown](https://aitinkerers.org/technologies/react.md)) — 220 public demos
- [Vector database](https://aitinkerers.org/technologies/vector-database) ([Markdown](https://aitinkerers.org/technologies/vector-database.md)) — 13 public demos
- [ChatGPT](https://aitinkerers.org/technologies/chatgpt) ([Markdown](https://aitinkerers.org/technologies/chatgpt.md)) — 83 public demos

## More Results

- Next: https://aitinkerers.org/technologies/rag.md?page=2
