# Inference Projects at AI Tinkerers

> Canonical HTML: https://aitinkerers.org/technologies/inference
> Markdown URL: https://aitinkerers.org/technologies/inference.md
> Technology record last updated: 2026-02-25T11:29:55Z
> Generated: 2026-09-21T02:52:22Z

Inference is the execution phase: a trained machine learning model processes new, unseen data to generate real-time predictions, like classifying an image or producing text.

Inference is where the value is realized: it’s the moment a trained model (e.g., a massive LLM like Llama 3) stops learning and starts working, applying its knowledge to real-world input. Unlike computationally intensive training, inference is a single, optimized forward pass. This process must be fast, often requiring millisecond latency for real-time applications (autonomous vehicles, live chatbots) or high throughput for batch processing. Hardware optimization is critical: specialized accelerators like NVIDIA GPUs, Google TPUs, or Groq's LPUs handle the matrix multiplications, ensuring the model delivers its prediction or output efficiently and cost-effectively at scale.

- Official technology site: https://huggingface.co/inference
- Public AI Tinkerers demos and talks: 9
- Result page: 1 of 1

## Recent Public Talks and Demos

### [How I spent $800 Building a Rat Terminator with ML (Instead of Hiring an Exterminator)](https://raleigh.aitinkerers.org/talks/rsvp_GQC-HE7VnGs)

There is a slide deck here: https://westcot.io/talks/you-dirty-rat/ Have a live version I may modify for in person demo that solves one of the problems with the stereo vision discussed in the deck. It's hard to do without slides as it's a narrative of starting with an automated trap and all the points (building bearings and movement with all the things I learned along the way).

- Event context: AI Tinkerers Raleigh Meetup — August 12, 2026 — 2026-08-12 — Raleigh
- Public talk page: https://raleigh.aitinkerers.org/talks/rsvp_GQC-HE7VnGs

### [Watch 1 hour highly techincal YouTubes in 5 minutes with AI!](https://seattle.aitinkerers.org/talks/rsvp_cjwx88z4AkE)

AG is an agent that watches YouTube podcasts for you so you know which ones to really dig into. With AG, see in 5 minutes a summary of the YouTube, key quotes, see key blackboard / slide / code sections, jump around key passages, and decide if you should spend the full time on the video. Break down highly techincal episodes from Dwarkesh, Lenny, AI Engineer, and more!

- Event context: AI Tinkerers Seattle Summer Bash — 2026-07-29 — Seattle
- Public talk page: https://seattle.aitinkerers.org/talks/rsvp_cjwx88z4AkE

### [Your Webcam Knows the Geometry - Real-Time Relighting &amp; Single-Shot Novel Views](https://lausanne.aitinkerers.org/talks/rsvp_cMMuinG2WV4)

Two single-view inverse-rendering systems that reconstruct scene geometry from one camera and re-render it under conditions never captured — one changes the light, the other changes the viewpoint, both fast enough to ship. The Relighting part decomposes a live webcam frame into geometry + material (normals/albedo/shading) and re-renders it under any lighting in real time (~24 FPS), showcasing a variety of lighting conditions (multiple colored lights, orbiting and rainbow-rotating lights) The Novel View Synthesis builds a 3D representation of the scene using Gaussian Splatting and renders it from a different angle. It currently runs at 5 FPS but we are hoping to get closer to real time soon.

- Event context: AI Tinkerers Lausanne June 2026 Meetup — 2026-06-25 — Lausanne
- Public talk page: https://lausanne.aitinkerers.org/talks/rsvp_cMMuinG2WV4

### [From Image to Structured Data: Building a Local AI Document OCR Platform for Administrative Workflows](https://tokyo.aitinkerers.org/talks/rsvp_Vs-o12h3f_U)

Administrative and compliance-heavy professions still rely heavily on paper and scanned documents. However, sending sensitive documents to cloud OCR or AI services is often not acceptable due to privacy, regulatory, or client confidentiality requirements. In this talk, I will present a professional web-based OCR processing platform designed for secure, local-first document handling — with a focus on real-world administrative document workflows such as those handled by 行政書士 professionals.

- Event context: AI Tinkerers Tokyo - Toranomon Meetup - February 19, 2026 — 2026-02-19 — Tokyo
- Public talk page: https://tokyo.aitinkerers.org/talks/rsvp_Vs-o12h3f_U

### [Oracle &amp; NVIDIA AI‑Q: A Blueprint for High‑Performance Research Automation](https://singapore.aitinkerers.org/talks/rsvp_M4Vy0ySvk_Y)

Oracle &amp; Nvidia will be presenting a technical deep-dive demo of an AI Research Assistant built using the NVIDIA AI‑Q blueprint approach, implemented with an enterprise-grade Oracle + NVIDIA stack—using Oracle Database 26ai as the system of record for vectors and retrieval, and NVIDIA’s accelerated AI software for embedding, retrieval optimization, and inference. At a high level, this presentation shows how to build and run a production-ready Retrieval-Augmented Generation (RAG) + agentic workflow where: Enterprise documents are ingested and embedded, using NVIDIA’s AI software stack (including NIM microservices and retrieval components). Embeddings are stored directly inside Oracle Database 26ai using its native VECTOR data type—so we don’t need a separate vector database. Semantic retrieval happens in Oracle Database 26ai using SQL (AI Vector Search), enabling vector similarity search combined with enterprise relational filters and governance. The retrieved context is then fed into a reasoning model to generate grounded answers and structured insights. What the Audience Will Learn / Take Away 1) How Oracle + NVIDIA changes enterprise AI architecture NVIDIA AI Enterprise is available natively through the OCI Console, reducing friction in provisioning AI software and accelerating adoption. Oracle and NVIDIA are co‑engineering deeper integrations, including NVIDIA NIM microservices support and NeMo Retriever integration with Oracle Database 26ai, enabling smoother RAG pipelines. 2) Why Oracle Database 26ai is a key differentiator for RAG Instead of deploying a separate vector DB, I’ll show how Oracle Database 26ai provides AI Vector Search directly inside the database, using the VECTOR data type—letting teams store embeddings next to business data and query semantically in SQL. 3) How to build a faster, simpler, more governable RAG workflow The demo will highlight architectural improvements that enterprise teams care about: - fewer moving parts, - less data duplication, - simpler security and governance patterns, - and a clean operational model that aligns with existing Oracle enterprise data platforms. What we’ll Show in the Demo (Step-by-Step) - Document ingestion (technical PDFs and enterprise materials) - Embedding generation using NVIDIA-optimized AI tooling (NIM microservices / retrieval stack). - Semantic retrieval using Oracle AI Vector Search in SQL (top‑K similarity search + filters). - Answer generation using reasoning over retrieved context

- Event context: AI Tinkerers - The Age of AI &amp; Infrastructure (Singapore) — 2026-02-11 — Singapore
- Public talk page: https://singapore.aitinkerers.org/talks/rsvp_M4Vy0ySvk_Y

### [Edge AI computer](https://seattle.aitinkerers.org/talks/rsvp_sDK4qXpxtng)

I'll just show the video of the Co-pilot plugged in a car &amp; driving me + the inference output. If there's an extra display or a laptop, I'd love to use that instead.

- Event context: AI Tinkerers Seattle - January 2025 Meetup — 2025-01-23 — Seattle
- Public talk page: https://seattle.aitinkerers.org/talks/rsvp_sDK4qXpxtng

### [Speeding up inference with Effort Engine](https://poland.aitinkerers.org/talks/rsvp__2MDM6B1P6s)

Founder of Effort Engine (new algorithm for LLM Inference)

- Event context: AI Tinkerers Poland - First Inaugural Meetup in Warsaw (November) — 2024-11-21 — Poland
- Public talk page: https://poland.aitinkerers.org/talks/rsvp__2MDM6B1P6s

### [Repeated inference in practice](https://hamburg.aitinkerers.org/talks/rsvp_WxCKoC_z4CQ)

We use LLMs to identify which elements on a website are most relevant for changes based on a specific marketing strategy. This is a complex task, and due to the size of the prompt, the structured output from the LLM can sometimes be unstable and may lack precision and recall. By running multiple inferences and combining and reranking the outputs, we can achieve better stability and quality in the results.

- Event context: AI Tinkerers Hamburg - September 12 — 2024-09-12 — Hamburg
- Public talk page: https://hamburg.aitinkerers.org/talks/rsvp_WxCKoC_z4CQ

### [Anatomy of a Thinking Machine](https://la.aitinkerers.org/talks/rsvp_MjNJg6eHsLw)

We all know that AI has a hardware problem. - What happens at a hardware level during inference? - What exactly are all these tools in the inference ecosystems from Nvidia, AMD...? - An early preview of Cortex, an open source tool that runs LLMs across multiple platforms Here's my cofounder Dan doing a similar talk, but I'm planning for this demo to be **shorter &amp; purely technical**: https://www.youtube.com/watch?v=orcPcUzSbOw&amp;ab_channel=HackerHouseTW

- Event context: June 25th - LA AI Tinkerers Meetup &amp; Demos — 2024-06-26 — Los Angeles
- Public talk page: https://la.aitinkerers.org/talks/rsvp_MjNJg6eHsLw

## Related Technologies

- [BERT](https://aitinkerers.org/technologies/bert) ([Markdown](https://aitinkerers.org/technologies/bert.md)) — 179 public demos
- [BLOOM](https://aitinkerers.org/technologies/bloom) ([Markdown](https://aitinkerers.org/technologies/bloom.md)) — 115 public demos
- [Data](https://aitinkerers.org/technologies/data) ([Markdown](https://aitinkerers.org/technologies/data.md)) — 8 public demos
- [GPT-3](https://aitinkerers.org/technologies/gpt-3) ([Markdown](https://aitinkerers.org/technologies/gpt-3.md)) — 191 public demos
- [GPT-4](https://aitinkerers.org/technologies/gpt-4) ([Markdown](https://aitinkerers.org/technologies/gpt-4.md)) — 529 public demos
- [Llama-2](https://aitinkerers.org/technologies/llama-2) ([Markdown](https://aitinkerers.org/technologies/llama-2.md)) — 227 public demos
- [PaLM 2](https://aitinkerers.org/technologies/palm-2) ([Markdown](https://aitinkerers.org/technologies/palm-2.md)) — 116 public demos
- [RoBERTa](https://aitinkerers.org/technologies/roberta) ([Markdown](https://aitinkerers.org/technologies/roberta.md)) — 118 public demos
- [ABBYY FineReader](https://aitinkerers.org/technologies/abbyy-finereader) ([Markdown](https://aitinkerers.org/technologies/abbyy-finereader.md)) — 3 public demos
- [AI models](https://aitinkerers.org/technologies/ai-models) ([Markdown](https://aitinkerers.org/technologies/ai-models.md)) — 6 public demos
- [Algorithm](https://aitinkerers.org/technologies/algorithm) ([Markdown](https://aitinkerers.org/technologies/algorithm.md)) — 2 public demos
- [Amazon Textract](https://aitinkerers.org/technologies/amazon-textract) ([Markdown](https://aitinkerers.org/technologies/amazon-textract.md)) — 5 public demos
- [AMD](https://aitinkerers.org/technologies/amd) ([Markdown](https://aitinkerers.org/technologies/amd.md)) — 1 public demo
- [Anthropic](https://aitinkerers.org/technologies/anthropic) ([Markdown](https://aitinkerers.org/technologies/anthropic.md)) — 36 public demos
- [Bambu labs printer](https://aitinkerers.org/technologies/bambu-labs-printer) ([Markdown](https://aitinkerers.org/technologies/bambu-labs-printer.md)) — 1 public demo
- [Claude](https://aitinkerers.org/technologies/claude) ([Markdown](https://aitinkerers.org/technologies/claude.md)) — 173 public demos
- [Cloud Vision API](https://aitinkerers.org/technologies/cloud-vision-api) ([Markdown](https://aitinkerers.org/technologies/cloud-vision-api.md)) — 3 public demos
- [Codex](https://aitinkerers.org/technologies/codex) ([Markdown](https://aitinkerers.org/technologies/codex.md)) — 44 public demos
