Technology

oMLX

oMLX is an open-source, macOS-native inference server that optimizes local LLM workflows on Apple Silicon using a two-tier, paged SSD KV cache.

Built specifically for Apple Silicon, oMLX bypasses unified memory limits by offloading inactive context blocks to your SSD in safetensors format. This two-tier KV cache architecture drops time-to-first-token (TTFT) from 90 seconds to under 5 seconds during long-context agentic workflows (like running Claude Code or Cursor). By combining continuous batching via mlx-lm with a native macOS menu bar app and an OpenAI-compatible API, it delivers up to 3x faster generation speeds on standard hardware without taxing your system's active RAM.

https://omlx.ai

What builders pair with oMLX

Projects using both technologies. Select a pairing to see a project.

Pairing: FastAPI

Photo from the event
Event photo

Scrape, sense, snipe: local LLMs reading Twitter to trade Polymarket

Zürich · July 1, 2026

Recent Talks & Demos

Showing 1-1 of 1

Members-Only

Sign in to see who built these projects