Technology
oMLX
oMLX is an open-source, macOS-native inference server that optimizes local LLM workflows on Apple Silicon using a two-tier, paged SSD KV cache.
Built specifically for Apple Silicon, oMLX bypasses unified memory limits by offloading inactive context blocks to your SSD in safetensors format. This two-tier KV cache architecture drops time-to-first-token (TTFT) from 90 seconds to under 5 seconds during long-context agentic workflows (like running Claude Code or Cursor). By combining continuous batching via mlx-lm with a native macOS menu bar app and an OpenAI-compatible API, it delivers up to 3x faster generation speeds on standard hardware without taxing your system's active RAM.
What builders pair with oMLX
Projects using both technologies. Select a pairing to see a project.
Pairing: FastAPI
Scrape, sense, snipe: local LLMs reading Twitter to trade Polymarket
Pairing: pgvector
Scrape, sense, snipe: local LLMs reading Twitter to trade Polymarket
Pairing: PostgreSQL
Scrape, sense, snipe: local LLMs reading Twitter to trade Polymarket
Pairing: React
Scrape, sense, snipe: local LLMs reading Twitter to trade Polymarket
Recent Talks & Demos
Showing 1-1 of 1