Technology

LiteRT-LM

LiteRT-LM is a production-ready orchestration layer for deploying high-performance Large Language Models directly to edge devices.

LiteRT-LM serves as the specialized GenAI infrastructure for Google's AI Edge ecosystem, built specifically to handle the complexities of on-device LLM deployment. It manages critical tasks like KV-cache optimization, prompt templating, and function calling while leveraging the foundational LiteRT runtime for hardware acceleration across CPUs, GPUs, and NPUs. Currently powering features in Chrome and the Pixel Watch, it supports a broad range of open models—including Gemma 4, Llama, and Phi-4—delivering up to 1.4x faster GPU performance than legacy TFLite delegates. By utilizing the .litertlm format, developers can implement multi-modal workflows and agentic tool-use capabilities on Android, iOS, and web platforms with minimal latency.

https://ai.google.dev/edge/litert

Recent Talks & Demos

Showing 1-0 of 0

Members-Only

Sign in to see who built these projects

No public projects found for this technology yet.