Technology

LPU architecture

The Language Processing Unit (LPU) architecture, developed by Groq, is an ASIC designed for ultra-low-latency AI inference, leveraging a deterministic, software-first approach.

Groq’s Language Processing Unit (LPU) architecture is a custom-built Application-Specific Integrated Circuit (ASIC) engineered to eliminate latency in AI inference, particularly for Large Language Models (LLMs). The LPU achieves its speed by employing a programmable assembly line architecture and integrating memory directly on-chip (SRAM), bypassing the traditional memory wall bottleneck that slows down GPUs. This design enables deterministic execution, meaning every step is predictable to the clock cycle, which allows for a massive effective memory bandwidth—up to 80 TB/s—and delivers real-world performance like running Llama-2 70B at over 300 tokens per second per user. It is purpose-built for the decode phase of generative AI, ensuring near-instant response times.

https://groq.com/

What builders pair with LPU architecture

Projects using both technologies. Select a pairing to see a project.

Pairing: GroqCloud

Photo from the event
Event photo

ScribeWizard and StockBot powered by Groq

Amsterdam · September 25, 2024

Recent Talks & Demos

Showing 1-1 of 1

Members-Only

Sign in to see who built these projects