Technology
SGLang
SGLang is a high-performance, open-source serving framework designed to accelerate large language model and multimodal inference with advanced prefix caching and structured output control.
Developed to optimize complex LLM workflows, SGLang delivers low-latency, high-throughput serving across hardware configurations ranging from a single local GPU to massive distributed clusters. The engine achieves its speed through a co-designed system: a flexible front-end language that simplifies structured generation, paired with a blazing-fast back-end runtime featuring RadixAttention for automatic KV cache reuse. By natively supporting continuous batching, speculative decoding, and tensor parallelism, SGLang enables developers to run frontier models (like Llama 3 and DeepSeek) at maximum hardware efficiency while maintaining strict control over JSON schemas and structured API outputs.
Recent Talks & Demos
Showing 1-0 of 0