Back to Resources
⚙️

SGLang

LLMs

High-performance serving framework for LLMs and multimodal models.

15kstars1.8kforksPython

About

SGLang pairs a fast runtime with a structured generation language, and has become a de-facto standard for high-throughput LLM inference — reportedly powering deployments across hundreds of thousands of GPUs.

Key Features

  • RadixAttention
  • Structured generation
  • High throughput
  • Multi-GPU

Tags

InferenceServingPerformance

Related Resources