Back to Resources
⚙️
SGLang
LLMsHigh-performance serving framework for LLMs and multimodal models.
15kstars1.8kforksPython
About
SGLang pairs a fast runtime with a structured generation language, and has become a de-facto standard for high-throughput LLM inference — reportedly powering deployments across hundreds of thousands of GPUs.
Key Features
- RadixAttention
- Structured generation
- High throughput
- Multi-GPU
Tags
InferenceServingPerformance