SGLang
LMSYS-hosted serving framework for large language models, aimed at agentic and large-scale serving.
0 members list SGLang, unchanged over the last 12 weeks, data as of 4 October 2026
Members who list SGLang in their Stack, by week.
About SGLang
SGLang is a serving framework designed for running large language models and multimodal models in production environments. Developers use it to deploy these models across single GPUs or distributed clusters while achieving low latency and high throughput. The framework supports a broad range of open-source models and various hardware platforms including NVIDIA and AMD GPUs, CPUs, and TPUs. It incorporates performance optimizations such as disaggregated prefill and decode operations, speculative decoding, and custom GPU kernels.
Developed by sgl-project, SGLang is distributed under the Apache 2.0 license. The framework aims to simplify deployment through a straightforward installation process and OpenAI-compatible API endpoints. It is built to handle large-scale inference scenarios and integrates support for diverse model architectures and hardware configurations, making it applicable to various deployment scenarios from research to production use.
Checked 4 October 2026.
Nobody has this tool in their stack yet.
Posts about SGLang
- Llama.cpp alternatives for running local agents: ds4, Magnitude and Strata against llama.cpp, vLLM and Ollama
ds4 runs big MoE models on 96 GB Macs, Magnitude tunes kernels to your device, Strata puts one model on a gaming PC. Every speed claim is self-reported.