TL;DR
Last updated September 30, 2026.
SGLang and vLLM are the two most used open-source LLM serving engines. Both are Apache 2.0 and both run an OpenAI-compatible server with continuous batching, prefix caching, speculative decoding, structured outputs, and multi-GPU parallelism. Pick SGLang when requests share long prefixes (agent loops, multi-turn chat, RL rollouts), because RadixAttention reuses that KV cache automatically. Pick vLLM when you want the widest model and hardware plugin coverage, multi-LoRA, or Anthropic Messages API compatibility. Current releases are vLLM v0.30.0 (September 22, 2026) and SGLang v0.5.20 (September 18, 2026).
SGLang wins on
Shared-prefix traffic. RadixAttention keeps finished requests' KV cache in a radix tree and reuses it for any later request with the same prefix, which is the common case in agents, few-shot prompting, and RL rollout sampling.
vLLM wins on
Breadth. PagedAttention memory management, the largest set of supported model architectures, hardware plugins from TPU to Gaudi to Ascend, efficient multi-LoRA, and an OpenAI server that also speaks the Anthropic Messages API and gRPC.
SGLang vs vLLM Comparison Table
Every row below comes from the two projects' own READMEs, documentation, and GitHub release pages, checked September 30, 2026.
| SGLang | vLLM | |
|---|---|---|
| Origin | LMSYS | Sky Computing Lab, UC Berkeley |
| License | Apache 2.0 | Apache 2.0 |
| Latest release | v0.5.20 (Sept 18, 2026) | v0.30.0 (Sept 22, 2026) |
| Core KV cache idea | RadixAttention (radix-tree prefix reuse) | PagedAttention (block-based KV memory) |
| Prefix caching | Automatic, built into RadixAttention | Yes (automatic prefix caching) |
| Continuous batching | Yes | Yes, plus chunked prefill |
| Speculative decoding | Yes, with SpecForge for training draft models | n-gram, suffix, EAGLE, DFlash |
| Parallelism | Tensor, pipeline, data, expert, context | Tensor, pipeline, data, expert, context |
| Structured outputs | JSON, regex, EBNF | xgrammar or guidance |
| APIs | OpenAI-compatible | OpenAI-compatible, Anthropic Messages, gRPC |
| LoRA serving | Yes, via OpenAI-compatible APIs | Multi-LoRA for dense and MoE layers |
| Disaggregated serving | Prefill/decode (PD) and EPD disaggregation | Prefill, decode, and encode |
| RL rollout integrations | Miles, slime, AReaL, Tunix, verl | RLHF weight-sync examples in docs |
| Install | uv pip install --prerelease=allow sglang | uv pip install vllm |
| Docker image | lmsysorg/sglang:latest | vllm/vllm-openai |
PagedAttention vs RadixAttention
vLLM's PagedAttention, described in the 2023 paper "Efficient Memory Management for Large Language Model Serving with PagedAttention" (Kwon et al.), treats the KV cache like virtual memory. It splits each sequence's cache into fixed-size blocks that do not have to be contiguous, so the server wastes almost no GPU memory on fragmentation and can fit more concurrent sequences into a batch.
SGLang's RadixAttention, introduced in the LMSYS post of January 17, 2024, solves a different problem. Instead of discarding a request's KV cache when it finishes, SGLang keeps it in a radix tree keyed by token sequence. When a new request arrives, the scheduler matches its longest cached prefix and only computes the remainder. Multi-turn chat, a shared system prompt, few-shot examples, and tree-of-thought search all reuse cache this way without any manual configuration.
vLLM now ships automatic prefix caching, and both engines support disaggregated prefill and decode. Prefix reuse is still the core of SGLang's scheduler, so prefix-heavy workloads remain its strongest case.
SGLang vs vLLM Performance and Benchmarks
LMSYS's launch benchmark reported SGLang reaching up to 5x the throughput of vLLM and Guidance on Llama-7B and Mixtral-8x7B on NVIDIA A10G GPUs, measured on chained LLM tasks with heavy prefix sharing. That number is from January 2024 and from the SGLang authors, so read it as the upper bound for prefix-heavy work, not a general result. vLLM has added prefix caching, CUDA graphs, torch.compile kernel generation, and disaggregated prefill and decode since then.
| Workload | Likely leader | Why |
|---|---|---|
| Agent loops with a long shared system prompt | SGLang | Radix-tree prefix reuse skips recomputing the shared tokens |
| RL rollouts, many samples per prompt | SGLang | Same prompt, many completions; native integrations with rollout frameworks |
| Multi-turn chat | SGLang | Earlier turns stay cached between requests |
| Single-shot prompts with no overlap | Close | Prefix reuse has nothing to reuse; kernels and batching dominate |
| Many LoRA adapters on one base model | vLLM | Efficient multi-LoRA for dense and MoE layers |
| Exotic accelerators (Gaudi, Spyre, Rebellions) | vLLM | Broader hardware plugin ecosystem |
Benchmark on your own traffic. Throughput between the two engines swings with the model, the GPU, batch size, input and output lengths, and above all how much prefix your requests share. Both projects document a benchmarking workflow for replaying a request mix against a running server.
Hardware Support
| Platform | SGLang | vLLM |
|---|---|---|
| NVIDIA GPUs | A100, H100/H200/H800/H20, B200/B300/GB200/GB300, RTX 30-50, DGX Spark, Jetson Orin | Yes |
| AMD GPUs | Instinct MI300X, MI325X, MI350X, MI355X | Yes |
| Google TPU | v6e, v7 (SGL-JAX, SGL-torchtpu) | Via hardware plugin |
| Intel | Arc / Arc Pro B-Series GPUs, CPU servers | Intel GPUs, Gaudi via plugin |
| Huawei Ascend | A2, A3, 950PR/DT NPUs | Via hardware plugin |
| Apple Silicon | Metal / MLX | Not listed in README |
| CPUs | CPU servers | x86, ARM, PowerPC |
APIs and Launch Commands
Both servers accept requests from the standard OpenAI client, so switching between them is mostly a base URL change.
# vLLM: OpenAI-compatible server on http://localhost:8000/v1
uv pip install vllm
vllm serve Qwen/Qwen3-4B
# SGLang: OpenAI-compatible server on http://localhost:30000/v1
uv pip install --prerelease=allow sglang
python3 -m sglang.launch_server --model-path Qwen/Qwen3-4B --host 0.0.0.0vLLM's server also implements the Anthropic Messages API, so tools that expect an Anthropic endpoint can point at a self-hosted model without a proxy. Both engines ship tool-calling and reasoning parsers per model family; a mismatched parser is the most common reason tool calls come back as plain text.
RL Rollouts and Agent Workloads
SGLang describes itself as optimized for agentic workloads, RL rollouts, and large-scale serving. Its README lists Miles, slime, AReaL, Tunix, and verl as training frameworks that call SGLang for rollout generation. RL sampling draws many completions from one prompt, and agent loops resend a growing transcript every step, so both reuse long prefixes on almost every request.
vLLM covers the same ground with RLHF examples for weight sync over HTTP, IPC, and NCCL. For a broader view of where engine choice sits among the other levers, see LLM inference optimization, continuous batching, speculative decoding, and prompt caching. To see how a minimal engine is put together, read nano-vLLM.
SGLang vs vLLM Pros and Cons
- SGLang: automatic prefix reuse via RadixAttention
- SGLang: first-class RL rollout integrations
- SGLang: Apple Silicon and TPU v6e/v7 support
- vLLM: widest model architecture coverage
- vLLM: Anthropic Messages API and gRPC alongside OpenAI
- vLLM: efficient multi-LoRA for dense and MoE
- SGLang: pip install needs --prerelease=allow
- vLLM: trailed SGLang by up to 5x on prefix-heavy tasks in LMSYS's 2024 benchmark
- Both: tool-call parsers must match the model family
Which One to Pick
Start from the shape of your traffic. If most requests share a long system prompt or transcript, or you are generating RL rollouts, run SGLang first. If you serve many different models, need dozens of LoRA adapters, depend on an accelerator that only has a vLLM plugin, or want an Anthropic-compatible endpoint, run vLLM first. In every case, replay a day of real requests against both before you standardize, since the right answer moves with each release.
If you would rather not operate either engine, Morph serves open models on its own kernels through an OpenAI-compatible API. See Morph Open Source Models and the dedicated inference calculator.
Frequently Asked Questions
What is the difference between SGLang and vLLM?
Both are open-source (Apache 2.0) LLM serving engines with OpenAI-compatible servers. vLLM, from UC Berkeley's Sky Computing Lab, is built around PagedAttention, a block-based KV cache memory manager, and has the widest model and hardware plugin coverage. SGLang, from LMSYS, is built around RadixAttention, which keeps finished requests' KV cache in a radix tree and reuses it for any later request with a matching prefix. SGLang targets agentic workloads, RL rollouts, and large-scale serving.
Is SGLang faster than vLLM?
It depends on the workload. When many requests share a long prefix (agent loops, multi-turn chat, few-shot prompts, RL rollouts with a shared system prompt), SGLang's RadixAttention reuses cached KV automatically and usually wins on throughput. LMSYS reported up to 5x higher throughput than vLLM on such tasks in its January 2024 launch post. vLLM has since added prefix caching too, so on single-shot prompts with little overlap the two are close and results depend on the model, GPU, and batch size. Benchmark both on your own traffic.
What are the latest versions of SGLang and vLLM?
As of September 30, 2026, the latest GitHub releases are vLLM v0.30.0, published September 22, 2026, and SGLang v0.5.20, published September 18, 2026. Both projects ship releases every few weeks.
Do SGLang and vLLM both support the OpenAI API?
Yes. vLLM runs an OpenAI-compatible server and also supports the Anthropic Messages API and gRPC. SGLang exposes OpenAI-compatible completions and chat completions endpoints, and by default listens on port 30000. vLLM's server listens on port 8000 by default.
Which hardware do SGLang and vLLM support?
vLLM supports NVIDIA, AMD, and Intel GPUs plus x86, ARM, and PowerPC CPUs, with hardware plugins for Google TPU, Intel Gaudi, IBM Spyre, Huawei Ascend, and others. SGLang supports NVIDIA (A100 through B300 and GB300, plus DGX Spark and Jetson Orin), AMD Instinct MI300X to MI355X, Google TPU v6e and v7, Intel Arc GPUs, CPU servers, Apple Silicon via Metal and MLX, Huawei Ascend NPUs, and Moore Threads.
Which engine is better for RL training?
SGLang. Its README lists Miles, slime, AReaL, Tunix, and verl as training frameworks that integrate SGLang for rollout generation, and RL rollouts sample many completions from the same prompt, which is exactly where RadixAttention prefix reuse pays off. vLLM is also used for RLHF and ships weight-sync examples, so both work, but SGLang is the more common rollout engine.
The fastest endpoints are private deployments
Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 100 billion tokens per day run this way.
Skip running the serving stack yourself
Morph serves open-weight models on custom kernels behind an OpenAI-compatible API, with dedicated deployments when you need guaranteed throughput.
Sources
- vLLM GitHub README (features, hardware, install)
- vLLM v0.30.0 release, September 22, 2026
- Kwon et al., Efficient Memory Management for LLM Serving with PagedAttention
- SGLang GitHub README (hardware, ecosystem, install)
- SGLang v0.5.20 release, September 18, 2026
- LMSYS: Fast and Expressive LLM Inference with RadixAttention and SGLang
- SGLang documentation (OpenAI-compatible APIs)