SGLang vs vLLM (2026): Features, Performance, Hardware, and Which to Pick

SGLang vs vLLM compared on current releases (vLLM v0.30.0, Sept 22 2026; SGLang v0.5.20, Sept 18 2026). PagedAttention vs RadixAttention, prefix caching, parallelism, hardware support, APIs, RL rollout use, launch commands, and when each one wins.

September 30, 2026 · 2 min read

TL;DR

Last updated September 30, 2026.

SGLang and vLLM are the two most used open-source LLM serving engines. Both are Apache 2.0 and both run an OpenAI-compatible server with continuous batching, prefix caching, speculative decoding, structured outputs, and multi-GPU parallelism. Pick SGLang when requests share long prefixes (agent loops, multi-turn chat, RL rollouts), because RadixAttention reuses that KV cache automatically. Pick vLLM when you want the widest model and hardware plugin coverage, multi-LoRA, or Anthropic Messages API compatibility. Current releases are vLLM v0.30.0 (September 22, 2026) and SGLang v0.5.20 (September 18, 2026).

SGLang wins on

Shared-prefix traffic. RadixAttention keeps finished requests' KV cache in a radix tree and reuses it for any later request with the same prefix, which is the common case in agents, few-shot prompting, and RL rollout sampling.

vLLM wins on

Breadth. PagedAttention memory management, the largest set of supported model architectures, hardware plugins from TPU to Gaudi to Ascend, efficient multi-LoRA, and an OpenAI server that also speaks the Anthropic Messages API and gRPC.

v0.30.0
vLLM latest release, Sept 22 2026
v0.5.20
SGLang latest release, Sept 18 2026
:8000
vLLM default server port
:30000
SGLang default server port

SGLang vs vLLM Comparison Table

Every row below comes from the two projects' own READMEs, documentation, and GitHub release pages, checked September 30, 2026.

SGLang vs vLLM feature comparison (September 2026)
SGLangvLLM
OriginLMSYSSky Computing Lab, UC Berkeley
LicenseApache 2.0Apache 2.0
Latest releasev0.5.20 (Sept 18, 2026)v0.30.0 (Sept 22, 2026)
Core KV cache ideaRadixAttention (radix-tree prefix reuse)PagedAttention (block-based KV memory)
Prefix cachingAutomatic, built into RadixAttentionYes (automatic prefix caching)
Continuous batchingYesYes, plus chunked prefill
Speculative decodingYes, with SpecForge for training draft modelsn-gram, suffix, EAGLE, DFlash
ParallelismTensor, pipeline, data, expert, contextTensor, pipeline, data, expert, context
Structured outputsJSON, regex, EBNFxgrammar or guidance
APIsOpenAI-compatibleOpenAI-compatible, Anthropic Messages, gRPC
LoRA servingYes, via OpenAI-compatible APIsMulti-LoRA for dense and MoE layers
Disaggregated servingPrefill/decode (PD) and EPD disaggregationPrefill, decode, and encode
RL rollout integrationsMiles, slime, AReaL, Tunix, verlRLHF weight-sync examples in docs
Installuv pip install --prerelease=allow sglanguv pip install vllm
Docker imagelmsysorg/sglang:latestvllm/vllm-openai

PagedAttention vs RadixAttention

vLLM's PagedAttention, described in the 2023 paper "Efficient Memory Management for Large Language Model Serving with PagedAttention" (Kwon et al.), treats the KV cache like virtual memory. It splits each sequence's cache into fixed-size blocks that do not have to be contiguous, so the server wastes almost no GPU memory on fragmentation and can fit more concurrent sequences into a batch.

SGLang's RadixAttention, introduced in the LMSYS post of January 17, 2024, solves a different problem. Instead of discarding a request's KV cache when it finishes, SGLang keeps it in a radix tree keyed by token sequence. When a new request arrives, the scheduler matches its longest cached prefix and only computes the remainder. Multi-turn chat, a shared system prompt, few-shot examples, and tree-of-thought search all reuse cache this way without any manual configuration.

The gap has narrowed

vLLM now ships automatic prefix caching, and both engines support disaggregated prefill and decode. Prefix reuse is still the core of SGLang's scheduler, so prefix-heavy workloads remain its strongest case.

SGLang vs vLLM Performance and Benchmarks

LMSYS's launch benchmark reported SGLang reaching up to 5x the throughput of vLLM and Guidance on Llama-7B and Mixtral-8x7B on NVIDIA A10G GPUs, measured on chained LLM tasks with heavy prefix sharing. That number is from January 2024 and from the SGLang authors, so read it as the upper bound for prefix-heavy work, not a general result. vLLM has added prefix caching, CUDA graphs, torch.compile kernel generation, and disaggregated prefill and decode since then.

Where each engine tends to lead, by workload shape
WorkloadLikely leaderWhy
Agent loops with a long shared system promptSGLangRadix-tree prefix reuse skips recomputing the shared tokens
RL rollouts, many samples per promptSGLangSame prompt, many completions; native integrations with rollout frameworks
Multi-turn chatSGLangEarlier turns stay cached between requests
Single-shot prompts with no overlapClosePrefix reuse has nothing to reuse; kernels and batching dominate
Many LoRA adapters on one base modelvLLMEfficient multi-LoRA for dense and MoE layers
Exotic accelerators (Gaudi, Spyre, Rebellions)vLLMBroader hardware plugin ecosystem

Benchmark on your own traffic. Throughput between the two engines swings with the model, the GPU, batch size, input and output lengths, and above all how much prefix your requests share. Both projects document a benchmarking workflow for replaying a request mix against a running server.

Hardware Support

Supported hardware, from each project's README
PlatformSGLangvLLM
NVIDIA GPUsA100, H100/H200/H800/H20, B200/B300/GB200/GB300, RTX 30-50, DGX Spark, Jetson OrinYes
AMD GPUsInstinct MI300X, MI325X, MI350X, MI355XYes
Google TPUv6e, v7 (SGL-JAX, SGL-torchtpu)Via hardware plugin
IntelArc / Arc Pro B-Series GPUs, CPU serversIntel GPUs, Gaudi via plugin
Huawei AscendA2, A3, 950PR/DT NPUsVia hardware plugin
Apple SiliconMetal / MLXNot listed in README
CPUsCPU serversx86, ARM, PowerPC

APIs and Launch Commands

Both servers accept requests from the standard OpenAI client, so switching between them is mostly a base URL change.

# vLLM: OpenAI-compatible server on http://localhost:8000/v1
uv pip install vllm
vllm serve Qwen/Qwen3-4B

# SGLang: OpenAI-compatible server on http://localhost:30000/v1
uv pip install --prerelease=allow sglang
python3 -m sglang.launch_server --model-path Qwen/Qwen3-4B --host 0.0.0.0

vLLM's server also implements the Anthropic Messages API, so tools that expect an Anthropic endpoint can point at a self-hosted model without a proxy. Both engines ship tool-calling and reasoning parsers per model family; a mismatched parser is the most common reason tool calls come back as plain text.

RL Rollouts and Agent Workloads

SGLang describes itself as optimized for agentic workloads, RL rollouts, and large-scale serving. Its README lists Miles, slime, AReaL, Tunix, and verl as training frameworks that call SGLang for rollout generation. RL sampling draws many completions from one prompt, and agent loops resend a growing transcript every step, so both reuse long prefixes on almost every request.

vLLM covers the same ground with RLHF examples for weight sync over HTTP, IPC, and NCCL. For a broader view of where engine choice sits among the other levers, see LLM inference optimization, continuous batching, speculative decoding, and prompt caching. To see how a minimal engine is put together, read nano-vLLM.

SGLang vs vLLM Pros and Cons

Strengths
  • SGLang: automatic prefix reuse via RadixAttention
  • SGLang: first-class RL rollout integrations
  • SGLang: Apple Silicon and TPU v6e/v7 support
  • vLLM: widest model architecture coverage
  • vLLM: Anthropic Messages API and gRPC alongside OpenAI
  • vLLM: efficient multi-LoRA for dense and MoE
Limitations
  • SGLang: pip install needs --prerelease=allow
  • vLLM: trailed SGLang by up to 5x on prefix-heavy tasks in LMSYS's 2024 benchmark
  • Both: tool-call parsers must match the model family

Which One to Pick

Start from the shape of your traffic. If most requests share a long system prompt or transcript, or you are generating RL rollouts, run SGLang first. If you serve many different models, need dozens of LoRA adapters, depend on an accelerator that only has a vLLM plugin, or want an Anthropic-compatible endpoint, run vLLM first. In every case, replay a day of real requests against both before you standardize, since the right answer moves with each release.

If you would rather not operate either engine, Morph serves open models on its own kernels through an OpenAI-compatible API. See Morph Open Source Models and the dedicated inference calculator.

Frequently Asked Questions

What is the difference between SGLang and vLLM?

Both are open-source (Apache 2.0) LLM serving engines with OpenAI-compatible servers. vLLM, from UC Berkeley's Sky Computing Lab, is built around PagedAttention, a block-based KV cache memory manager, and has the widest model and hardware plugin coverage. SGLang, from LMSYS, is built around RadixAttention, which keeps finished requests' KV cache in a radix tree and reuses it for any later request with a matching prefix. SGLang targets agentic workloads, RL rollouts, and large-scale serving.

Is SGLang faster than vLLM?

It depends on the workload. When many requests share a long prefix (agent loops, multi-turn chat, few-shot prompts, RL rollouts with a shared system prompt), SGLang's RadixAttention reuses cached KV automatically and usually wins on throughput. LMSYS reported up to 5x higher throughput than vLLM on such tasks in its January 2024 launch post. vLLM has since added prefix caching too, so on single-shot prompts with little overlap the two are close and results depend on the model, GPU, and batch size. Benchmark both on your own traffic.

What are the latest versions of SGLang and vLLM?

As of September 30, 2026, the latest GitHub releases are vLLM v0.30.0, published September 22, 2026, and SGLang v0.5.20, published September 18, 2026. Both projects ship releases every few weeks.

Do SGLang and vLLM both support the OpenAI API?

Yes. vLLM runs an OpenAI-compatible server and also supports the Anthropic Messages API and gRPC. SGLang exposes OpenAI-compatible completions and chat completions endpoints, and by default listens on port 30000. vLLM's server listens on port 8000 by default.

Which hardware do SGLang and vLLM support?

vLLM supports NVIDIA, AMD, and Intel GPUs plus x86, ARM, and PowerPC CPUs, with hardware plugins for Google TPU, Intel Gaudi, IBM Spyre, Huawei Ascend, and others. SGLang supports NVIDIA (A100 through B300 and GB300, plus DGX Spark and Jetson Orin), AMD Instinct MI300X to MI355X, Google TPU v6e and v7, Intel Arc GPUs, CPU servers, Apple Silicon via Metal and MLX, Huawei Ascend NPUs, and Moore Threads.

Which engine is better for RL training?

SGLang. Its README lists Miles, slime, AReaL, Tunix, and verl as training frameworks that integrate SGLang for rollout generation, and RL rollouts sample many completions from the same prompt, which is exactly where RadixAttention prefix reuse pays off. vLLM is also used for RLHF and ships weight-sync examples, so both work, but SGLang is the more common rollout engine.

Private deployments

The fastest endpoints are private deployments

Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 100 billion tokens per day run this way.

Talk to us about a private deployment

Skip running the serving stack yourself

Morph serves open-weight models on custom kernels behind an OpenAI-compatible API, with dedicated deployments when you need guaranteed throughput.

Sources