vLLM
As of 14 Sep 2026, vLLM is the RightSignal pick for best inference serving.
The like-for-like Spheron runs on Llama 3.3 70B FP8 on a single H100 (three-engine, March 2026; two-engine, June 2026) put SGLang 1-4% ahead of vLLM on unique-prompt traffic, which Spheron itself calls run-to-run variance.
Near-parity overall and faster on agentic/structured workloads, with major production adoption (xAI, AMD, LinkedIn, Cursor).
Why it’s the challengercurrent reign
held off
reigns
Why vLLM is the pick
- The like-for-like Spheron runs on Llama 3.3 70B FP8 on a single H100 (three-engine, March 2026; two-engine, June 2026) put SGLang 1-4% ahead of vLLM on unique-prompt traffic, which Spheron itself calls run-to-run variance.
- Both ran vLLM v0.18.0 on the legacy model runner: Model Runner V2 only became the default for dense models in v0.25.0 (11 July 2026) and for every model in v0.29.0 (9 September 2026), so neither run reflects what ships by default today.
- Pick vLLM for breadth, hardware coverage and release discipline rather than for any current throughput number.
Evidence
The sources behind this title’s record.
- vllm.ai/blog/2026-05-11-vllm-tops-artificial-analysis
- vllm.ai/blog/2026-07-16-keeping-vllm-production-quality
- pytorch.org/blog/pytorch-foundation-welcomes-vllm/
- vllm.ai/blog/2025-01-27-v1-alpha-release
- github.com/sgl-project/sglang
- vllm.ai/blog/2024-09-05-perf-update
- vllm.ai/blog/2023-06-20-vllm
Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.
Caveats & challengers
- Less of a two-horse race than it looks: TensorRT-LLM led the same March 2026 Spheron H100 70B run by 8-13% at 1-50 concurrent requests and about 16% at 100, once its engine was compiled.
- SGLang's edge is prefix-driven rather than general: on Spheron's June 2026 run with an 80% shared 512-token prefix it cut TTFT p50 by 37% at 50 concurrent requests, while on unique prompts the two sat within 1-4% of each other, and enabling vLLM's --enable-prefix-caching narrowed the TTFT gap to roughly 15-18% at 10 concurrent requests, though at 50 Spheron still puts APC-on vLLM at 2,100 effective tok/s against SGLang's 2,550.
- Measure your own prefix overlap before committing, and re-run on current releases: those runs used vLLM v0.18.0 on the legacy runner, whereas Model Runner V2 has been the default for dense models since v0.25.0 and for every model since v0.29.0 (9 September 2026).
Near-parity overall and faster on agentic/structured workloads, with major production adoption (xAI, AMD, LinkedIn, Cursor).
At a glance
- Title
- Best inference serving
- Licence
- Apache-2.0
- Pick since
- 20 Jun 2023
- Last reviewed
- 13 Sep 2026
- Title holders to date
- 1
- Official page
- github.com/vllm-project/vllm
Title history
Every change, on the record.
The like-for-like Spheron runs on Llama 3.3 70B FP8 on a single H100 (three-engine, March 2026; two-engine, June 2026) put SGLang 1-4% ahead of vLLM on unique-prompt traffic, which Spheron itself calls run-to-run variance.
1182 daysas pick