right signal
Pick Best inference serving

vLLM

As of 14 Sep 2026, vLLM is the RightSignal pick for best inference serving.

The like-for-like Spheron runs on Llama 3.3 70B FP8 on a single H100 (three-engine, March 2026; two-engine, June 2026) put SGLang 1-4% ahead of vLLM on unique-prompt traffic, which Spheron itself calls run-to-run variance.

Current pick
since 20 Jun 2023
1182
days as pick
Next best (challenger)
SGLang

Near-parity overall and faster on agentic/structured workloads, with major production adoption (xAI, AMD, LinkedIn, Cursor).

Why it’s the challenger
Reign history
1182days
current reign
1challenge
held off
20 JUN 2023became pick
0previous
reigns
Last reviewed: 13 Sep 2026 Reviews are continuous. This pick can change when the evidence changes.

Why vLLM is the pick

  • The like-for-like Spheron runs on Llama 3.3 70B FP8 on a single H100 (three-engine, March 2026; two-engine, June 2026) put SGLang 1-4% ahead of vLLM on unique-prompt traffic, which Spheron itself calls run-to-run variance.
  • Both ran vLLM v0.18.0 on the legacy model runner: Model Runner V2 only became the default for dense models in v0.25.0 (11 July 2026) and for every model in v0.29.0 (9 September 2026), so neither run reflects what ships by default today.
  • Pick vLLM for breadth, hardware coverage and release discipline rather than for any current throughput number.
Judged on throughput benchmarks · day-0 model support · ecosystem · operational maturity

Evidence

The sources behind this title’s record.

Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.

Caveats & challengers

  • Less of a two-horse race than it looks: TensorRT-LLM led the same March 2026 Spheron H100 70B run by 8-13% at 1-50 concurrent requests and about 16% at 100, once its engine was compiled.
  • SGLang's edge is prefix-driven rather than general: on Spheron's June 2026 run with an 80% shared 512-token prefix it cut TTFT p50 by 37% at 50 concurrent requests, while on unique prompts the two sat within 1-4% of each other, and enabling vLLM's --enable-prefix-caching narrowed the TTFT gap to roughly 15-18% at 10 concurrent requests, though at 50 Spheron still puts APC-on vLLM at 2,100 effective tok/s against SGLang's 2,550.
  • Measure your own prefix overlap before committing, and re-run on current releases: those runs used vLLM v0.18.0 on the legacy runner, whereas Model Runner V2 has been the default for dense models since v0.25.0 and for every model since v0.29.0 (9 September 2026).
Challenger SGLang

Near-parity overall and faster on agentic/structured workloads, with major production adoption (xAI, AMD, LinkedIn, Cursor).

At a glance

Title
Best inference serving
Licence
Apache-2.0
Pick since
20 Jun 2023
Last reviewed
13 Sep 2026
Title holders to date
1
Official page
github.com/vllm-project/vllm

Title history

Every change, on the record.

View full changelog
Pick 20 Jun 2023 – presentvLLM Current

The like-for-like Spheron runs on Llama 3.3 70B FP8 on a single H100 (three-engine, March 2026; two-engine, June 2026) put SGLang 1-4% ahead of vLLM on unique-prompt traffic, which Spheron itself calls run-to-run variance.

1182 days
as pick