MasterNodeAI
news

Model Serving Runtimes for MoE: vLLM vs SGLang vs TensorRT-LLM Performance Benchmarks

Maximize AI speed and ROI by discovering the best inference engine for your Mixture of Experts deployment in this MoE model serving runtime comparison.

news

Model Serving Runtimes for MoE: vLLM vs SGLang vs TensorRT-LLM Performance Benchmarks

Choosing the wrong runtime for a Mixture-of-Experts workload doesn't just leave performance on the table — it changes your unit economics entirely. The MoE model serving runtime comparison between vLLM, SGLang, and TensorRT-LLM reveals throughput spreads as wide as 2.5x on identical hardware, which at H100 on-demand rates of $8–$12/hour translates directly into cost-per-token penalties that compound at production scale. Since DeepSeek-V3's release drove enterprise adoption of self-hosted MoE inference sharply upward through Q1 2025, the decision has become unavoidable for any team running Mixtral 8x22B, DeepSeek-V3 (671B total / ~37B active), or Qwen2.5-MoE in a real production environment.

This analysis is for ML engineers and infrastructure leads who need to match a runtime to their GPU fleet — H100, H200, or L40S — given their team's operational capacity and whether they optimize primarily for throughput, deployment speed, or long-term flexibility. By the end, you'll have a clear first deployment choice and an honest picture of the trade-offs you're accepting.


The Five Dimensions That Actually Differentiate MoE Runtimes

Generic LLM benchmarks are nearly useless for this decision. They underrepresent the expert-routing bottleneck that is specific to MoE architectures — top-k expert routing alone adds 5–15% latency overhead that every runtime must absorb before delivering a single useful token to your application.

The evaluation framework that matters covers five areas: raw MoE throughput on standardized models; memory efficiency under the unique VRAM pressure MoE creates (all expert weights must reside in memory regardless of which experts are active per token); multi-GPU scaling given the all-to-all communication costs that emerge at expert parallelism boundaries; operational complexity including time-to-first-request and toolchain dependencies; and ecosystem extensibility covering quantization, speculative decoding, and multi-tenant serving capabilities.

GPU utilization is the hidden variable connecting most of these dimensions. MoE workloads characteristically land at 40–60% GPU utilization due to irregular memory access patterns from expert routing, compared to 70–85% for dense models. The runtime's job is to minimize how much of that structural inefficiency it amplifies through poor scheduling or suboptimal kernel selection.


vLLM: The Generalist with MoE Growing Pains

vLLM's expert parallelism support, added in v0.6.x, delivered up to 1.8x throughput improvement for Mixtral 8x7B on multi-GPU configurations — a meaningful advance that made it viable for MoE workloads rather than merely functional. Its PagedAttention memory management handles MoE's irregular access patterns reasonably well, though the architecture was designed for dense models and the seams show under heavy MoE load.

The honest number to anchor on: vLLM benchmarks show 1.5–3x lower throughput than SGLang on DeepSeek-class models. The gap has two primary causes. First, vLLM's prefix caching is less aggressive than SGLang's RadixAttention implementation, which matters significantly for multi-turn or shared-prefix workloads. Second, vLLM's expert kernel optimization has not received the same focused engineering attention as its dense-model path. GPU utilization for MoE workloads on vLLM tends toward the lower end of that 40–50% range.

Where vLLM earns its place is ecosystem breadth. It runs the widest range of models, has the most comprehensive documentation, provides an OpenAI-compatible API server out of the box, and requires the least friction to go from Hugging Face weights to a running endpoint. If your team already operates vLLM for dense models and MoE is an addition rather than a primary workload, the migration cost of switching stacks likely exceeds the throughput benefit — particularly on mixed-workload infrastructure where a single serving layer simplifies operations considerably.

At H100 on-demand pricing, the throughput gap versus SGLang represents a real cost-per-token penalty at scale. On a 100M token/day production workload, the difference between 40% and the throughput SGLang achieves on the same hardware is measurable in thousands of dollars monthly. That math changes the calculus once MoE becomes your primary serving workload.


SGLang: The Current Open-Source Throughput Leader for Large MoE

SGLang is the runtime to beat for pure MoE throughput in the open-source space. January 2025 benchmarks demonstrated 2.5x higher throughput than vLLM on DeepSeek-V3, and LMSYS reported serving DeepSeek-R1 (671B total / ~37B active) at 52 tokens/second on 8×H200 — numbers that set the competitive baseline for what optimized MoE serving looks like.

The architectural differentiator is RadixAttention-based prefix caching. For applications with repeated system prompts, multi-turn conversations, or RAG pipelines that prepend retrieved context, RadixAttention avoids recomputing the KV cache for shared prefixes across requests. In practice, this means SGLang's throughput advantage compounds with workload complexity: a simple completion API gets 2.5x better throughput; a multi-tenant chat application with shared system prompts can see the gap widen further. This makes SGLang particularly well-suited for inference APIs serving multiple tenants from a single model instance.

The grouped GEMM optimizations and expert parallelism implementation in SGLang reflect focused engineering on MoE-specific kernel performance rather than bolted-on support. Combined with the H200's higher memory bandwidth, SGLang on 8×H200 hardware positions teams to reach the $0.05–$0.10 per million token cost floor that optimized DeepSeek-V3 clusters can achieve.

The real trade-offs are operational. SGLang has a steeper setup curve than vLLM, thinner third-party integration coverage, and documentation that consistently lags the development pace. Its battle-testing outside the DeepSeek model family is less established — teams running Mixtral 8x22B in production on SGLang have a smaller community to draw on when debugging edge cases. The runtime is open-source with no licensing fees, so hardware remains the cost driver, but the engineering hours required to configure and maintain SGLang correctly are a non-trivial input cost for smaller teams.


TensorRT-LLM: Maximum Performance, Maximum Commitment

TensorRT-LLM represents the performance ceiling for MoE serving on NVIDIA hardware, and the ceiling is meaningfully higher than the alternatives. NVIDIA's own benchmarks reported DeepSeek-V3 inference at approximately 60 tokens/second on 8×H200 — above SGLang's published 52 tokens/second — and approximately 1.5x lower latency versus vLLM baseline on Mixtral 8x7B running on H100s. Native MoE support added in late 2024 brought expert parallelism and hardware-specific grouped GEMM kernels tuned for Hopper and Ada architectures.

The compiled-engine model is the source of both TensorRT-LLM's performance advantage and its operational burden. Rather than executing dynamically, TensorRT-LLM compiles a model into an engine file optimized for a specific hardware configuration. That compilation extracts maximum throughput from the target GPU's capabilities — H100, H200, and B200 each get kernels designed for their specific memory hierarchies. W4A16 expert quantization in TensorRT-LLM delivers approximately 50% memory footprint reduction with less than 1% quality degradation, which directly addresses the VRAM pressure that constrains MoE serving on any runtime.

The operational cost is real. Every model update requires recompiling an engine, extending iteration cycles substantially. Teams that ship model updates frequently — A/B testing fine-tunes, swapping to new checkpoint versions, experimenting with quantization levels — will feel this friction acutely. There is no AMD or alternative hardware path; the tighter the NVIDIA coupling, the harder the eventual exit if your hardware strategy changes. TensorRT-LLM is open-source under Apache 2.0, but the practical lock-in comes from the NVIDIA toolchain dependencies, not the license.

The right profile for TensorRT-LLM is a dedicated NVIDIA H100/H200/B200 cluster, a model that updates infrequently, and a latency SLA that justifies the engineering investment in compilation and operations overhead. Enterprises with active NVIDIA support agreements get additional value from the support path. For real-time inference applications — sub-100ms time-to-first-token requirements, high-concurrency production APIs — TensorRT-LLM's compiled performance ceiling is worth the operational investment.


Head-to-Head Comparison

DimensionvLLMSGLangTensorRT-LLM
DeepSeek-V3 throughput (8×H200)Baseline~2.5x vLLM (52 tok/s, R1)~60 tok/s (NVIDIA benchmark)
Mixtral 8x7B latency vs. vLLMBaseline~1.5–2x better~1.5x lower latency
MoE expert parallelismYes (v0.6.x+)Yes (optimized)Yes (native, Hopper/Ada tuned)
Prefix cachingBasicRadixAttention (best-in-class)Limited
W4A16 expert quantizationPartialPartialFull support (~50% VRAM reduction)
GPU utilization (MoE)40–50%50–60%60–75%
Speculative decodingIn developmentSupportedSupported
Time to first requestLow (hours)Medium (days)High (days–weeks, compilation)
Model update agilityHighHighLow (recompile required)
Hardware lock-inNoneNoneNVIDIA only
Multi-tenant / shared prefixModerateExcellentLimited
Best hardware tierH100, L40S, H200H200, H100H100, H200, B200
LicenseApache 2.0Apache 2.0Apache 2.0

The Decision Framework

Three questions determine the correct answer for your team.

Is DeepSeek-V3 or DeepSeek-R1 your primary model? Deploy SGLang first. The 2.5x throughput advantage over vLLM on these architectures is verified and consistent, and the multi-tenant prefix caching story is directly valuable for chat and RAG applications. Accept the setup investment; it pays back within weeks at production volume.

Do you have a dedicated NVIDIA H100/H200/B200 cluster with stable models and hard latency SLAs? TensorRT-LLM's compiled performance ceiling is the right target. Budget the engineering time for compilation workflows, treat model updates as planned events rather than continuous iterations, and the operational complexity becomes a fixed cost against a real performance return.

Are you running mixed workloads — dense models alongside MoE — or is operational simplicity a genuine constraint? Start with vLLM. The throughput penalty is real but bounded, the documentation and community will save your team hours per week, and the path from Hugging Face weights to production is the shortest available. Migrate to SGLang for MoE-specific deployments once the workload justifies the infrastructure specialization.

The $0.05–$0.15 per million token cost floor on optimized H200 clusters is achievable — but only if the runtime is matched to both the model architecture and the team's capacity to extract performance from it. Hardware cost is fixed; how efficiently your runtime uses that hardware is the variable you control.