Vllm performance tuning
Skill air-gapped/skills/.claude/skills/vllm-performance-tuning
Claude Code plugin marketplace — 58 installable reference skills across vLLM/SGLang inference, Kubernetes & Harvester, GPU host bring-up, observability, security, and agent workflows.
npx -y skills add air-gapped/skills --skill vllm-performance-tuningAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
vLLM performance-tuning operator reference — tuning workflow (baseline → bottleneck → knob → re-bench), fused-MoE kernel autotune (`benchmark_moe.py` generates `E=N,N=M,device_name=X.json` configs), DeepEP all-to-all + expert parallelism + EPLB, CUDA graph modes (FULL_AND_PIECEWISE default), torch.compile AOT + compile cache, scheduler knobs (`--max-num-batched-tokens`, `--max-num-seqs`, `--async-scheduling`), TP/EP/DP/PP decision tree, NCCL/DCGM on H100/H200/B200/GB200, PD disaggregation (Nixl/Mooncake/LMCache), known regressions + vendor quirks (v0.14→0.15.1 MiniMax, MI300X FP8<BF16, DeepGEMM M<128 TTFT).
SKILL.md
14.5 KB, ~4.1k tokens by cl100k_base, as published. Nobody here has run it
vLLM performance tuning
Target: operators deploying models on new hardware, chasing throughput / latency / goodput SLOs, or diagnosing perf regressions. Current through v0.25.1 (2026-07-14). Last freshened 2026-07-21.
Companion skills: vllm-benchmarking (measure), vllm-caching (KV), vllm-nvidia-hardware (GPU/GEMM), vllm-configuration (env vars), vllm-observability (metrics).
Tuning levers (apply by goal, not in fixed order)
Always first — characterize the workload. ISL / OSL / req/s / concurrency / SLO (P95 TTFT, P95 TPOT, P95 ITL). "Goodput" = tok/s/GPU under SLO, not raw tok/s. Everything below is keyed off these numbers.
Parallelism + MoE kernels (biggest single wins):
- Pick parallelism (see
references/moe-and-ep.md) — model-fits-1-GPU → TP=1 + replicas (DP); MoE MLA (DeepSeek/Kimi-K2) → DP-attn + EP; multi-node → TP intra + PP inter OR Wide-EP. - MoE on a new SKU → run
benchmark_moe.py --tune— generatesE=*,N=*,device_name=*.jsonconfigs. Without tuned configs vLLM logs "Using default MoE config. Performance might be sub-optimal!" = 20-40% throughput loss. - Wide-EP (
--enable-expert-parallel --enable-eplb --enable-dbo) for DeepSeek/Qwen3/Kimi-K2 at ≥16 GPUs.
Throughput / batching:
auto_tune.sh(benchmarks/auto_tune/) sweepsmax_num_seqs × max_num_batched_tokens.--gpu-memory-utilization— raise from 0.90 toward 0.95 until steady OOM margin, then back off. MoE: cap at 0.85 (all-to-all buffers not in accounting).- Chunked prefill (always on in V1) — raise
--max-num-batched-tokens(default 2048 since PR #10544) if TTFT > SLO; lower if ITL > SLO.
Latency / graph + compile:
- CUDA graphs — keep
FULL_AND_PIECEWISE(default); align--cuda-graph-sizeswithmax_num_seqs*2. --async-scheduling— default-on in recent releases unless using spec-dec / PP / unsupported MM path.- Compile cache — pre-bake
$VLLM_CACHE_ROOT/torch_compile_cacheon a representative pod; mount as PVC / bake into OCI layer.
Distributed (last resort, only when the lever above is exhausted):
- NCCL — on well-configured clouds do nothing. Bare-metal IB:
NCCL_IB_HCA,NCCL_IB_GID_INDEX,NCCL_NET_GDR_LEVEL. NeverNCCL_CUMEM_ENABLE=0on GB200. - PD disagg — reach for only after the above are exhausted and prefill interference is the actual bottleneck.
Triage tree ("why is it slow")
From Red Hat's 5-step triage (2026-03-09):
| Symptom | Look at | Common cause |
|---|---|---|
| TTFT high, queue empty | compute-bound prefill | chunked-prefill budget too low, no prefix cache, bad parallelism |
| TTFT high, queue growing | capacity | raise replicas, raise max_num_seqs, check preemption rate |
| TPOT high, TTFT fine | decode-bound | MoE kernel not tuned, wrong attention backend, async sched off |
| ITL spikes | CUDA-graph miss | batch sizes fall outside captured buckets |
| Preemptions climbing | KV thrashing | raise --swap-space, lower --max-num-seqs, or add replicas |
num_running < configured concurrency | scheduler stall | check async-sched blockers, multimodal path, structured output |
DCGM signals (not GPU_UTIL): DCGM_FI_PROF_SM_OCCUPANCY, DCGM_FI_PROF_PIPE_TENSOR_ACTIVE. Low tensor-core active on a GEMM-bound workload = memory-bound.
Quick-answer router
MoE tuning + expert parallelism + DeepEP + EPLB + parallelism decision matrix → references/moe-and-ep.md
Scheduler knobs + CUDA graphs + torch.compile + compile cache → references/scheduler-and-compile.md
NCCL / InfiniBand / DCGM + PD disaggregation (Nixl/Mooncake/LMCache) → references/distributed.md
Known regressions + vendor quirks (AMD / Ascend / XPU) → references/regressions.md
Full citation anchors → references/sources.md
Top 10 operator mistakes this skill exists to prevent
-
Not running
benchmark_moe.pyon a new GPU SKU. Shipped configs cover common combos (H100, H200, A100, MI300X for Mixtral). B200 / B300 / GB200 / MI325X / Jetson Thor / RTX Pro Blackwell almost always need re-tuning. Symptom:Using default MoE config. Performance might be sub-optimal!in logs. Fix:python benchmarks/kernels/benchmark_moe.py --model <moe-model> --tp-size <N> --enable-expert-parallel --tune --save-dir ./configs, thenexport VLLM_TUNED_CONFIG_FOLDER=./configs. Expected: 20-40% throughput recovery. -
gpu_memory_utilization=0.95on MoE. All-to-all staging buffers (DeepEP, NVSHMEM) aren't in the memory accountant. OOM at high concurrency. Fix: cap at 0.85 for MoE, 0.92 for dense. -
TP not divisible by head count. Model has 32 heads, TP=7 → shape mismatch. Rule:
num_heads % TP == 0ANDhidden_size % TP == 0. -
DeepSeek-V3.2 at TP=8 on H100/H200/B200/B300. FlashMLA-Sparse only uses 16 heads per rank, padded to 64 → overhead. Fix:
DP=8, EP=8, TP=1. (DeepSeek-V3.2 recipe) -
Llama-4-Maverick with
--enable-expert-parallel. Activation density 0.78% (1/128) — AllToAll overhead exceeds parallelism win. EP hurts 7-12% vs TP-only for Maverick. DeepSeek-R1 (3.13%) and Qwen3-235B (6.25%) benefit from EP. Rule: only enable EP when(experts_per_token / total_experts) > 2%. -
MLA model with TP=8. Single KV head, TP duplicates ~84.5 GB KV cache per rank. Fix: DP-attention + EP-MoE splits KV to ~0.125 GB/GPU/request.
-
--async-schedulingwith unsupported path. Structured outputs (fixed #26866), spec-dec (#24799 fixed), PP/struct-out/spec-dec/MM umbrella tracker #27679 closed 2025-12-29 — all sub-PRs merged. The vllm-ascend v0.11.0rc2 precision bug (#4649) is fixed — closed 2026-03-13; upgrade rather than disabling async-sched. Symptom: stall, latency regression, or precision loss. Fix: upgrade; only disable if you've reproduced the issue on your version. -
Shipping
VLLM_MOE_USE_DEEP_GEMM=1blindly on H200. Betweend83f3f7and5a84b76the DeepGEMM MoE M<128 restriction was removed; H200 DeepSeek-R1 EP at concurrency ≤8 regressed 1.5× TTFT. #28882 closed 2026-04-21 — upgrade to v0.19.1+ or current main. For pre-v0.19.1 deployments, the workaround wasVLLM_MOE_USE_DEEP_GEMM=0+ FlashInfer FP8 for low-concurrency decode; re-benchmark after upgrading before removing the override. -
Skipping compile cache on K8s. First-pod torch.compile = 5-15 min on large models. Fix: pre-bake
$VLLM_CACHE_ROOT/torch_compile_cacheon one pod, mount as PVC / OCI layer. Llama-4 specifically needsVLLM_DISABLE_COMPILE_CACHE=1— stale-cache bug. -
NCCL_CUMEM_ENABLE=0on GB200. Disables multi-node NVLink, forces TCP/IB fallback. Nvidia's rule: "users should not need to tune NCCL environment variables" on modern clouds. GB200 setNCCL_NET_GDR_C2C=1. PR #16992 fixed vLLM's defaults.
Operator cheat sheet
MoE tune on new hardware (canonical recipe)
# Step 1 — run the tuner (uses Ray to parallelize across local GPUs)
python benchmarks/kernels/benchmark_moe.py \
--model deepseek-ai/DeepSeek-V3 \
--tp-size 8 --enable-expert-parallel --dtype fp8_w8a8 \
--tune --save-dir ./moe_configs
# Step 2 — point vLLM at the configs
export VLLM_TUNED_CONFIG_FOLDER=./moe_configs
# Step 3 — serve + verify no "default MoE config" warning in logs
vllm serve deepseek-ai/DeepSeek-V3 --tensor-parallel-size 8 \
--enable-expert-parallel --enable-eplb --enable-dbo \
--gpu-memory-utilization 0.85
Parallelism first-pick table
| Model family | Small-scale | Large-scale (≥16 GPUs) |
|---|---|---|
| Dense (Llama, Qwen3-dense) | TP=N, DP=replicas | TP=8 intra-node + PP=nodes OR TP=8 + DP=N |
| MoE non-MLA (Mixtral, Qwen3-MoE) | TP=N, EP off | TP + EP: EP = E / TP |
| MoE MLA (DeepSeek-V3/R1, Kimi-K2) | DP=N + EP | Wide-EP: DP-attn + EP-MoE, --enable-eplb |
| Llama-4-Maverick (0.78% density) | TP only | TP only (EP hurts) |
| DeepSeek-V3.2 (FlashMLA-Sparse) | DP=8, EP=8, TP=1 | same |
Concurrency crossover (8× MI300X benchmarks): ≤128 concurrent → TP wins, ≥512 → DP wins, 256-512 mixed. (AMD MoE playbook)
Scheduler first-pass by workload
| Scenario | max_num_batched_tokens | max_num_seqs | Other |
|---|---|---|---|
| Throughput-heavy (batch decode) | 4096-16384 | 256-512 | async sched on |
| Latency-heavy (chat) | 1024-2048 | 64-128 | async sched on, --stream-interval 1 |
| Long-context RAG | 8192-16384 | 32-64 | --enable-prefix-caching, --long-prefill-token-threshold |
| Wide-EP DeepSeek | 8192 | 256 | --enable-expert-parallel --enable-eplb --enable-dbo, FULL_AND_PIECEWISE |
Compile-level shorthand
| Flag | Effect |
|---|---|
-O0 | No compile, no CUDA graphs (= --enforce-eager) |
-O1 | Simple compile + PIECEWISE graphs |
-O2 | default — full compile + FULL_AND_PIECEWISE + fusions (AllReduce+RMSNorm +15%, SP+Async-TP +10%, Attention+Quant FP8 +7%) |
-O3 | reserved (currently = -O2) |
What changed under you, v0.22.0 → v0.25.1
Four minors of execution-path change. These move the baseline a re-tune is measured against — re-benchmark across any of these boundaries rather than comparing to numbers taken before them.
- Model Runner V2 became the default execution path, in three steps. Qwen3 (v0.22.0) → + Llama and Mistral dense models (#43458, v0.23.0) → all dense models (#44443, v0.25.0). MRv2 also gained a FlashInfer sampler (#42472), breakable CUDA graphs (#44050), pipeline-parallel bubble elimination (#42187), and full-CUDA-graph-compatible dynamic speculative decoding (#45953). A throughput figure measured on a pre-MRv2 build is not comparable to one taken after the switch for the same model.
- DeepEP v2 replaced v1 (#41183, v0.24.0), with follow-on token-bound and topk-index fixes (#46404, #46432).
- Async EPLB is on by default (#43219, v0.23.0). v0.24.0 then made NCCL-based EPLB rejected in combination with async EPLB (#44978) — a config that used to start will now fail fast.
- Sequence parallelism no longer requires DP (#47070, v0.25.0), +1.9–5.0% E2E throughput — a lever that was previously unavailable in non-DP deployments.
CUDA_VISIBLE_DEVICESis no longer set internally; use the newdevice_idsargument (#45026, v0.24.0). Device-pinning wrappers need review.- PagedAttention was removed entirely (#47361, v0.25.0).
- The Transformers modeling backend is now as fast as native vLLM (#47187, v0.25.0) — the historical "always convert to a native implementation for speed" reflex is worth re-testing.
- Notable kernel wins if hunting single-digit percentages: batch-invariant
Cutlass FP8 +28.9% E2E (#40408, v0.22.0), CutlassFP8 padding
pre-processing +13.5% TTFT (#42651), SM90 CUTLASS FP8 odd-M
swap_ab180–290% kernel (#44572, v0.24.0), reduce-scatter MoE all-reduce +3.1–3.2% E2E (#46635, v0.25.0).
Key numbers to memorize
| Metric | Value |
|---|---|
Default max_num_batched_tokens (since PR #10544) | 2048 (was 512) |
Default max_num_seqs | 256 |
| Default CUDA-graph sizes | [1,2,4] + range(8,256,8) + range(256,max,16), cap min(max_num_seqs*2, 512) |
| H200 Wide-EP DeepSeek-R1 throughput | 2.2k tok/s/GPU vs ~1.5k baseline (vllm.ai/blog/large-scale-serving) |
| GB200 Wide-EP DeepSeek-R1 | 26.2K TPGS prefill, 10.1K TPGS decode, 3-5× H200 (vllm.ai/blog/dsr1-gb200-part1) |
| MLPerf v5.1 Blackwell Ultra | 5,842 tok/s/GPU offline, 2,907 server (NVIDIA blog) |
| DeepEP dispatch (FP8, Azure H100 IB400) | 45.9 GB/s RDMA, 149.8 GB/s NVLink (Azure blog) |
| Activation density cutoff for EP win | > 2% (below: TP wins) |
Source policy
All claims cite file:line, release-note PR refs, or issue IDs. Full anchor list + vendor-specific sources in references/sources.md. Compiled 2026-04-18 against v0.19.0; freshened 2026-05-28 against v0.21.0. Last freshened 2026-07-21 against v0.25.1, covering the v0.22-v0.25 execution-path changes and a re-probe of every tracked issue.
Treat a CLOSED issue as unfixed until you read why it closed. This pass found #31475 (MI300X FP8 slower than BF16) and #25538 (preempt/resume thrashing) both closed NOT_PLANNED by the inactivity bot, and #35048 stale-marked and heading the same way — none of them fixed. Only #29539 and #34249 closed against real fixes, and #38971 closed with a usable answer (--moe-backend).
Next refresh when v0.26.x ships, or when the Wide-EP GB200 Part II blog lands (still Part I only as of 2026-07-21, not re-probed this pass).
What ships with it: 6 files
57.1 KB alongside SKILL.md
references/
- distributed.md7.5 KB
- improvement-backlog.md5.1 KB
- moe-and-ep.md12.5 KB
- regressions.md9.0 KB
- scheduler-and-compile.md8.6 KB
- sources.md14.3 KB
Gives 0 of the 12 instructions most performance cost skills give in ~4.1k tokens
Counted across 803 of the 1,058 authors here whose files we hold, read 2026-08-07
- Keep skill files under 500 lines or tokensin 82 of 803, across 16 files
- Use imperative form in instructionsin 80 of 803, across 9 files
- Draft assertions while test runs are in progressin 75 of 803, across 9 files
- Create two to three realistic test promptsin 74 of 803, across 9 files
- Write skill descriptions to be pushyin 72 of 803, across 7 files
- Save test cases to evals JSONin 72 of 803, across 6 files
- Ask questions about edge cases and input formatsin 72 of 803, across 7 files
- Save timing data immediately when runs completein 70 of 803, across 5 files
- Include all trigger conditions in the skill descriptionin 69 of 803, across 3 files
- Launch all test runs in a single turn or simultaneouslyin 69 of 803, across 3 files
- Capture intent before writing a skillin 67 of 803, across 1 file
- Import directly instead of barrel filesin 52 of 803, across 15 files
Said here and by no other author read
- characterize workload metrics before tuning
- select parallelism strategy based on model family
- run benchmark_moe.py on new gpu SKUs
- cap gpu_memory_utilization at 0.85 for MoE
- align cuda-graph-sizes with max_num_seqs doubled
- pre-bake torch_compile_cache for kubernetes pods
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.