Inference dcgm correlate
Skill cfregly/gpu-perf-tune/plugins/profile-and-optimize/skills/inference-dcgm-correlate
Correlate DCGM Prometheus byte-traffic counters with an inference-perf-bench sweep window to compute byte-grounded workload-level Speed-of-Light. The third tier of the SoL rigor hierarchy (after zymtrace sample-share and ncu per-kernel arithmetic intensity). Reads the sweep window from `inference_perfbench_v1.json.bench`, queries Prometheus via the Prometheus MCP for the DCGM PROF group (`DRAM_ACTIVE`, `NVLINK_TX/RX_BYTES`, `PIPE_TENSOR_ACTIVE`, `PIPE_FP16_ACTIVE`), falls back to `DCGM_FI_DEV_*` counter-tier metrics when PROF is not exported, and writes `<cell>/dcgm_correlation.json` with per-resource %SoL = real byte traffic / (peak * duration * n_gpus). Triggers on "dcgm correlate", "dcgm sol", "byte-grounded sol", "workload %sol", "dram active over sweep", "nvlink bytes", "tensor pipe active", "real-vs-peak workload bandwidth", or combinations of "dcgm / prometheus" with "sweep / window / sol / workload / byte-traffic".From its SKILL.md
npx -y skills add cfregly/gpu-perf-tune --skill inference-dcgm-correlateAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
15.3 KB, ~3.9k tokens by cl100k_base, as published. Nobody here has run it
inference-dcgm-correlate
Purpose
Lift workload-level Speed-of-Light from "first-principles estimate" to
"measured byte traffic over the sweep window". The methodology canon
(docs/METHODOLOGY.md "Speed-of-light framing") names three levels of SoL rigor:
Steady-state window: per-(c) DCGM correlation is only meaningful when the bench cell sustained steady state, i.e.
num_prompts >= 2*c. Atnum=c+4the window is ramp/drain- dominated so BOTH the throughput AND the high-c DCGM utilization read low (seedocs/METHODOLOGY.md"Capture hygiene").import_roofline_sweepWARNs on any cell< 2c.
- Sample-share proxy (page 4 in the perf-report PDF) - zymtrace
per-category time-share read as a coarse upper bound on category
busyness. (zymtrace flushes to ClickHouse asynchronously, so an empty L1
right after the window is ingest lag, not absence - wait + requery for
the freshest data. See
server/docs/zymtrace-query-hygiene.md.) - ncu per-kernel arithmetic intensity (page 5) - proper roofline scatter from ncu DRAM bytes + SM FLOPS counters.
- DCGM workload-level byte traffic (page 6) - this skill. Real GB transferred across NVLink / HBM / Tensor pipe during the drive_load sweep, divided by peak × duration × n_gpus.
This skill is the page-6 producer.
When to use
- After a campaign's
drive_load.pysweep completes and the bundle has aninference_perfbench_v1.json.bench.captured_at + duration_effective_spair recording the sweep window. - When you want a workload-level %SoL number anchored in measured byte counters, not a first-principles HBM-roofline estimate.
- When refreshing a campaign's
sol-summary.mdwith byte-grounded numbers replacing the time-share proxies.
Do not use for:
- Per-kernel arithmetic-intensity questions - that's
inference-kernel-ncu-profile's domain (DCGM has no per-kernel attribution). - Real-time cluster health - DCGM scrape interval (~10-30 s) is too
coarse for sub-minute incident response. Use
prometheus-anchored-querydirectly with the regular dashboard panels.
Prerequisites
- The campaign / cell directory exists at
campaigns/<campaign>/cells/<cell>/. - The source
inference_perfbench_v1.jsonis reachable (either inside the cell dir or passed via--bundle-pathexplicitly). - Prometheus is reachable via the Prometheus MCP server
(
prometheus_mcp). Thequery_observability_knowledge_basetool is used FIRST to confirm the DCGM metrics exist with the expected cardinality. - The cluster's DCGM exporter is configured to export the
DCGM_FI_PROF_*group. If not, the skill falls back toDCGM_FI_DEV_*counter-tier and flags the result withdcgm_group_level: "counter".
Workflow
Phase 0 - pre-flight (knowledge-base probe)
from tools.perf_tune_report.dcgm_correlate import (
DcgmCorrelateInputs,
correlate,
read_sweep_window_from_bundle,
)
- Resolve the bundle path + cell directory.
- Load
configs/sol-ceilings.yaml. On a GB300 cluster usehw_key=gb300_nvl72,n_gpus=the deploy's TP (GB300 node = 4, NOT 8), and the tensorpeak_key=nvfp4_dense_pflopsfor NVFP4 weights (bf16_dense_pflops/fp8_dense_pflopsotherwise). Stamp these from the deploy, not from habit: theb200_sm100/n_gpus: 8/bf16defaults apply only to B200 clusters and will mis-scale a GB300/NVFP4 %SoL if reused. - Call
query_observability_knowledge_basefor the metrics listed indcgm_config.prof_group_probe_metrics. Confirm:- Each metric exists on the target cluster.
- Labels include at least
{namespace, pod, gpu, device}(the default DCGM label set). - Cardinality is bounded (e.g. < 1000 series per metric across the target deploy).
Phase 1 - sweep window
Read (start_utc, end_utc) either by calling
read_sweep_window_from_bundle(bundle_path) or from explicit
input (when the bundle's captured_at is the END of the sweep, not the
start - older bundles).
Phase 2 - build queries (dry-run)
inputs = DcgmCorrelateInputs(
bundle_path=bundle,
cell_dir=cell,
sweep_start=start,
sweep_end=end,
hw_key="b200_sm100", # GB300: "gb300_nvl72"
pod_label_selector="app=basic-inference",
namespace="inference",
expected_n_gpus=8, # GB300 node = 4 (the deploy TP), NOT 8
)
# Dry-run first to print the PromQL the correlator WILL fire:
result = correlate(inputs, ceilings, prom_client, dry_run=True)
for q in result.queries:
print(q["peak_key"], "->", q["promql"])
Phase 3 - execute the correlation
result = correlate(inputs, ceilings, prom_client, dry_run=False)
out_path = write_correlation(result, inputs.cell_dir)
The result's resources list has one row per peak that mapped to a
DCGM metric, each with:
measured_bytes_total,measured_bytes_per_s(bandwidth peaks) ormeasured_tflops_avg(compute peaks)sol_pct(the headline number)notes[](short-sweep / missing-data flags)
Phase 4 - patch sol-summary.md
Update the campaign's sol-summary.md "Workload-level SoL" table to
reference the byte-grounded numbers, replacing the previous
first-principles estimate. Cite the dcgm_correlation.json path so
future readers can re-derive the math.
Phase 5 - re-render + re-publish to raise sol_rigor to L3
Emitting dcgm_correlation.json per cell is not the end - the campaign's
published lake row + report PDF only reflect the byte-grounding after a
re-render then re-publish:
perftunereport report_render --campaign <slug> # draws pages 6 + 6b; sets dcgm_grounded + sol_rigor=L3
perftunereport publish_to_lake --campaign <slug> --if-exists overwrite
Byte-grounding RAISES sol_rigor to L3 - it is RECORDED, not a gate
(always-publish policy). A sol_complete=true campaign that is
dcgm_grounded=false (no dcgm_correlation.json, pages 6/6b absent) still
publishes at sol_rigor=L1 (zymtrace proxy) - the gap is RECORDED on the
campaign_v1 row + warned, never a refusal. dcgm_grounded + sol_rigor flow
report_status.json -> report_render envelope -> campaign_v1 columns. Run
this skill (or the CLI verb below) for every plot-ready cell, then
re-render + re-publish so the lake row is dcgm_grounded=true / sol_rigor=L3
- a tighter roofline. Pass
publish_to_lake --strictonly when you want andcgm_grounded=falsecampaign to refuse instead of land.
Phase 5b - offline / CI path: the dcgm_correlate CLI verb + frozen snapshot
The live correlate() python path above needs a PrometheusClient (the
agent wires it to the Prometheus MCP). For an offline / re-runnable /
CI context that cannot reach Prometheus, capture the DCGM means into a
frozen YAML (schema dcgm_frozen_v1) once, then fold it in
deterministically with the CLI verb:
perftunereport dcgm_correlate --campaign <slug> --cell-id <cell> \
--frozen-yaml <cell>/dcgm-frozen.yaml \
[--kernels-json <cell>/kernels.json] # default: the cell's own kernels.json (page 6b)
This wraps correlate_from_frozen + write_correlation and is the path the
campaign orchestrator's step_dcgm_correlate runs (it consumes a
cells/<id>/dcgm-frozen.yaml per cell). Always snapshot a frozen YAML even
when you used the live path, so the byte-grounding is reproducible offline
and survives deploy teardown (the DCGM time-series may age out of Prometheus
retention, but the frozen means do not).
Full-context reporting (no bare numbers)
Per the canon "Every performance number carries its full context (no bare numbers)"
(docs/METHODOLOGY.md "Full-context reporting"): every number this
skill emits MUST carry its full measurement-context descriptor, and every comparison MUST be
matched on it. A bare tok/s / TPOT / BW / %SoL / speedup is a defect - it cannot set a
default, ship a config, or appear in a report.
- Identity: model (+HF path), hardware (exact ceiling token
GB300/B200), quant, kv-cache dtype. - Parallelism: TP, DP (replicas), PP, EP, parallel_strategy.
- Serving cfg: max-num-seqs, max-num-batched-tokens, gpu-memory-utilization, max-model-len, cudagraph_mode/enforce_eager, async_scheduling, prefix-caching.
- Workload: dataset, ISL/OSL (or mean in/out tokens), concurrency, num-prompts.
- Regime: warm vs cold. Latency vs throughput tier.
- Stack: image/vllm commit, bench backend, serving engine.
- Grounding:
%SoL(+ ceiling key fromconfigs/sol-ceilings.yaml- never inline a peak), sol_rigor (L1-L4), trials n (mean±std), same-node, baseline named. - Per-number exact shape (no smoothing): when reporting more than one number, keep EACH with its own exact shape (ISL/OSL, concurrency, dataset, regime) - never normalize a set to one uniform descriptor that hides per-point variation (e.g.
c=1 @ ISL1024/OSL256+c=64 @ ISL4096/OSL512, NOT one shared "random").
This skill IS the byte-grounded SoL producer. Its output is the
authoritative third-tier evidence per docs/METHODOLOGY.md
"Speed-of-light framing":
- Peaks live in
configs/sol-ceilings.yaml. This skill reads them by key (b200_sm100.hbm3e_tbps,gb300_nvl72.nvfp4_dense_pflops, etc.) - never inline. - DCGM metric anchors live in the same YAML under each peak's
dcgm_metric/dcgm_metrics_bytes/dcgm_fallback_metricfields. - The renderer's page 6 (
dcgm_sol.py) consumes the emitteddcgm_correlation.jsonand draws workload-level resource bars showing measured-vs-peak × duration × n_gpus. - This skill's output drives the
dcgm_groundedflag + the campaign'ssol_rigor: with adcgm_correlation.jsonthe campaign isdcgm_grounded=true/sol_rigor=L3(orL4if ncu is also present), without one it isdcgm_grounded=false/sol_rigor=L1. Under the always-publish policy this is RECORDED, not a gate -publish_to_lakelands the campaign either way, with the gap on thecampaign_v1row + a loud warning. Run this skill per cell to raise rigor to L3 (a tighter roofline), passpublish_to_lake --strictonly when you want an ungrounded campaign to refuse instead of land.
Next lever / BREAKTHROUGH (Grind Mandate)
If this skill emits a measured result, its output MUST end by naming the next perf lever,
its expected unlock (direction + rough magnitude), and the gate that proves/refutes it,
per docs/METHODOLOGY.md "Always be grinding (next-lever framing)". A
measured win is the new floor, not the finish -- so do everything we can to find the next
BREAKTHROUGH: the highest-EV unlock toward Speed-of-Light (a new champion / kernel / router /
quant / parallelism / spec-decode win, or an unblocked stack), not just the next micro-lever.
Rank the candidate breakthrough levers by value x cost (the GRIND FRONTIER, perftunereport value_view), pursue the top, bank the rest with evidence. Record WHY a refuted lever loses,
update the standing frontier in the active bundle's HANDOFF.md. Never conclude
"exhausted/optimal/done" without an explicit next-lever frontier (an empty frontier AND a
documented SoL wall only). Delete this section ONLY if the skill produces no measurements.
Safety
- Read-only. Every Prometheus call is a read. No mutation of any cluster state.
- Knowledge-base FIRST. Skill MUST call
query_observability_knowledge_basebefore anyquery_prometheusto confirm cardinality bounds. This is the standing prometheus-anchored-query pattern. - Bundle write only inside the supplied cell directory.
No writes outside
<campaign>/cells/<cell>/. - Provenance preserved. Every PromQL invocation is recorded in the
output's
queriesarray. Thedcgm_correlation.jsoncarriesschema_version,sweep_start_utc,sweep_end_utc,n_gpus,dcgm_group_level. It also carries the re-query provenancenodes(distinct host(s) the DCGM series carried),namespace, andpod_label_selector.DCGM_FI_DEV_POWER_USAGEis per-node, so without the node a window-only capture cannot be re-queried later fortokens_per_watt- the livecorrelate()now auto-captures the node from the series labels (Hostname/node/exported_node/instance), and a frozen YAML SHOULD recordnodes:(+ optionalnamespace:/pod_label_selector:) so the byte-grounding stays re-queryable after the time-series ages out.
Pairs with
inference-kernel-ncu-profile- per-kernel arithmetic intensity. Run BOTH for a complete SoL picture: ncu surfaces the per-kernel %SoL, this skill surfaces the workload-level %SoL aggregated over the same window.
inference-perf-bench- the drive_load.py sweep that produces the window this skill correlates against.prometheus-anchored-query- the general anchored-PromQL primitive. This skill is the DCGM-specific specialisation.
analyze-zymtrace-workload- the time-share proxy view (level 1 of SoL hierarchy). This skill is its level-3 upgrade.
Source-of-truth references
- Tool:
server/tools/perf_tune_report/dcgm_correlate.py. - Tests:
server/tools/perf_tune_report/test_dcgm_correlate.py - fake-Prometheus-client coverage of ratio/byte-rate aggregation, PROF/counter/absent fallback, and short-sweep warning.
- Renderer page 6:
server/tools/perf_tune_report/renderer/dcgm_sol.py(consumes the emitteddcgm_correlation.json). docs/METHODOLOGY.md"Speed-of-light framing" - the standing three-level rigor hierarchy this skill operationalises.
Contact
Open an issue on this repository.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.