Inference dcgm correlate
Skill cfregly/gpu-perf-tune/plugins/profile-and-optimize/skills/inference-dcgm-correlate
31 GPU inference profiling and optimization skills for Claude Code, with a bundled MCP server
npx -y skills add cfregly/gpu-perf-tune --skill inference-dcgm-correlateAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Correlate DCGM Prometheus byte-traffic counters with an inference-perf-bench sweep window to compute byte-grounded workload-level Speed-of-Light. The third tier of the SoL rigor hierarchy (after zymtrace sample-share and ncu per-kernel arithmetic intensity). Reads the sweep window from `inference_perfbench_v1.json.bench`, queries Prometheus via the Prometheus MCP for the DCGM PROF group (`DRAM_ACTIVE`, `NVLINK_TX/RX_BYTES`, `PIPE_TENSOR_ACTIVE`, `PIPE_FP16_ACTIVE`), falls back to `DCGM_FI_DEV_*` counter-tier metrics when PROF is not exported, and writes `<cell>/dcgm_correlation.json` with per-resource %SoL = real byte traffic / (peak * duration * n_gpus). Triggers on "dcgm correlate", "dcgm sol", "byte-grounded sol", "workload %sol", "dram active over sweep", "nvlink bytes", "tensor pipe active", "real-vs-peak workload bandwidth", or combinations of "dcgm / prometheus" with "sweep / window / sol / workload / byte-traffic".
SKILL.md
15.3 KB, as published. Nobody here has run it
inference-dcgm-correlate
Purpose
Lift workload-level Speed-of-Light from "first-principles estimate" to
"measured byte traffic over the sweep window". The methodology canon
(docs/METHODOLOGY.md "Speed-of-light framing") names three levels of SoL rigor:
Steady-state window: per-(c) DCGM correlation is only meaningful when the bench cell sustained steady state, i.e.
num_prompts >= 2*c. Atnum=c+4the window is ramp/drain- dominated so BOTH the throughput AND the high-c DCGM utilization read low (seedocs/METHODOLOGY.md"Capture hygiene").import_roofline_sweepWARNs on any cell< 2c.
- Sample-share proxy (page 4 in the perf-report PDF) - zymtrace
per-category time-share read as a coarse upper bound on category
busyness. (zymtrace flushes to ClickHouse asynchronously, so an empty L1
right after the window is ingest lag, not absence - wait + requery for
the freshest data. See
server/docs/zymtrace-query-hygiene.md.) - ncu per-kernel arithmetic intensity (page 5) - proper roofline scatter from ncu DRAM bytes + SM FLOPS counters.
- DCGM workload-level byte traffic (page 6) - this skill. Real GB transferred across NVLink / HBM / Tensor pipe during the drive_load sweep, divided by peak × duration × n_gpus.
This skill is the page-6 producer.
When to use
- After a campaign's
drive_load.pysweep completes and the bundle has aninference_perfbench_v1.json.bench.captured_at + duration_effective_spair recording the sweep window. - When you want a workload-level %SoL number anchored in measured byte counters, not a first-principles HBM-roofline estimate.
- When refreshing a campaign's
sol-summary.mdwith byte-grounded numbers replacing the time-share proxies.
Do not use for:
- Per-kernel arithmetic-intensity questions - that's
inference-kernel-ncu-profile's domain (DCGM has no per-kernel attribution). - Real-time cluster health - DCGM scrape interval (~10-30 s) is too
coarse for sub-minute incident response. Use
prometheus-anchored-querydirectly with the regular dashboard panels.
Prerequisites
- The campaign / cell directory exists at
campaigns/<campaign>/cells/<cell>/. - The source
inference_perfbench_v1.jsonis reachable (either inside the cell dir or passed via--bundle-pathexplicitly). - Prometheus is reachable via the Prometheus MCP server
(
prometheus_mcp). Thequery_observability_knowledge_basetool is used FIRST to confirm the DCGM metrics exist with the expected cardinality. - The cluster's DCGM exporter is configured to export the
DCGM_FI_PROF_*group. If not, the skill falls back toDCGM_FI_DEV_*counter-tier and flags the result withdcgm_group_level: "counter".
Workflow
Phase 0 - pre-flight (knowledge-base probe)
from tools.perf_tune_report.dcgm_correlate import (
DcgmCorrelateInputs,
correlate,
read_sweep_window_from_bundle,
)
- Resolve the bundle path + cell directory.
- Load
configs/sol-ceilings.yaml. On a GB300 cluster usehw_key=gb300_nvl72,n_gpus=the deploy's TP (GB300 node = 4, NOT 8), and the tensorpeak_key=nvfp4_dense_pflopsfor NVFP4 weights (bf16_dense_pflops/fp8_dense_pflopsotherwise). Stamp these from the deploy, not from habit: theb200_sm100/n_gpus: 8/bf16defaults apply only to B200 clusters and will mis-scale a GB300/NVFP4 %SoL if reused. - Call
query_observability_knowledge_basefor the metrics listed indcgm_config.prof_group_probe_metrics. Confirm:- Each metric exists on the target cluster.
- Labels include at least
{namespace, pod, gpu, device}(the default DCGM label set). - Cardinality is bounded (e.g. < 1000 series per metric across the target deploy).
Phase 1 - sweep window
Read (start_utc, end_utc) either by calling
read_sweep_window_from_bundle(bundle_path) or from explicit
input (when the bundle's captured_at is the END of the sweep, not the
start - older bundles).
Phase 2 - build queries (dry-run)
inputs = DcgmCorrelateInputs(
bundle_path=bundle,
cell_dir=cell,
sweep_start=start,
sweep_end=end,
hw_key="b200_sm100", # GB300: "gb300_nvl72"
pod_label_selector="app=basic-inference",
namespace="inference",
expected_n_gpus=8, # GB300 node = 4 (the deploy TP), NOT 8
)
# Dry-run first to print the PromQL the correlator WILL fire:
result = correlate(inputs, ceilings, prom_client, dry_run=True)
for q in result.queries:
print(q["peak_key"], "->", q["promql"])
Phase 3 - execute the correlation
result = correlate(inputs, ceilings, prom_client, dry_run=False)
out_path = write_correlation(result, inputs.cell_dir)
The result's resources list has one row per peak that mapped to a
DCGM metric, each with:
measured_bytes_total,measured_bytes_per_s(bandwidth peaks) ormeasured_tflops_avg(compute peaks)sol_pct(the headline number)notes[](short-sweep / missing-data flags)
Phase 4 - patch sol-summary.md
Update the campaign's sol-summary.md "Workload-level SoL" table to
reference the byte-grounded numbers, replacing the previous
first-principles estimate. Cite the dcgm_correlation.json path so
future readers can re-derive the math.
Phase 5 - re-render + re-publish to raise sol_rigor to L3
Emitting dcgm_correlation.json per cell is not the end - the campaign's
published lake row + report PDF only reflect the byte-grounding after a
re-render then re-publish:
perftunereport report_render --campaign <slug> # draws pages 6 + 6b; sets dcgm_grounded + sol_rigor=L3
perftunereport publish_to_lake --campaign <slug> --if-exists overwrite
Byte-grounding RAISES sol_rigor to L3 - it is RECORDED, not a gate
(always-publish policy). A sol_complete=true campaign that is
dcgm_grounded=false (no dcgm_correlation.json, pages 6/6b absent) still
publishes at sol_rigor=L1 (zymtrace proxy) - the gap is RECORDED on the
campaign_v1 row + warned, never a refusal. dcgm_grounded + sol_rigor flow
report_status.json -> report_render envelope -> campaign_v1 columns. Run
this skill (or the CLI verb below) for every plot-ready cell, then
re-render + re-publish so the lake row is dcgm_grounded=true / sol_rigor=L3
- a tighter roofline. Pass
publish_to_lake --strictonly when you want andcgm_grounded=falsecampaign to refuse instead of land.
Phase 5b - offline / CI path: the dcgm_correlate CLI verb + frozen snapshot
The live correlate() python path above needs a PrometheusClient (the
agent wires it to the Prometheus MCP). For an offline / re-runnable /
CI context that cannot reach Prometheus, capture the DCGM means into a
frozen YAML (schema dcgm_frozen_v1) once, then fold it in
deterministically with the CLI verb:
perftunereport dcgm_correlate --campaign <slug> --cell-id <cell> \
--frozen-yaml <cell>/dcgm-frozen.yaml \
[--kernels-json <cell>/kernels.json] # default: the cell's own kernels.json (page 6b)
This wraps correlate_from_frozen + write_correlation and is the path the
campaign orchestrator's step_dcgm_correlate runs (it consumes a
cells/<id>/dcgm-frozen.yaml per cell). Always snapshot a frozen YAML even
when you used the live path, so the byte-grounding is reproducible offline
and survives deploy teardown (the DCGM time-series may age out of Prometheus
retention, but the frozen means do not).
Full-context reporting (no bare numbers)
Per the canon "Every performance number carries its full context (no bare numbers)"
(docs/METHODOLOGY.md "Full-context reporting"): every number this
skill emits MUST carry its full measurement-context descriptor, and every comparison MUST be
matched on it. A bare tok/s / TPOT / BW / %SoL / speedup is a defect - it cannot set a
default, ship a config, or appear in a report.
- Identity: model (+HF path), hardware (exact ceiling token
GB300/B200), quant, kv-cache dtype. - Parallelism: TP, DP (replicas), PP, EP, parallel_strategy.
- Serving cfg: max-num-seqs, max-num-batched-tokens, gpu-memory-utilization, max-model-len, cudagraph_mode/enforce_eager, async_scheduling, prefix-caching.
- Workload: dataset, ISL/OSL (or mean in/out tokens), concurrency, num-prompts.
- Regime: warm vs cold. Latency vs throughput tier.
- Stack: image/vllm commit, bench backend, serving engine.
- Grounding:
%SoL(+ ceiling key fromconfigs/sol-ceilings.yaml- never inline a peak), sol_rigor (L1-L4), trials n (mean±std), same-node, baseline named. - Per-number exact shape (no smoothing): when reporting more than one number, keep EACH with its own exact shape (ISL/OSL, concurrency, dataset, regime) - never normalize a set to one uniform descriptor that hides per-point variation (e.g.
c=1 @ ISL1024/OSL256+c=64 @ ISL4096/OSL512, NOT one shared "random").
This skill IS the byte-grounded SoL producer. Its output is the
authoritative third-tier evidence per docs/METHODOLOGY.md
"Speed-of-light framing":
- Peaks live in
configs/sol-ceilings.yaml. This skill reads them by key (b200_sm100.hbm3e_tbps,gb300_nvl72.nvfp4_dense_pflops, etc.) - never inline. - DCGM metric anchors live in the same YAML under each peak's
dcgm_metric/dcgm_metrics_bytes/dcgm_fallback_metricfields. - The renderer's page 6 (
dcgm_sol.py) consumes the emitteddcgm_correlation.jsonand draws workload-level resource bars showing measured-vs-peak × duration × n_gpus. - This skill's output drives the
dcgm_groundedflag + the campaign'ssol_rigor: with adcgm_correlation.jsonthe campaign isdcgm_grounded=true/sol_rigor=L3(orL4if ncu is also present), without one it isdcgm_grounded=false/sol_rigor=L1. Under the always-publish policy this is RECORDED, not a gate -publish_to_lakelands the campaign either way, with the gap on thecampaign_v1row + a loud warning. Run this skill per cell to raise rigor to L3 (a tighter roofline), passpublish_to_lake --strictonly when you want an ungrounded campaign to refuse instead of land.
Next lever / BREAKTHROUGH (Grind Mandate)
If this skill emits a measured result, its output MUST end by naming the next perf lever,
its expected unlock (direction + rough magnitude), and the gate that proves/refutes it,
per docs/METHODOLOGY.md "Always be grinding (next-lever framing)". A
measured win is the new floor, not the finish -- so do everything we can to find the next
BREAKTHROUGH: the highest-EV unlock toward Speed-of-Light (a new champion / kernel / router /
quant / parallelism / spec-decode win, or an unblocked stack), not just the next micro-lever.
Rank the candidate breakthrough levers by value x cost (the GRIND FRONTIER, perftunereport value_view), pursue the top, bank the rest with evidence. Record WHY a refuted lever loses,
update the standing frontier in the active bundle's HANDOFF.md. Never conclude
"exhausted/optimal/done" without an explicit next-lever frontier (an empty frontier AND a
documented SoL wall only). Delete this section ONLY if the skill produces no measurements.
Safety
- Read-only. Every Prometheus call is a read. No mutation of any cluster state.
- Knowledge-base FIRST. Skill MUST call
query_observability_knowledge_basebefore anyquery_prometheusto confirm cardinality bounds. This is the standing prometheus-anchored-query pattern. - Bundle write only inside the supplied cell directory.
No writes outside
<campaign>/cells/<cell>/. - Provenance preserved. Every PromQL invocation is recorded in the
output's
queriesarray. Thedcgm_correlation.jsoncarriesschema_version,sweep_start_utc,sweep_end_utc,n_gpus,dcgm_group_level. It also carries the re-query provenancenodes(distinct host(s) the DCGM series carried),namespace, andpod_label_selector.DCGM_FI_DEV_POWER_USAGEis per-node, so without the node a window-only capture cannot be re-queried later fortokens_per_watt- the livecorrelate()now auto-captures the node from the series labels (Hostname/node/exported_node/instance), and a frozen YAML SHOULD recordnodes:(+ optionalnamespace:/pod_label_selector:) so the byte-grounding stays re-queryable after the time-series ages out.
Pairs with
inference-kernel-ncu-profile- per-kernel arithmetic intensity. Run BOTH for a complete SoL picture: ncu surfaces the per-kernel %SoL, this skill surfaces the workload-level %SoL aggregated over the same window.
inference-perf-bench- the drive_load.py sweep that produces the window this skill correlates against.prometheus-anchored-query- the general anchored-PromQL primitive. This skill is the DCGM-specific specialisation.
analyze-zymtrace-workload- the time-share proxy view (level 1 of SoL hierarchy). This skill is its level-3 upgrade.
Source-of-truth references
- Tool:
server/tools/perf_tune_report/dcgm_correlate.py. - Tests:
server/tools/perf_tune_report/test_dcgm_correlate.py - fake-Prometheus-client coverage of ratio/byte-rate aggregation, PROF/counter/absent fallback, and short-sweep warning.
- Renderer page 6:
server/tools/perf_tune_report/renderer/dcgm_sol.py(consumes the emitteddcgm_correlation.json). docs/METHODOLOGY.md"Speed-of-light framing" - the standing three-level rigor hierarchy this skill operationalises.
Contact
Open an issue on this repository.