agentsclimarketplace

Rust perf

Skill decebal/curated-claude-skills/skills/stack-agnostic/rust-perf

Twelve Claude skills, curated by exclusion. Each does what a one-line prompt can't — memory, proof, or a real binary.

Install
npx -y skills add decebal/curated-claude-skills --skill rust-perf

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Detect and fix performance issues in Rust projects using a modern 2026 toolchain. Use when the user asks to "profile rust", "optimize rust binary", "make it faster", "flamegraph", "pgo", "bolt", "find hotspots", "why is this slow", or mentions latency/throughput regressions in Rust code. Adapts to what's already in the repo (hotpath, criterion, custom allocator) before suggesting new tools. Covers CPU sampling, in-process pprof, deterministic benchmarks, heap profiling, causal profiling, and the PGO+BOLT pipeline.

SKILL.md

12.7 KB, as published. Nobody here has run it

Rust Performance Skill

A phased playbook for diagnosing and fixing Rust perf issues. Each phase answers a different question; do not skip ahead. Never run PGO/BOLT until a hotspot is confirmed and a representative workload exists — otherwise you optimize for the wrong path.

Phase 0 — Repo detection (always run first)

Before suggesting any tool, grep the repo for what's already wired up. Adapt recommendations to what exists.

DetectCommandIf present
hotpath instrumentationgrep -l 'hotpath' Cargo.toml crates/*/Cargo.toml apps/*/Cargo.toml 2>/dev/nullStart triage with --features hotpath. See Phase 1a.
Custom allocatorgrep -r '#\[global_allocator\]' src/ crates/ apps/ 2>/dev/nullSkip the allocator swap recommendation.
criterion benchesCheck for [[bench]] in Cargo.toml, benches/ dirPropose iai-callgrind alongside it for CI gates, not as replacement.
tracing in use`grep -r 'tracing::' --include='*.rs' -lhead -1`
HTTP serverAxum/actix/hyper importsOffer pprof-rs endpoint for in-prod profiling.
CI perf gates.github/workflows/*perf*.yml or *profile*.ymlExtend the existing gate rather than making a new one.
Build runnertest -f Makefile || test -f Taskfile.yml || test -f justfileUse the matching template in Phase 6a (makefile-perf.mk / taskfile-perf.yml / justfile-perf.just). Don't add a second runner.

State what you found in one line before moving on: "Detected: hotpath (8 crates), criterion benches, no custom allocator, no pprof endpoint."

Phase 1 — Triage (which layer is slow?)

Ask the user one question before profiling: "What workload should reproduce the issue?" A command, a benchmark name, or a concrete request pattern. If they can't answer, stop — you cannot profile an abstraction.

1a. If hotpath is wired up

Run with hotpath first; it's zero-overhead when disabled and already instruments the project's known hot paths.

cargo run -p <app> --release --features hotpath
# or for p95/p99 with allocation tracking:
cargo run -p <app> --release --features hotpath-alloc

Read the report. If the hot function is in the instrumented list, jump to Phase 3 (measure) or Phase 4 (heap) depending on whether the issue is time or allocations.

If the top frame in the workload is not instrumented — that's a signal. Drop to Phase 2 sampling.

1b. If hotpath is not wired up

Skip to Phase 2. Do not propose adding hotpath purely for one-off investigation; it only pays off with CI gates.

Phase 2 — Sample (find uninstrumented hotspots)

Pick ONE sampler based on the environment. Do not run all three.

2a. Local dev (default)

samply — cross-platform (macOS/Linux/Windows), no root, opens Firefox Profiler UI.

cargo install samply
cargo build --release
samply record ./target/release/<binary> <workload-args>

cargo flamegraph is the blog-standard choice but Linux-only and needs perf perms. Prefer samply unless the user explicitly wants SVG output for a PR.

2b. Production / long-running server

pprof-rs — in-process sampler, no restart, expose behind a debug-gated endpoint. See template templates/pprof-endpoint-axum.rs for a drop-in axum handler that returns pprof protobuf (consumable by go tool pprof and pprof.me).

[features]
profiling = ["dep:pprof"]

[dependencies]
pprof = { version = "0.14", default-features = false, features = ["flamegraph", "prost-codec", "cpp"], optional = true }
// behind a feature flag or admin auth
#[cfg(feature = "profiling")]
use pprof::protos::Message as _;

#[cfg(feature = "profiling")]
async fn profile_handler(Query(p): Query<ProfileParams>) -> impl IntoResponse {
    let guard = pprof::ProfilerGuardBuilder::default()
        .frequency(100)
        .blocklist(&["libc", "libgcc", "pthread", "vdso"])
        .build()
        .unwrap();
    tokio::time::sleep(Duration::from_secs(p.seconds.unwrap_or(30))).await;
    let report = guard.report().build().unwrap();
    let profile = report.pprof().unwrap();
    let mut body = Vec::new();
    profile.encode(&mut body).unwrap();
    ([(header::CONTENT_TYPE, "application/octet-stream")], body)
}

Gate on a feature flag or admin auth — never expose unauthenticated. Use prost-codec, not protobuf-codec — the method signatures differ.

2c. Free signal if tracing is already used

tracing-flame — converts existing tracing spans to flamegraph without new instrumentation.

use tracing_flame::FlameLayer;
let (flame_layer, _guard) = FlameLayer::with_file("./tracing.folded").unwrap();
tracing_subscriber::registry().with(flame_layer).init();

Then inferno-flamegraph < tracing.folded > flame.svg.

Follow-up after sampling

When the top frame in a flamegraph is not in hotpath's instrumented list (check docs/guides/HOTPATH_PROFILING.md or similar), propose adding #[hotpath::measure] to that function so CI catches future regressions. Sampling finds it once; hotpath keeps it found.

Phase 3 — Measure (deterministic baselines before optimizing)

Flamegraphs are directional, not quantitative. Before any optimization, establish a measurable baseline.

ToolUse whenTemplate
divanNew microbenchmarks. ~3× faster than criterion, built-in allocation counters, nicer output.templates/bench-divan.rs
iai-callgrindCI regression gate. Deterministic instruction counts (zero variance) — catches sub-1% regressions criterion misses.templates/bench-iai-callgrind.rs
criterion-perf-eventsKeeping existing criterion but want cache-miss / branch-mispredict counters alongside wall time.
criterionOnly if already in the repo and the team is invested. Don't add it fresh.

For CI perf gates, iai-callgrind is strictly better than criterion because criterion's variance floor (~3%) makes small regressions undetectable. Wire it into the existing perf workflow rather than a parallel one — see templates/ci-perf-regression.yml for a reusable GitHub Actions workflow that runs iai-callgrind and fails on configurable regression thresholds.

Phase 4 — Heap / allocations (the blog skips this; often the biggest win)

Allocation pressure commonly dominates CPU time in server workloads. Check here before PGO.

4a. Swap the global allocator (5-minute change, frequently 10-20% wins)

If the repo has no #[global_allocator], propose one. Templates: templates/allocator-mimalloc.rs (default) and templates/allocator-jemalloc.rs (long-running server, high-churn workloads).

[dependencies]
mimalloc = { version = "0.1", default-features = false }
#[global_allocator]
static GLOBAL: mimalloc::MiMalloc = mimalloc::MiMalloc;

mimalloc is the default recommendation. jemallocator is solid on Linux but inferior on macOS. Measure before/after with the same workload from Phase 1.

Do this first — before PGO, before BOLT, before restructuring code. It's the highest-ROI change in the entire playbook for allocation-heavy workloads.

4b. Find which call sites allocate

dhat-rs — Valgrind DHAT-compatible profiler as a crate. No Valgrind required. Template: templates/dhat-heap.rs.

#[cfg(feature = "dhat-heap")]
#[global_allocator]
static ALLOC: dhat::Alloc = dhat::Alloc;

fn main() {
    #[cfg(feature = "dhat-heap")]
    let _profiler = dhat::Profiler::new_heap();
    // ... workload
}

Produces dhat-heap.json; open in DHAT Viewer. Look for unnecessary .clone(), Vec that could be SmallVec/stack arrays, repeated String allocations in hot loops, Arc::clone on things that could be &.

4c. Production heap (Linux only)

bytehound — timeline view, low overhead, suitable for long-running prod services.

Phase 5 — Causal profiling (optional, highest-leverage question)

A flamegraph answers where is time spent. A causal profiler answers where would optimization actually matter — those are different questions.

coz-rs wraps coz. Useful when the obvious hotspot has been optimized and further wins are unclear.

use coz;
fn hot_function() {
    coz::scope!("hot_function");
    // ...
}

Run under coz run --- ./target/release/<bin>. Output ranks functions by speedup impact on total runtime, not self-time. Often surprising.

Skip this phase for simple cases. Reach for it when flamegraphs have become unhelpful.

Phase 6 — Optimize (only after Phases 1-4)

Preconditions before this phase:

  • ✅ A confirmed hotspot (not a hunch)
  • ✅ A reproducible workload
  • ✅ A deterministic baseline (iai-callgrind or divan)
  • ✅ Allocator already chosen (Phase 4a done)

Skipping any of these means optimizing blindly.

6a. PGO + BOLT via cargo-pgo

Do not hand-roll -Cprofile-generate / -Cprofile-use / llvm-profdata / perf2bolt. Use the wrapper. Pick the runner already in the repo — templates available for all three:

RunnerTemplateDetect by
maketemplates/makefile-perf.mkMakefile exists
task (go-task)templates/taskfile-perf.ymlTaskfile.yml exists
justtemplates/justfile-perf.justjustfile exists

All three expose the same 7 targets (perf-baseline, perf-alloc-check, perf-flamegraph, perf-heap, perf-pgo, perf-bolt, perf-regression) parameterized on PACKAGE, BENCH, WORKLOAD, THRESHOLD. Don't add a second runner just to use the template.

cargo install cargo-pgo
rustup component add llvm-tools-preview

# Make (most common in existing repos):
make PACKAGE=allsource-core WORKLOAD="./target/release/core --ingest-bench" perf-pgo
make PACKAGE=allsource-core WORKLOAD="./target/release/core --ingest-bench" perf-bolt

# Task (go-task):
task perf:pgo PACKAGE=allsource-core WORKLOAD="./target/release/core --ingest-bench"

# Just:
just package=allsource-core workload="./target/release/core --ingest-bench" perf-pgo

# Raw cargo-pgo commands (reference):
cargo pgo build
cargo pgo run -- <representative-workload>
cargo pgo optimize build

cargo pgo bolt build --with-pgo
cargo pgo bolt run --with-pgo -- <representative-workload>
cargo pgo bolt optimize --with-pgo

Typical wins on a hot server binary: PGO 5-15%, BOLT an additional 2-5%. Diminishing returns if the workload isn't representative.

6b. Verify

Re-run the iai-callgrind benchmark and the flamegraph. Confirm the hot frames shifted and instruction counts dropped. If not, the profile workload was not representative — iterate.

Anti-patterns to refuse

  • "Profile my app" with no workload specified — ask for one.
  • Adding hotpath instrumentation during investigation — only add it with CI gates in mind.
  • PGO without a representative workload — worse than no PGO.
  • Swapping allocators without measuring — commit the benchmark result next to the swap.
  • Criterion for CI regression gates — variance makes small regressions invisible; use iai-callgrind.
  • cargo flamegraph on macOS/Windows — use samply.
  • Exposing pprof-rs endpoint without auth — never.
  • Running all samplers (samply + pprof + tracing-flame) at once — pick one per investigation.

Tool selection cheat-sheet

QuestionTool
Which of my instrumented functions regressed?hotpath
Where is time going, across everything?samply (dev) / pprof-rs (prod) / tracing-flame (if tracing)
Microbenchmark this functiondivan
CI regression gate, <1% sensitivityiai-callgrind
How much memory am I allocating and where?dhat-rs
Global allocator swapmimalloc
What would actually matter to optimize?coz-rs
Release binary optimizationcargo-pgo (PGO + BOLT)

Output format

When reporting findings back to the user, always structure as:

  1. Detected state — what's already in the repo (one line).
  2. Workload used — the exact command (one line).
  3. Hotspot — file:line or function name, with % of total time.
  4. Proposed fix — specific code change, not "optimize this".
  5. Measured impact — before/after from iai-callgrind or divan. No fix ships without this.

If step 5 can't be produced, say so explicitly rather than claiming the fix works.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.