agentsclimarketplace

Perf budget

Skill vindreshsingh/engineering-skills/skills/perf-budget

Optimizes performance by measuring first and fixing the real bottleneck. Use when something is slow, before optimizing, when setting latency or bundle targets, or reviewing changes for performance traps — not for speculative micro-tuning.From its SKILL.md

Install
npx -y skills add vindreshsingh/engineering-skills --skill perf-budget

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

10.5 KB, ~2.5k tokens by cl100k_base, as published. Nobody here has run it

Perf Budget

Performance work starts with a number and a measurement, not a hunch. Set a budget, profile the real workload, fix the dominant cost, measure again, and add a guard so regression can't sneak back. Optimizing the wrong layer — micro-tuning a cold path while N+1 queries dominate — wastes effort and adds complexity.

A perf budget is an explicit target: p95 latency, bundle kilobytes, query time, memory ceiling. Without one, "faster" is unmeasurable and every optimization is premature.

Run the performance checklist alongside this process for common traps — only after you've confirmed the path is hot. Pair with [[observability]] for production numbers, [[data-modeling]] for query/index work, [[caching-strategy]] only after read cost is measured, [[react-patterns]] for React/Next.js specifics, and [[browser-checks]] for perceived slowness in the UI.

When to Use

  • A page, API, query, job, or interaction is measurably slow
  • Setting or enforcing targets — SLOs, bundle budgets, CI perf gates
  • Before optimizing — confirm where time actually goes
  • Reviewing PRs for obvious traps ([[review-gate]]) — N+1, waterfalls, unbounded work
  • After a launch when latency or error budget regressed ([[launch-readiness]])

Skip micro-optimization on cold paths, one-time setup, or code with no user-facing latency impact. Don't demand perf work on every PR — apply when the path is hot or the budget is at risk.

Not a substitute for [[caching-strategy]] — measure first; cache only when read cost justifies it.

Process

Work in order. No code changes until you have a baseline number.

1. Define the budget — make "fast" concrete

Write the target before profiling:

SurfaceExample budgets
APIp95 < 200ms, p99 < 500ms for POST /orders
PageLCP < 2.5s, TTI < 3s on 4G ([[browser-checks]])
Query< 50ms p95 at 10M rows; examines < 1% of table
JobProcess 10k msgs/min; p95 handle < 2s
BundleMain chunk < 200KB gzip; route chunk < 80KB
MemoryWorker steady < 512MB under peak load

Include:

  • Metric — what (latency, size, throughput)
  • Statistic — p50 vs p95 vs p99 (tail matters for UX)
  • Workload — realistic data volume, concurrency, payload size
  • Environment — prod-like, not empty dev DB

Align with SLOs in [[observability]] — perf budgets feed alerts and error budgets.

Bad:  "Checkout should be fast"
Good: "POST /orders p95 < 500ms at 100 RPS, prod-like cart (20 items)"

2. Measure baseline — production truth beats guesswork

Prefer production or prod-like measurements ([[observability]]):

  • APM traces — p95/p99 per endpoint, span breakdown
  • DB slow query log, EXPLAIN ANALYZE on realistic volume ([[data-modeling]])
  • Browser Performance tab, Lighthouse, Web Vitals ([[browser-checks]])
  • Load test at expected concurrency — find saturation, not just single-request time

Profile the dominant path:

LayerTools / signals
BackendTrace waterfall, CPU profiler, flame graph
DatabaseQuery plan, rows examined vs returned, lock wait
FrontendNetwork waterfall, React profiler, bundle analyzer
BatchPer-stage timing, queue lag

Record baseline numbers in the ticket/PR — you'll need them for after.

Intuition is wrong — the slow function you remember is often 2% of request time; one unindexed query or serial await chain is the real killer.

3. Find the bottleneck — one dominant cost

Ask: what single change would move the budget most?

Rank costs from profile/trace:

POST /orders p95 840ms breakdown (example):
  payment API     520ms  ← dominant
  DB insert        45ms
  tax calculation  12ms
  JSON serialize    3ms

Fix payment first — not JSON. Shaving 1ms off serialize when payment is 520ms is waste.

Hot path check — does this code run per request at scale? One-off admin export ≠ checkout API.

If multiple similar costs, fix the easiest high-impact first — sometimes parallelizing two 200ms calls beats optimizing one 50ms call.

4. Fix structural wins before micro-optimizations

Priority order (stop when budget met):

PriorityWinSkills
1Remove N+1; batch/join queries[[data-modeling]]
2Add/fix indexes for real query patterns[[data-modeling]]
3Parallelize independent I/O; kill waterfalls[[react-patterns]]
4Paginate/limit unbounded work[[data-modeling]]
5Algorithm fix — O(n²) → O(n), better structure[[simplify]]
6Cache expensive measured reads[[caching-strategy]]
7Lazy load / split bundle[[react-patterns]]
8Micro-tune hot loopOnly with profiler proof

Caching last among structural options — confirm read cost on hot path first ([[caching-strategy]]). A better query often removes the need entirely.

One change at a time when validating — or isolate in benchmark — so you know what moved the number ([[fault-recovery]] discipline).

5. Measure again — prove it worked or revert

After each meaningful change:

  • Re-run same workload / benchmark / trace comparison
  • Compare same statistic (p95, not p50 only)
  • Check regressions elsewhere — faster checkout but 2× DB load?
Before: p95 840ms | After: p95 310ms | Budget: 500ms ✓

If the number didn't move — revert the change unless it bought clarity or mandatory fix. Don't keep "optimizations" that profiler can't justify.

Document before/after in PR — reviewers and future-you need proof.

6. Account for trade-offs — speed isn't free

Every perf win can cost something — name it:

OptimizationTrade-off
CacheStaleness, invalidation bugs ([[caching-strategy]])
ParallelismConnection pool pressure, harder debugging
DenormalizationWrite complexity, consistency ([[data-modeling]])
Aggressive bundlingCacheability, deploy granularity
Lower samplingLess observability detail

Spend complexity only where measurement shows user-visible gain. A 5ms win on a 2s page isn't worth opaque code.

7. Add guards — prevent silent regression

Lock the win:

GuardWhen
Benchmark in CICritical pure logic; threshold assert
Bundle size budgetFrontend CI fails if chunk grows > X KB
Load test gateRelease pipeline p95 check ([[pipeline-ops]])
SLO alertProduction p95 regression ([[observability]])
Query plan checkEXPLAIN in test on representative query

Guards should match the same metric you optimized — not a different proxy that drifts.

Re-baseline when workload changes — budgets aren't forever.

8. Layer-specific playbooks

Slow API endpoint

Trace → rank spans → N+1 / missing index / external call → fix dominant → re-trace → SLO alert.

Slow SQL

EXPLAIN ANALYZE on prod stats → rows examined → index or rewrite → measure at volume ([[data-modeling]]).

Slow page load

Network waterfall → LCP element → bundle size + critical path → lazy load / server fetch parallel ([[react-patterns]], [[browser-checks]]).

Slow React interaction

Profiler → unnecessary re-renders vs expensive render → fix render or defer work → 60fps input check.

Background job behind

Queue depth + per-message timing → slow handler step → batch DB writes → throughput metric.

"Feels slow" user report

Reproduce with trace → compare p95 vs p50 (tail issue?) → [[browser-checks]] on real device/network.

Pre-launch perf pass ([[launch-readiness]])

Define budget → load test staging → fix blockers → dashboard + alert before 100% ramp.

9. When NOT to optimize

  • Cold path (admin export once a day)
  • Already within budget with headroom
  • Hypothetical scale ("might have 1M users someday") without current pain
  • Readability collapse for unmeasured micro-gains
  • Caching before measuring read cost
  • Optimizing staging with 3 rows when prod has 10M

Ship the feature; measure in prod; optimize what prod proves is slow ([[observability]]).

Common Rationalizations

  • "This code is obviously the slow part." — Profile it; obvious culprits are often innocent.
  • "Micro-optimizing can't hurt." — Complexity and risk without measured gain hurts maintainability.
  • "We don't have time to measure." — Less time than optimizing the wrong thing twice.
  • "Add a cache, it'll be faster." — Measure first; invalidation bugs cost more than slow reads.
  • "p50 looks fine." — Users hit p99 tail; tail is UX for many flows.
  • "Staging was fast enough." — Data volume and concurrency differ in prod.
  • "We'll add perf tests later." — Regressions ship in the meantime.
  • "Parallel always helps." — Pool exhaustion and contention can slow the system.

Red Flags

  • Optimizing before profiling or tracing
  • Fixing cold path while hot path still over budget
  • No before/after numbers in PR
  • Cache added without measured read cost or invalidation ([[caching-strategy]])
  • N+1 or full table scan visible in trace — unaddressed
  • Bundle grew with no budget check
  • Micro-optimization with no profiler evidence
  • Load test never run before high-tier launch
  • p50 cited when users complain about sporadic slowness
  • Perf "fix" that increases error rate or memory without acknowledgment
  • CI perf guard disabled to green the build

Verification

  • Concrete budget defined — metric, statistic, workload, environment
  • Baseline measured on hot path — trace/profile/query plan, not guesswork ([[observability]])
  • Dominant bottleneck identified and fixed first (structural before micro)
  • Performance checklist applied to changed hot paths
  • Before/after numbers documented; budget met or explicit acceptance of gap
  • Trade-offs named — cache, parallelism, consistency, complexity
  • Regression guard added or existing SLO/alert covers the metric ([[pipeline-ops]])
  • Caching only if read cost measured and invalidation designed ([[caching-strategy]])
  • Frontend: bundle and interaction checked if UI path ([[react-patterns]], [[browser-checks]])

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most performance cost skills give in ~2.5k tokens

Counted across 803 of the 1,058 authors here whose files we hold, read 2026-08-07

  • Keep skill files under 500 lines or tokensin 82 of 803, across 16 files
  • Use imperative form in instructionsin 80 of 803, across 9 files
  • Draft assertions while test runs are in progressin 75 of 803, across 9 files
  • Create two to three realistic test promptsin 74 of 803, across 9 files
  • Write skill descriptions to be pushyin 72 of 803, across 7 files
  • Save test cases to evals JSONin 72 of 803, across 6 files
  • Ask questions about edge cases and input formatsin 72 of 803, across 7 files
  • Save timing data immediately when runs completein 70 of 803, across 5 files
  • Include all trigger conditions in the skill descriptionin 69 of 803, across 3 files
  • Launch all test runs in a single turn or simultaneouslyin 69 of 803, across 3 files
  • Capture intent before writing a skillin 67 of 803, across 1 file
  • Import directly instead of barrel filesin 52 of 803, across 15 files

Said here and by no other author read

  • define a concrete performance budget before profiling
  • measure a baseline on a production-like environment
  • profile the dominant cost before fixing
  • fix structural bottlenecks before micro-optimizations
  • make one change at a time during validation
  • revert optimizations that lack profiler proof

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 326,851. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.