agentsclimarketplace

Sota performance

Skill martinholovsky/SOTA-skills/skills/sota-performance

State-of-the-Art (2026) AI/LLM engineering skills/agents for building and auditing software — 40+ domain & language skills, BUILD/AUDIT modes, audit checklists.

Install
npx -y skills add martinholovsky/SOTA-skills --skill sota-performance

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 8 stars8 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

State-of-the-art performance engineering for building fast systems and auditing existing code for bottlenecks. Use when the task involves performance, optimization, latency, profiling, slow code, memory usage, caching, or throughput — designing latency budgets, fixing N+1 and accidental-quadratic patterns, tuning allocation/GC pressure, network and I/O efficiency, cache architecture, Core Web Vitals, or setting up benchmarks and regression gates. Not for concurrency correctness (races, deadlocks, cancellation) — use sota-async-concurrency. Trigger keywords: performance, optimization, latency, profiling, slow, memory usage, caching, throughput, bottleneck, p99, flamegraph, Core Web Vitals.

SKILL.md

9.8 KB, ~2.3k tokens by cl100k_base, as published. Nobody here has run it

SOTA Performance Engineering

Purpose

Make systems fast by default and find why they are slow by evidence. This skill encodes two disciplines that share one rule set:

  1. BUILD — write code whose performance characteristics are known, budgeted, and protected by regression tests before it ships.
  2. AUDIT — read existing code and telemetry to locate bottlenecks, rank them by user-facing impact, and prescribe fixes with expected gains.

Core doctrine: measure first, but fix known pathologies on sight. Profiling is mandatory before micro-optimization; it is NOT required to remove an O(n²) loop, an N+1 query, or an unbounded cache. "Premature optimization" never excuses shipping a known pathology.

BUILD mode

When writing new code or features:

  1. Set a budget before writing. Define the latency budget (p99, not average) and decompose it across hops. An endpoint with a 200 ms p99 budget that calls auth (10 ms) + 2 DB queries (2×15 ms) + serialization (5 ms) has 155 ms of headroom — spend it consciously. See rules/01-methodology.md.
  2. Choose data structures by access pattern, not habit. Know the n. n < 100: anything works. n unbounded: complexity class is the design. See rules/02-algorithms-data-structures.md.
  3. Batch and stream at every boundary. One round trip per collection, not per item. Stream large results; never materialize unbounded data.
  4. Control allocation in hot paths. Pre-size collections, reuse buffers, avoid per-iteration allocation in loops that run > 10⁴ times per second. See rules/03-memory.md.
  5. Make I/O cheap by construction. Pooled connections, keep-alive, buffered writes, compression chosen per payload type. See rules/04-io-network.md.
  6. Cache deliberately or not at all. Every cache ships with: key schema, TTL + jitter, eviction policy, invalidation path, stampede protection, and a hit-ratio metric. A cache missing any of these is a future incident. See rules/05-caching.md.
  7. Protect the win. Add a benchmark or perf test in CI for any code with a budget. A perf improvement without a regression gate is a loan, not an asset.
  8. Frontend ships against Core Web Vitals budgets (LCP ≤ 2.5 s, INP ≤ 200 ms, CLS ≤ 0.1 at p75). See rules/06-frontend-web.md.

AUDIT mode

How to find performance issues by reading code

Work outside-in, hottest path first:

  1. Identify the hot paths. Entry points with highest traffic or strictest SLO: request handlers, queue consumers, render loops, cron jobs over large datasets. Audit those first; ignore cold admin paths until the end.
  2. Grep for pathology signatures (high hit rate, low effort):
    • Loops containing await/network/DB calls → N+1 (rules/02)
    • String/array concatenation inside loops → accidental quadratic (rules/02)
    • .includes/in list/linear find inside a loop → O(n·m) (rules/02)
    • SELECT *, queries without LIMIT, missing pagination (rules/02, rules/04)
    • Caches/maps with insert but no eviction or TTL → leak (rules/03, rules/05)
    • addEventListener/subscribe without matching removal (rules/03)
    • Sequential awaits on independent operations → serialized latency (rules/04)
    • New client/connection per request instead of pooled (rules/04)
    • Sync file/crypto/compression calls on async event loops (rules/04)
    • JSON.parse/serialize of large payloads in hot loops (rules/03)
  3. Check the boundaries. Most production latency lives at boundaries: process↔kernel (syscalls), service↔DB, service↔service, server↔browser. Count round trips per user action; > 3 sequential round trips is a finding.
  4. Check resource lifecycle. Anything created per-request that is expensive to create (connections, TLS sessions, regexes, compiled templates, clients) should be created once and reused.
  5. Check what's missing: no timeouts, no pagination, no backpressure, no pool bounds, no cache eviction — absent code is the most common perf bug.

What to measure (when you can run the system)

  • Latency distribution: p50/p95/p99 per endpoint — never averages (rules/01). Compare p99 to p50; ratio > 10× means contention, GC, or stampedes, not slow code.
  • USE per resource (Utilization, Saturation, Errors): CPU, memory, disk, network, pools, queues. RED per service (Rate, Errors, Duration).
  • Where time goes: CPU flamegraph for compute, off-CPU/wall profile for waiting. A request that is slow with idle CPU is blocked on I/O or locks.
  • Allocation rate and GC pause time for managed runtimes.
  • Cache hit ratios and DB round trips per request.
  • Frontend: field CWV (CrUX/RUM) at p75, not lab-only Lighthouse.

Severity conventions (by user-facing impact)

SeverityCriteria
CriticalActive or imminent user-facing failure: unbounded growth (memory leak, unpaginated scan) that will OOM/timeout at production scale; O(n²)+ on user-controlled input; stampede-capable cache in front of a fragile origin; p99 SLO breached now.
HighMeasurable user-facing degradation: N+1 on a hot path; missing pool/keep-alive adding RTTs per request; blocking call on event loop; CWV in "poor" band; hot-path complexity that degrades super-linearly with organic growth.
MediumWasteful but currently within budget: avoidable allocations in warm paths; missing compression; suboptimal cache TTLs; sequential awaits worth ~10–50 ms; bundle over budget but CWV still "needs improvement".
LowHygiene: micro-inefficiencies in cold paths, style-level fixes, missing benchmarks for non-critical code.

Escalate one level if the code path is on the critical user journey (checkout, login, search) or if growth is super-linear with data/users.

Finding format

[SEVERITY] <one-line title>
Location: <file:line(s)>
Pattern: <pathology name, e.g. "N+1 query", "unbounded cache">
Evidence: <code excerpt or metric>
Impact: <quantified or estimated user-facing effect, with the math>
Fix: <specific change, with expected gain>
Verify: <how to confirm the fix: benchmark, profile, metric to watch>

Estimate impact with arithmetic, not adjectives: "200 items × 1 query × ~1 ms RTT = ~200 ms added per page view" beats "this is slow".

Rules index

FileRead this when...
rules/01-methodology.mdYou need to profile, benchmark, set latency budgets, interpret percentiles, apply USE/RED, decide what's worth optimizing (Amdahl), or set up CI perf regression gates.
rules/02-algorithms-data-structures.mdAuditing loops and data access: N+1, accidental quadratics, repeated scans, hash-vs-tree choices, batching, streaming vs materializing, and high-level DB pointers (indexes, SELECT *, chatty transactions).
rules/03-memory.mdDealing with allocation pressure, GC pauses, object pooling, arenas, cache locality, SoA vs AoS, or hunting memory leaks (closures, listeners, unbounded caches) per runtime.
rules/04-io-network.mdAnything crossing a syscall or the wire: buffering, zero-copy, connection pooling, HTTP/2/3, compression choice (zstd/brotli), TLS resumption, CDN, request coalescing, pagination over the wire.
rules/05-caching.mdDesigning or auditing any cache: hierarchy placement, key design, invalidation, stampede protection (singleflight, jitter, soft TTL), negative caching, and when caching is the wrong fix.
rules/06-frontend-web.mdWeb performance: Core Web Vitals thresholds, bundle budgets, code splitting, image formats (AVIF/WebP), font loading, hydration cost, edge rendering.

Top-10 non-negotiables

  1. Measure before optimizing; fix pathologies on sight. Profile before micro-tuning. But N+1, O(n²) on unbounded input, unbounded caches, and sync-blocking the event loop need no profiler — fix them when you see them.
  2. Percentiles, never averages. Report and budget p50/p95/p99. An average hides the 1% of users who hit every cache miss and GC pause.
  3. No I/O inside a loop over a collection. Batch it, join it, or parallelize it with a bound. One round trip per item is always a finding.
  4. Every cache has bounded size, TTL with jitter, an invalidation path, and stampede protection. Otherwise it's a memory leak with a hit ratio.
  5. Pool expensive resources. Connections, TLS sessions, threads, compiled regexes, HTTP clients: create once, reuse always, bound the pool.
  6. Stream unbounded data; never load "all rows" into memory. Paginate with cursors, process in chunks, set LIMITs.
  7. Never block an async event loop with sync file I/O, crypto, compression, or CPU-heavy work. Offload to workers or use async variants.
  8. Set timeouts and bounds on everything: requests, queries, pools, queues, retries (with backoff + jitter). Missing bounds turn slowness into outage.
  9. Sequential awaits on independent work are stolen latency. Run independent I/O concurrently; the latency of the batch is the max, not sum.
  10. Protect every win with a regression gate. A benchmark in CI with variance-aware thresholds, or the regression returns within a quarter.

What ships with it: 6 files

80.0 KB alongside SKILL.md

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.