Standards go performance
Skill pecigonzalo/agent-skills/skills/standards-go-performance
Use this skill when diagnosing Go performance issues, tuning garbage collection, reviewing performance, or working on allocation-sensitive code. Provides pprof and trace workflows, allocation control, GC tuning guardrails, and runtime metrics patterns.From its SKILL.md
npx -y skills add pecigonzalo/agent-skills --skill standards-go-performanceAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 25 days oldThe repository was created 25 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
18.0 KB, ~4.4k tokens by cl100k_base, as published. Nobody here has run it
Go Performance Standards
Provides: pprof/trace profiling workflows, allocation control patterns, escape analysis guidance, GC tuning guardrails, and runtime metrics for production observability.
Primary references: Go diagnostics and GC guides, plus core Go style references.
Quick Reference
Golden Rule: Profile first. Never optimize without data.
Do (✅):
- Collect a focused profile before changing anything
- Preallocate slices with known capacity:
make([]T, 0, n) - Use
strings.Builder/bytes.Builderin hot paths - Reuse buffers with
sync.Poolwhen benchmarks justify it - Use
strconv(notfmt) for int/float↔string in hot loops - Hoist
[]byte("constant")outside loops; preallocate maps withmake(map[K]V, n) - Set
GOMEMLIMITin containers with ≥10% headroom - Export
runtime/metricsand alert on GC CPU fraction and goroutine count - Document before/after benchmark numbers for every change
Don't (❌):
- Optimize without a pprof or
benchmemprofile showing the bottleneck - Collect CPU and heap profiles simultaneously (they interfere)
- Use
fmt.Sprintfor+concatenation in hot loops - Box values into
any/interface{}in allocation-sensitive code - Set
GOMEMLIMITtoo tight (causes GC thrashing) - Disable GC (
GOGC=off) in shared or multi-tenant environments - Use
unsafetricks to force stack allocation
Key Commands:
go test -cpuprofile cpu.out -memprofile mem.out ./pkg
go tool pprof -http=:0 ./pkg.test cpu.out # flame graph UI
go test -bench . -benchmem ./... # allocs/op
go test -trace trace.out ./pkg && go tool trace trace.out
go build -gcflags=all='-m' ./... # escape analysis
GODEBUG=gctrace=1 ./myservice # live GC output
go tool pprof http://localhost:6060/debug/pprof/profile?seconds=30 # live CPU profile
go build -pgo=default.pgo ./cmd/myservice # PGO build
Pattern 1: Measure First: pprof Workflow
Use pprof to locate real bottlenecks before touching any code.
When: CPU hot paths, latency regressions, or unexplained memory growth.
How:
- Collect one targeted profile at a time: CPU or heap, not both simultaneously
- Benchmark:
go test -cpuprofile cpu.out -bench BenchmarkFoo ./pkg - Analyze interactively:
go tool pprof -http=:0 ./pkg.test cpu.out - Focus on
top(cumulative/flat),list FunctionName, and the flame graph - Change exactly one thing; re-collect the profile; compare
- Record baseline and post-change numbers in the PR or commit message
Pitfalls:
- Optimizing based on code reading alone: the bottleneck is rarely where you expect
- Collecting multiple profiles in a single run (timing interference skews results)
- Treating micro-benchmark results as production truth (cache, data size, and concurrency differ)
Verify: Documented before/after numbers exist; flame graph or top confirms the hot path moved.
Pattern 1b: Profiling Long-Running Services
For servers and daemons, enable pprof via HTTP rather than start/stop in main.
When: Production services, long-lived daemons, or any process where you cannot easily add start/stop calls around the interesting work.
How:
- Import
_ "net/http/pprof"in yourmainpackage (side-effect import registers handlers onhttp.DefaultServeMux). - Serve on a separate, internal-only port:
import (
_ "net/http/pprof"
"net/http"
)
// In main or server init: bind to loopback or internal network only
go http.ListenAndServe("localhost:6060", nil)
- Collect profiles on demand:
# 30-second CPU profile from a running service
go tool pprof http://localhost:6060/debug/pprof/profile?seconds=30
# Heap snapshot
go tool pprof http://localhost:6060/debug/pprof/heap
# Goroutine dump
go tool pprof http://localhost:6060/debug/pprof/goroutine
- For programmatic collection (e.g., triggered by a signal or health check):
import (
"os"
"runtime/pprof"
)
func writeCPUProfile(path string, fn func()) error {
f, err := os.Create(path)
if err != nil { return err }
defer f.Close()
if err := pprof.StartCPUProfile(f); err != nil { return err }
fn()
pprof.StopCPUProfile()
return nil
}
func writeHeapProfile(path string) error {
f, err := os.Create(path)
if err != nil { return err }
defer f.Close()
runtime.GC() // get up-to-date statistics
return pprof.WriteHeapProfile(f)
}
Pitfalls:
- Exposing
/debug/pprofon a public-facing port is a security risk: bind to loopback or a firewall-protected internal address only. - CPU and heap profiling simultaneously skews both: collect one at a time.
runtime.GC()before a heap snapshot forces a collection; do not call it in production hot paths.
Verify: curl -s http://localhost:6060/debug/pprof/ returns the index page; profiles download and open with go tool pprof.
Pattern 2: Allocation Control
Reduce allocs/op to lower GC pressure and improve throughput.
When: go test -bench . -benchmem shows high allocs/op, or GC CPU fraction is elevated in production.
How:
- Preallocate slices:
make([]T, 0, n)when approximate size is known - Build strings in hot paths with
strings.Builder; build byte slices withbytes.Builder - Pool short-lived, large buffers:
var bufPool = sync.Pool{
New: func() any { return new(bytes.Buffer) },
}
func process(data []byte) {
buf := bufPool.Get().(*bytes.Buffer)
buf.Reset()
defer bufPool.Put(buf)
// use buf ...
}
- Avoid boxing small values into
any/interface{}on hot paths
Hot-Path Allocation Tricks
Use these only when profiles show allocation pressure:
- Prefer
strconvoverfmtin hot conversion loops. - Hoist constant
[]byte("...")conversions out of loops. - Preallocate slices/maps for append/populate loops.
- Avoid unnecessary pointers for small value-like types when mutation is not required.
Pitfalls:
- Adding
sync.Poolbefore profiling shows allocation as the bottleneck: pool overhead is non-zero - Sharing a pooled buffer across goroutines without external synchronization
sync.Poolentries are GC'd between benchmark iterations; reset state (buf.Reset()) before use- Hoisting
[]byte("const")out of a loop only helps if the loop body runs frequently: verify withbenchmem
Verify: go test -bench . -benchmem shows reduced allocs/op; heap profile (-memprofile) confirms fewer live objects.
Pattern 2b: Concurrency Optimization
Reduce lock contention and avoid unnecessary synchronization overhead.
When: Mutex contention appears in a mutex profile (/debug/pprof/mutex) or a benchmark shows throughput bottleneck on shared state.
How:
Atomic operations for single-value counters:
Use sync/atomic typed operations (Go 1.19+) for simple shared scalars: faster than a mutex for non-compound operations:
import "sync/atomic"
type Counter struct{ n atomic.Int64 }
func (c *Counter) Inc() { c.n.Add(1) }
func (c *Counter) Value() int64 { return c.n.Load() }
Limit atomics to simple load/store/add on a single field. For compound operations (read-then-write, multi-field updates) use a mutex.
Sharded locks for high-contention maps:
When a single sync.RWMutex on a map is the bottleneck, shard the map across N independent mutexes:
const numShards = 256
type ShardedMap struct {
shards [numShards]struct {
mu sync.RWMutex
items map[string]any
}
}
func (m *ShardedMap) shard(key string) *struct {
mu sync.RWMutex
items map[string]any
} {
h := fnv.New32()
h.Write([]byte(key))
return &m.shards[h.Sum32()%numShards]
}
func (m *ShardedMap) Get(key string) (any, bool) {
s := m.shard(key)
s.mu.RLock()
defer s.mu.RUnlock()
v, ok := s.items[key]
return v, ok
}
Pitfalls:
- Reaching for atomics or sharding before the mutex profile confirms contention: premature complexity.
- Sharding with a poor hash distributes unevenly; benchmark with a representative key distribution.
- Compound atomic operations (e.g., check-then-set) require a mutex or
compare-and-swaploop.
Verify: Mutex profile (go tool pprof .../debug/pprof/mutex) shows reduced contention; throughput benchmark (go test -bench . -benchmem) confirms improvement.
Pattern 3: Escape Analysis & Inlining
Keep small, short-lived values on the stack by understanding how the compiler decides.
When: Heap profile shows unexpected allocations for small structs or values that look stack-local.
How:
- Inspect compiler decisions:
go build -gcflags=all='-m' ./... - Keep functions small to enable inlining (inlined callees don't need heap-allocated frames)
- Avoid returning a pointer to a local when the caller doesn't need the address; pass by value for small structs
- Interfaces cause escapes: avoid interface-boxing in loops
# Sample output: look for "escapes to heap" on hot allocations
go build -gcflags=all='-m -m' ./pkg 2>&1 | grep 'escapes to heap'
Pitfalls:
- Refactoring for stack placement without confirming the allocation is a measured bottleneck
- Relying on specific escape decisions across Go versions: the compiler may change heuristics
- Using
unsafeto force stack allocation (undefined behavior, never do this)
Verify: -gcflags=-m output no longer shows the target allocation escaping; benchmem allocs/op drops.
Pattern 3b: Algorithmic Improvements
Structural code changes often yield larger gains than low-level micro-optimizations.
When: Profile shows time spread across a loop body rather than concentrated in one call; or throughput is limited by N×M operations.
How:
Early returns and fast paths:
// ✅ Skip work early: avoid expensive computation for common no-op cases
func process(items []Item) {
for _, item := range items {
if !item.IsValid() {
continue // cheap check first
}
if item.IsSimple() {
handleSimple(item) // fast path
continue
}
handleComplex(item) // slow path only when needed
}
}
Batch operations: Reduce per-call overhead (network round-trips, syscalls, lock acquisitions) by processing multiple items in one call:
// ❌ N round-trips
for _, item := range items { db.Insert(ctx, item) }
// ✅ 1 round-trip
db.BatchInsert(ctx, items)
Zero-copy buffer reuse: Reset a slice to zero length while keeping its backing array, avoiding reallocation:
type Parser struct{ buf []byte }
func (p *Parser) Parse(data []byte) {
p.buf = p.buf[:0] // reset length, keep capacity
p.buf = append(p.buf, data...)
// process p.buf ...
}
Stream directly to writers instead of building an intermediate []byte:
// ✅ Write directly: no intermediate allocation
json.NewEncoder(w).Encode(value)
// ❌ Allocates intermediate []byte
b, _ := json.Marshal(value)
w.Write(b)
Pitfalls:
- Fast-path branches that are never exercised add dead code and maintenance burden: verify with coverage.
- Batch sizes that are too large increase latency and memory pressure: benchmark with realistic data.
- Reusing
buf[:0]is only safe when the caller does not retain a reference to the previous content.
Verify: Benchmark shows improvement proportional to reduced calls or allocations; profile confirms the hot path moved.
Pattern 4: Tracing for Latency & Parallelism
Use the execution tracer to diagnose scheduler stalls, GC pauses, and goroutine blocking.
When: Tail latency spikes, work not spreading across cores, or GC STW pauses suspected.
How:
- Collect:
go test -trace trace.out ./pkg && go tool trace trace.out - In the trace UI, inspect:
- Goroutine states: runnable (waiting for P) vs blocked (syscall/chan/mutex)
- GC events: STW mark setup, mark termination pause durations
- Network poller: time goroutines spend waiting on I/O
- For in-process tracing: use
runtime/tracepackage withtrace.Start/trace.Stop
Pitfalls:
- Using the tracer to find CPU hotspots: pprof is the right tool for that
- Long-running trace sessions produce large files that are slow to load
- Interpreting trace without understanding the Go scheduler (M:N model; P = logical processor)
Verify: Trace shows the blocking or pause event; fix is validated with a second trace or a latency histogram (p99 drops).
Pattern 5: Runtime Metrics & Health Signals
Export Go runtime telemetry for production visibility into GC behavior, heap size, and goroutine health.
When: Setting up production monitoring, investigating memory growth, or designing GC-related alerts.
How:
- Use
runtime/metrics(Go 1.16+) for structured, stable metric names:
import "runtime/metrics"
samples := []metrics.Sample{
{Name: "/gc/cycles/total:events"},
{Name: "/memory/classes/heap/objects:bytes"},
{Name: "/sched/goroutines:goroutines"},
}
metrics.Read(samples)
- Key signals to expose:
/gc/cycles/total:events: GC frequency/memory/classes/heap/objects:bytes: live heap/memory/classes/heap/released:bytes: memory returned to OS/sched/goroutines:goroutines: goroutine count (leak detection)/cpu/classes/gc/total:cpu-seconds: GC CPU cost
- Wire to Prometheus or your metrics backend; see
standards-observabilityfor infrastructure patterns
Pitfalls:
- Monitoring only RSS/VSZ: these miss GC dynamics and can lag heap behavior
- Ignoring
heap/released: rising live heap with no release indicates growth - Not tracking GC CPU fraction separately; it can dominate application CPU silently
Verify: Dashboards show stable trends under load; oncall runbook references these metrics and their alert thresholds.
Pattern 5b: Runtime Tuning
Adjust Go runtime parameters when defaults are a measured bottleneck: not before.
When: CPU profiling shows scheduler overhead or goroutine stalls; or you need to tune memory behavior beyond GOMEMLIMIT.
How:
- Keep
GOMAXPROCSat defaults unless traces/profiles justify a change. - Use
GOMEMLIMITandGOGConly with measured evidence and sufficient headroom. - Keep process-wide runtime knob changes in
main, not libraries.
Pitfalls:
runtime.GOMAXPROCScalled in library code affects the entire process: leave it to the application'smain.runtime.ReadMemStatstakes a STW snapshot: avoid calling it on hot paths or in tight loops.debug.SetGCPercentchanges are process-wide; calling it from multiple goroutines simultaneously can produce unexpected results if the intent is to temporarily change and restore the value: serialize such calls.
Verify: Trace shows reduced scheduler stalls or lower P-idle time; benchmark confirms throughput improvement.
Pattern 6: GC Tuning Guardrails
Tune the GC only after profiling shows GC is the bottleneck; wrong settings cause more harm than good.
Caveat: GC guide advice applies to the standard
gctoolchain (~Go 1.19+); behavior is not guaranteed by the language spec. For core language idioms see Effective Go (https://go.dev/doc/effective_go). For runtime internals, always measure on your workload.
When: GC CPU fraction dominates the CPU budget (visible in trace or gctrace) or running in memory-constrained containers.
How:
- Start with defaults (
GOGC=100). Most services do not need tuning. - Increase
GOGC(e.g.,GOGC=200) only when: memory headroom exists and GC CPU is the confirmed bottleneck. Higher GOGC → fewer GC cycles → larger heap. - Set
GOMEMLIMITin containers to cap heap growth:
# Allow up to 800 MiB; container limit is 1 GiB (≥10% headroom)
GOMEMLIMIT=838860800 ./myservice
- Validate with
GODEBUG=gctrace=1:
gc 14 @12.345s 2%: 0.1+2.1+0.3 ms clock, 0.8+1.5/2.0/0.1+2.4 ms cpu, 512->512->256 MB, 512 MB goal
- Field 3 is GC CPU%; field 7 shows heap before→after→live; confirm pause times are acceptable.
Pitfalls:
- Setting
GOMEMLIMITtoo close to the container limit: the GC will thrash trying to stay under it - Setting
GOGC=offin shared environments: unbounded heap growth will OOM neighbors - Tuning before measuring:
GOGCandGOMEMLIMITinteract; always profile first
Verify: gctrace shows acceptable pause times (<1 ms STW for most services) and GC CPU % drops; no OOM or thrashing under sustained load.
Pattern 7: Profile-Guided Optimization (PGO)
Read the PGO guide only after representative profiles exist and measured hotspots remain.
Skill Loading Triggers
| Situation | Also load |
|---|---|
| Latency / trace analysis | standards-observability |
| Writing performance benchmarks | standards-go-testing |
| Reviewing allocation-sensitive code | role-code-review |
| Concurrency lock contention | standards-go-concurrency |
Verification Checklist
For baseline formatting, vet, and go test checks see
standards-go.
- Performance change justified by pprof or benchmark evidence (not intuition)
- Before/after benchmark numbers documented;
go test -bench . -benchmemrun - Only one profile type collected per benchmark run (CPU and heap not simultaneous)
- Preallocation / builder / conversion optimizations are applied where hot paths justify them
-
sync.Poolusage is benchmark-justified and safely reset - Escape analysis (
-gcflags=-m) reviewed for unexpected heap escapes -
GOMEMLIMITset with >=10% headroom below container memory limit -
GOGCchanged only when GC CPU fraction is the measured bottleneck; gctrace reviewed after - Runtime metrics exported; goroutine count alerting configured
- pprof HTTP endpoint bound to loopback/internal address only
- PGO profile collected from representative production traffic (not synthetic load)
Note: file list is sampled.
What ships with it: 1 file
1.0 KB alongside SKILL.md