agentsclimarketplace

Capacity planning

Skill stevancris/sre-ai-agent/skills/capacity-planning

AI agent that accumulates SRE knowledge from every incident — built on Agent Skills spec for Claude Code

Install
npx -y skills add stevancris/sre-ai-agent --skill capacity-planning

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Forecast infrastructure capacity needs and plan for growth. Use when planning for a product launch, reviewing resource utilization, handling autoscaling limits, responding to cost spikes from over-provisioning, or building a quarterly infrastructure roadmap. Trigger keywords: capacity, scaling, growth, traffic forecast, infrastructure sizing, autoscaling, resource limits, nodes, pods, instances, headroom, peak traffic, launch, viral growth, capacity crunch, running out of, max replicas, HPA limit, scale up.

SKILL.md

5.0 KB, as published. Nobody here has run it

Capacity Planning Skill

Setup Check

Before loading context files, check if context/CONTEXT.md exists in the current directory.

If context/CONTEXT.md exists — read it and proceed normally.

If context/CONTEXT.md does not exist — this skill was installed standalone (e.g. via npx skills add). Ask the user these questions before proceeding:

  1. Rolejunior-sre / senior-sre / sre-manager (shapes output depth and tone)
  2. Cloud provideraws / gcp / azure / on-prem / hybrid
  3. Observability stack — e.g. Datadog, Prometheus+Grafana, New Relic
  4. Company name and primary services affected (if relevant to this task)

Use the answers inline for this session. For persistent setup across all skills, suggest:

pipx install sre-agent
sre-agent init

Instructions

Step 1: Load Context

Read context/CONTEXT.md and context/company/tech-stack.md for cloud provider, container platform, and company stage.

Step 2: Gather Current Baseline

Ask the user for (or help them find):

  • Current resource utilization: CPU %, memory %, storage, network throughput
  • Current traffic: RPS at P50 / P95 / P99 / peak
  • Current pod/instance counts per service
  • Autoscaling configuration (min, max, target utilization)
  • Cost baseline (monthly cloud spend)

Step 3: Project Growth

Choose the appropriate growth model:

Linear growth model (steady, predictable business)

projected_traffic = current_traffic × (1 + monthly_growth_rate)^months

Exponential growth model (early startup, high growth)

projected_traffic = current_traffic × growth_multiplier^quarters

Event-driven model (product launch, seasonal peak)

peak_traffic = baseline_traffic × peak_multiplier
# Common peak multipliers: 2x (normal launch), 5x (viral launch), 10x (major campaign)

Ask the user:

  • What is the expected monthly traffic growth rate?
  • Are there any planned events (launches, campaigns, partnerships) that will spike traffic?
  • What is the planning horizon (3 months, 6 months, 1 year)?

Step 4: Identify Bottlenecks

For each resource dimension, calculate when it hits its limit:

Time to limit = (limit - current_value) / monthly_growth_rate

Example:
  CPU: currently at 45%, limit at 80% (with autoscaling headroom)
  Monthly traffic growth: 15%
  CPU grows proportionally to traffic
  Months to limit: (80 - 45) / (45 × 0.15) = 5.2 months

Bottleneck analysis table:

ResourceCurrentLimitGrowth RateMonths to Limit
CPU45%80%15%/mo5.2 months
Memory60%85%12%/mo2.1 months ← FIRST BOTTLENECK
Storage30%90%8%/mo7.5 months
Network20%80%20%/mo3.0 months

Step 5: Sizing Recommendations

For each bottleneck, provide concrete recommendations:

Kubernetes / containers:

  • Adjust resource requests and limits
  • Increase max replicas in HPA
  • Evaluate cluster node count and instance types
  • Consider vertical pod autoscaling for memory-bound services

Databases:

  • Read replica count for read-heavy scaling
  • Connection pool sizing
  • Instance class upgrade timeline
  • Sharding consideration for write-heavy growth

Caching:

  • Cache hit rate improvement (reduce origin load)
  • Cache cluster sizing

Step 6: Capacity Timeline

Produce a month-by-month action plan:

Month 1: Increase memory limits for payment-api (bottleneck at current rate)
Month 2: Add 5 nodes to prod EKS cluster (before projected limit)
Month 3: Upgrade RDS instance class (replication lag increasing)
Month 4: Add read replica for user-service DB
Month 6: Evaluate sharding strategy for orders table (projected 500GB)

Step 7: Cost Projection

Estimate cost of recommended capacity changes:

ChangeMonthly Cost ImpactOne-Time Cost
5 additional EKS nodes (m5.xlarge)+$1,200/mo
RDS instance upgrade (db.r6g.2xlarge)+$800/mo
Additional read replica+$600/mo
Total+$2,600/mo

Guidelines

  • Always include a "do nothing" scenario to show the cost of inaction (incident, degradation).
  • Build in a 30% headroom buffer beyond projected peak — forecasts are always wrong.
  • Autoscaling is not infinite: always check max replica counts and node pool limits.
  • Persona (sre-manager): include team bandwidth to execute the capacity plan and whether hiring is needed to manage increased infrastructure complexity.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.