agentsclimarketplace

Shipping and launch

Skill celestialdust/achilles-skills/skills/shipping-and-launch

Prepares production releases and authors the launch runbook. Use the moment you are preparing to ship to production, batching merged PRs into a release, or anyone asks for a pre-launch checklist, a feature-flag rollout, a staged/canary rollout plan, monitoring setup, or a rollback strategy. In the v1 autonomous run this AUTHORS the runbook only — it never fires deploy/rollout/rollback commands itself.From its SKILL.md

Install
npx -y skills add celestialdust/achilles-skills --skill shipping-and-launch

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

13.5 KB, ~3.1k tokens by cl100k_base, as published. Nobody here has run it

Shipping and Launch

Purpose

Stage: Ship — the release workhorse (post-merge, release-level; may batch slices).

Ship with confidence. The goal is not just to deploy — it's to deploy safely, with monitoring in place, a rollback plan ready, and a clear understanding of what success looks like. Every launch should be reversible, observable, and incremental.

v1 autonomy fence. This skill AUTHORS the release runbook; it does not execute deploy, rollout, or rollback commands. In the autonomous run the agent's span ends at the open PR (pull-request workhorse) — every DEPLOY / ENABLE / canary-advance / ROLL BACK step below is a runbook the human runs post-merge, never an action the agent fires unattended. Auto-deploy and auto-merge are out of v1 (branch-naming.md: never commit to main). When this skill runs inside the orchestrator, its output is a written plan in release.md, not a side effect.

When to use / when to skip

  • Deploying a feature to production for the first time
  • Releasing a significant change to users
  • Migrating data or infrastructure
  • Opening a beta or early access program
  • Any deployment that carries risk (all of them)
  • Skip when nothing has merged yet (there is no release to prepare) or for a docs-only / config-only change with no user-visible runtime effect.

Inputs

Consumes (refuse to run if absent):

  • The shipped PR(s) from pull-request — one or more design-anchored, risk-banded PRs that are merged or queued for the human's async merge. A release batches the slices that ship together. If no merged/approved PR exists, there is nothing to release — STOP.
  • environment.md (typed-kind manifest; read-only, value-blind) — names the production + staging targets and the feature-flag / monitoring services as typed rows. It carries NO values and NO commands; read it for what exists, never for secrets or shell strings.
  • prd.md + referenced ADRs — the success definition the rollout decision thresholds bind to (which business metrics matter, what "Good" means for this feature).

Refuse-to-run if: no merged/approved PR (nothing to ship); environment.md is missing or its rows are unprovisioned (run preflight-readiness first); or no rollback path is identifiable for the change.

The Pre-Launch Checklist

Code Quality

  • All tests pass (unit, integration, e2e)
  • Build succeeds with no warnings
  • Lint and type checking pass
  • Code reviewed and approved
  • No TODO comments that should be resolved before launch
  • No console.log debugging statements in production code
  • Error handling covers expected failure modes

Security

  • No secrets in code or version control
  • npm audit shows no critical or high vulnerabilities
  • Input validation on all user-facing endpoints
  • Authentication and authorization checks in place
  • Security headers configured (CSP, HSTS, etc.)
  • Rate limiting on authentication endpoints
  • CORS configured to specific origins (not wildcard)

Performance

  • Core Web Vitals within "Good" thresholds
  • No N+1 queries in critical paths
  • Images optimized (compression, responsive sizes, lazy loading)
  • Bundle size within budget
  • Database queries have appropriate indexes
  • Caching configured for static assets and repeated queries

Accessibility

  • Keyboard navigation works for all interactive elements
  • Screen reader can convey page content and structure
  • Color contrast meets WCAG 2.1 AA (4.5:1 for text)
  • Focus management correct for modals and dynamic content
  • Error messages are descriptive and associated with form fields
  • No accessibility warnings in axe-core or Lighthouse

Infrastructure

  • Environment variables set in production
  • Database migrations applied (or ready to apply)
  • DNS and SSL configured
  • CDN configured for static assets
  • Logging and error reporting configured
  • Health check endpoint exists and responds

Documentation

  • README updated with any new setup requirements
  • API documentation current
  • ADRs written for any architectural decisions
  • Changelog updated
  • User-facing documentation updated (if applicable)

Feature Flag Strategy

Ship behind feature flags to decouple deployment from release:

// Feature flag check
const flags = await getFeatureFlags(userId);

if (flags.taskSharing) {
  // New feature: task sharing
  return <TaskSharingPanel task={task} />;
}

// Default: existing behavior
return null;

Feature flag lifecycle:

1. DEPLOY with flag OFF     → Code is in production but inactive
2. ENABLE for team/beta     → Internal testing in production environment
3. GRADUAL ROLLOUT          → 5% → 25% → 50% → 100% of users
4. MONITOR at each stage    → Watch error rates, performance, user feedback
5. CLEAN UP                 → Remove flag and dead code path after full rollout

Rules:

  • Every feature flag has an owner and an expiration date
  • Clean up flags within 2 weeks of full rollout
  • Don't nest feature flags (creates exponential combinations)
  • Test both flag states (on and off) in CI

Staged Rollout

The Rollout Sequence

1. DEPLOY to staging
   └── Full test suite in staging environment
   └── Manual smoke test of critical flows

2. DEPLOY to production (feature flag OFF)
   └── Verify deployment succeeded (health check)
   └── Check error monitoring (no new errors)

3. ENABLE for team (flag ON for internal users)
   └── Team uses the feature in production
   └── 24-hour monitoring window

4. CANARY rollout (flag ON for 5% of users)
   └── Monitor error rates, latency, user behavior
   └── Compare metrics: canary vs. baseline
   └── 24-48 hour monitoring window
   └── Advance only if all thresholds pass (see table below)

5. GRADUAL increase (25% -> 50% -> 100%)
   └── Same monitoring at each step
   └── Ability to roll back to previous percentage at any point

6. FULL rollout (flag ON for all users)
   └── Monitor for 1 week
   └── Clean up feature flag

Rollout Decision Thresholds

Use these thresholds to decide whether to advance, hold, or roll back at each stage:

MetricAdvance (green)Hold and investigate (yellow)Roll back (red)
Error rateWithin 10% of baseline10-100% above baseline>2x baseline
P95 latencyWithin 20% of baseline20-50% above baseline>50% above baseline
Client JS errorsNo new error typesNew errors at <0.1% of sessionsNew errors at >0.1% of sessions
Business metricsNeutral or positiveDecline <5% (may be noise)Decline >5%

When to Roll Back

Roll back immediately if:

  • Error rate increases by more than 2x baseline
  • P95 latency increases by more than 50%
  • User-reported issues spike
  • Data integrity issues detected
  • Security vulnerability discovered

Monitoring and Observability

What to Monitor

Application metrics:
├── Error rate (total and by endpoint)
├── Response time (p50, p95, p99)
├── Request volume
├── Active users
└── Key business metrics (conversion, engagement)

Infrastructure metrics:
├── CPU and memory utilization
├── Database connection pool usage
├── Disk space
├── Network latency
└── Queue depth (if applicable)

Client metrics:
├── Core Web Vitals (LCP, INP, CLS)
├── JavaScript errors
├── API error rates from client perspective
└── Page load time

Error Reporting

// Set up error boundary with reporting
class ErrorBoundary extends React.Component {
  componentDidCatch(error: Error, info: React.ErrorInfo) {
    // Report to error tracking service
    reportError(error, {
      componentStack: info.componentStack,
      userId: getCurrentUser()?.id,
      page: window.location.pathname,
    });
  }

  render() {
    if (this.state.hasError) {
      return <ErrorFallback onRetry={() => this.setState({ hasError: false })} />;
    }
    return this.props.children;
  }
}

// Server-side error reporting
app.use((err: Error, req: Request, res: Response, next: NextFunction) => {
  reportError(err, {
    method: req.method,
    url: req.url,
    userId: req.user?.id,
  });

  // Don't expose internals to users
  res.status(500).json({
    error: { code: 'INTERNAL_ERROR', message: 'Something went wrong' },
  });
});

Post-Launch Verification

In the first hour after launch:

1. Check health endpoint returns 200
2. Check error monitoring dashboard (no new error types)
3. Check latency dashboard (no regression)
4. Test the critical user flow manually
5. Verify logs are flowing and readable
6. Confirm rollback mechanism works (dry run if possible)

Rollback Strategy

Every deployment needs a rollback plan before it happens:

## Rollback Plan for [Feature/Release]

### Trigger Conditions
- Error rate > 2x baseline
- P95 latency > [X]ms
- User reports of [specific issue]

### Rollback Steps
1. Disable feature flag (if applicable)
   OR
1. Deploy previous version: `git revert <commit> && git push`
2. Verify rollback: health check, error monitoring
3. Communicate: notify team of rollback

### Database Considerations
- Migration [X] has a rollback: `npx prisma migrate rollback`
- Data inserted by new feature: [preserved / cleaned up]

### Time to Rollback
- Feature flag: < 1 minute
- Redeploy previous version: < 5 minutes
- Database rollback: < 15 minutes

See Also

  • For the project-wide Definition of Done that every change must clear before this checklist, see ../../references/definition-of-done.md
  • For security pre-launch checks, see ../../references/security-checklist.md
  • For performance pre-launch checklist, see ../../references/performance-checklist.md
  • For accessibility verification before launch, see ../../references/accessibility-checklist.md

Rationalizations

RationalizationReality
"It works in staging, it'll work in production"Production has different data, traffic patterns, and edge cases. Monitor after deploy.
"We don't need feature flags for this"Every feature benefits from a kill switch. Even "simple" changes can break things.
"Monitoring is overhead"Not having monitoring means you discover problems from user complaints instead of dashboards.
"We'll add monitoring later"Add it before launch. You can't debug what you can't see.
"Rolling back is admitting failure"Rolling back is responsible engineering. Shipping a broken feature is the failure.

Red flags

  • Deploying without a rollback plan
  • No monitoring or error reporting in production
  • Big-bang releases (everything at once, no staging)
  • Feature flags with no expiration or owner
  • No one monitoring the deploy for the first hour
  • Production environment configuration done by memory, not code
  • "It's Friday afternoon, let's ship it"
  • Agent (not the human) executing a DEPLOY / canary-advance / ROLL BACK command unattended — the v1 fence makes deploy a human-post-merge action.

Verification (ending criteria)

Before deploying:

  • Pre-launch checklist completed (all sections green)
  • Feature flag configured (if applicable)
  • Rollback plan documented
  • Monitoring dashboards set up
  • Team notified of deployment

After deploying:

  • Health check returns 200
  • Error rate is normal
  • Latency is normal
  • Critical user flow works
  • Logs are flowing
  • Rollback tested or verified ready

Outputs & handoff contract

Emits the release artifact → release.md under docs/features/<slug>/ (or docs/releases/<date>-<slug>.md for a batched multi-feature release). Stable sections a consumer depends on — change the shape, update the consumer in the same commit:

  • Pre-launch checklist — every section green (Code Quality · Security · Performance · Accessibility · Infrastructure · Documentation) with the gating items checked.
  • Feature-flag plan — flag name, owner, expiration date, and the OFF → team → canary → 100% lifecycle.
  • Staged-rollout plan — the rollout sequence plus the decision-threshold table (advance / hold / roll-back per metric).
  • Rollback plan — trigger conditions, steps, DB considerations, and the time-to-rollback budget (the ## Rollback Plan for [Feature] template, filled).
  • Monitoring setup — the dashboards/alerts to watch + the first-hour post-launch verification list.
  • Deploy fence — an explicit Executed by: human, post-merge banner. In v1 the agent authors this runbook but executes no deploy/rollout/rollback command; every action is fenced behind the human's async merge (auto-deploy out of v1).

STATE.md: shipping-and-launch is release-level and post-merge — it does NOT own a slice row and never flips a slice gate to agent for a deploy action. Once the feature's slices are done (PRs merged), record release.md under that feature's Artifacts cell. If a CRITICAL/HIGH security finding or a secret surfaces while preparing the release → hard STOP, fire PushNotification, open no release (security.md).

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 3 of the 12 instructions most ship operate skills give in ~3.1k tokens

Counted across 779 of the 1,178 authors here whose files we hold, read 2026-08-07

  • Document a rollback plan before deploymenthere, and in 41 of 779, across 22 files
  • Update the changelogin 21 of 779, across 19 files
  • Run the test suitein 20 of 779
  • Create an annotated git tagin 20 of 779
  • Clean up feature flags after full rollouthere, and in 18 of 779, across 10 files
  • Verify deployment health after launchin 18 of 779, across 10 files
  • Test both feature flag statesin 17 of 779, across 9 files
  • Verify the working tree is cleanin 17 of 779
  • Make database migrations backward-compatiblein 16 of 779, across 8 files
  • Set up error monitoring before launchhere, and in 15 of 779, across 7 files
  • Monitor metrics at each rollout stagein 14 of 779, across 5 files
  • Create a GitHub releasein 14 of 779

Said here and by no other author read

  • batch merged slices into a single release
  • stop if no merged pull request exists
  • write the release plan to the release file

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 326,512. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.