agentsclimarketplace

Coordinate landing

Skill a-novel-kit/stack/.agents/skills/coordinate-landing

Development tools backing a-novel and a-novel-kit. Home of a-novel CLI and AI skills.

Install
npx -y skills add a-novel-kit/stack --skill coordinate-landing

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Vocabulary, invariants, and operator runbook for the cross-repo **landing saga**: how a multi-repo Epic lands atomically, recovers from a partial landing, rolls back, and releases through the `a-novel-kit/workflows` actions (merge-gate, epic-freeze, epic-rollback, release-train, the AGENT_KILL_SWITCH halt). Load it when a change spans several repos under one Epic. It owns the saga vocabulary and the **Epic Atomicity Rule**, and defers version mechanics to `manage-versions`, per-repo release mechanics to `prepare-release`.

SKILL.md

20.3 KB, as published. Nobody here has run it

The landing saga

A single feature often spans several repos — a golib change plus the two services that consume it. Under the a-novel model those repos are independently versioned and independently merged, so "land them together" is not free: it is a saga, a coordinated sequence with compensating actions when a step fails. This skill maps that saga: what the pieces are called, the one invariant they all serve, and the operator procedures that need a human.

The contributor-facing half — what a held or frozen Pull Request means to the person who opened it — is published in a-novel-kit/.github › docs/board-lifecycle.md. This skill is the operator's view.

The saga is enforced by the a-novel-kit/workflows governance actions, driven by two triggers: the per-PR / per-merge-group events, and a reconcile sweep that runs every ~15 minutes as a level-triggered, self-healing floor. Nothing here is a bespoke distributed-transaction engine — it is a small set of GitHub-native mechanisms (a required check, the merge queue, check-runs, a dispatch action) composed to make a multi-repo landing behave like one.


The Epic Atomicity Rule (INV-1)

All of an Epic's member PRs land together, or none of them do.

An Epic is a planning issue; its member PRs each carry the epic:<N> label. INV-1 is the invariant the whole saga exists to protect: a consumer must never merge while the dependency it needs is still unmerged (or vice-versa), because that leaves a repo un-buildable. Every mechanism below is either an enforcer of INV-1 (merge-gate, the merge queue) or a compensator for a violation of it (epic-freeze recovers, epic-rollback undoes). When a saga decision is unclear, ask which side of INV-1 it serves.

INV-1 is what makes the release train safe to run in any order: once an Epic has landed atomically, its repos are mutually consistent, so releasing them is order-independent. Cross-Epic ordering — Epic B depends on Epic A's release — is not INV-1's job; that is manage-versions.


Vocabulary — freeze it here, use it everywhere

Use these terms identically in code, PRs, issues, and conversation. No synonyms; never one word for two things. The vocabulary is a contract, and a reused name is a future bug.

TermMeaning
EpicThe planning issue grouping a multi-repo change; its number is N.
epic:<N> membershipThe label binding a PR to Epic N. Author-role-gated (only a maintainer can add it), so membership is trusted. It defines the set until activation freezes it — see Activation snapshot.
Activation snapshotThe member set frozen into the Epic issue body once it has held merge-ready long enough to settle. From then on the snapshot is the set: a PR de-labelled, closed, or relabelled afterwards stays a member.
WaveOne frozen set landing. A PR labelled after the freeze belongs to the next wave, and waits — the snapshot retires when every member is terminal, and the next ready set freezes its own.
Atomic landingAll member PRs merging together — INV-1 satisfied.
merge-gateThe required status check that holds an epic:<N> PR until the whole member set is ready + approved, then lets the merge queue land them together. A standalone (unlabelled) PR fast-paths to pass. On engagement of the halt it posts failure on every PR.
merge queueGitHub's native queue. The merge-gate re-evaluates over each frozen gh-readonly-queue/... head so the set greens and commits together.
Partial landingAn Epic where some members merged but others did not — an INV-1 violation (e.g. a member left the queue and did not re-enter).
epic-freeze / partial-landing detectorThe action + sweep pass that detects a partial landing and freezes every surviving sibling (posts failure on their heads + dequeues live groups) so no further member lands.
Grace windowThe 45-minute interval (3× the 15-min sweep) a stray sibling has to re-enter the queue before the freeze trips — absorbs normal queue churn.
Roll-forwardRe-enqueuing a landable stray within grace (enable auto-merge). Recovery forward, not a rollback — the preferred repair.
epic-rollbackThe human-triggered, admin-gated, VCS-layer compensator: reconstruct the merged epic:<N> ledger, git revert each squash newest-first, group the reverts under a fresh rollback-Epic, and land them in reverse through the unchanged merge-gate.
Release trainOne admin dispatch that releases every repo an Epic landed in — derive each repo's bump, drive its release.yaml, record the tag.
ReceiptThe tag a repo's release cut, recorded on the Epic. The release train's output and its idempotent-resume ledger.
AGENT_KILL_SWITCHThe org-wide fail-safe emergency halt (an org variable).
automation:pausedThe per-Epic pause label on the Epic issue — holds THAT Epic's automation, best-effort.
Blast-capThe per-Epic distinct-repo cap that trips a loud alert (freeze) or a hard abort (rollback) when an Epic's fan-out is suspiciously wide.
SagaThe whole coordinated land → detect → recover → release → archive lifecycle.
CoordinatorThe set of governance actions that drive the saga (the reconcile sweep + the per-event workflows).

Why membership freezes

Before landing starts, membership is the live epic:<N> label search. That works until a member is de-labelled mid-landing: GitHub indexes no "ever carried label X", so the PR vanishes from the set, and the detector reads the survivors as a clean landing while a member sits abandoned. The set is therefore captured into the Epic body once it has held merge-ready long enough for the label index to settle, and from that point the snapshot is the authority.

Two consequences worth knowing before operating an Epic:

  • A de-labelled member is still a member. Removing the label does not remove a PR from an in-flight landing. To stop the landing, use automation:paused on the Epic.
  • The snapshot names members; it does not authorise them. Each member must still show it carried epic:<N> — currently, or in its own label history. A named PR that never did voids the whole snapshot, because the Epic body is editable and its members become freeze targets.

What a merge group is not

A queued PR is validated as a merge group — a synthetic gh-readonly-queue/<base>/pr-<N>-<sha> ref — not as the pull request. The group has no author and no labels, so every PR-level exemption evaporates there. A ruleset bypass_actors entry is one: a bot allowance that lets a dependency PR merge does not survive the queue, and the base branch's required checks come back as hard requirements. Never assume a bypass covers the queue.

GitHub offers no way to scope a required status check to pull requests only. A flaky third party therefore cannot be exempted for the queue; the fix is to drop it as a required check and read its result from inside a check you control.

When a group sits pending, separate the two failure modes first: commits/<sha>/check-runs (Actions jobs) and commits/<sha>/status (third-party commit statuses) are different APIs. A green job list beside an empty status list means your upload job worked and the provider never reported back. The group then waits on a status that will never arrive, blocking every PR queued behind it until the queue's check-response timeout expires.


The saga lifecycle

author → LAND (merge-gate + queue, atomic) → DETECT (reconcile sweep)
                                                 ├─ whole?  → RELEASE (release train) → ARCHIVE
                                                 └─ partial → RECOVER (freeze + roll-forward within grace)
                                                                └─ unrecoverable → ROLL BACK (human-approved)
  1. Author. A maintainer labels each member PR epic:<N>.
  2. Land. The merge-gate holds each member until the whole set is ready + approved, then the merge queue lands them together — INV-1 satisfied atomically. A standalone PR is unaffected.
  3. Detect. The reconcile sweep (every ~15 min, level-triggered) re-derives each open Epic's state from live GitHub truth — it never trusts a stored flag, so it self-heals after any missed webhook.
  4. Recover. On a detected partial landing it freezes the surviving siblings and rolls forward any landable stray within the grace window. A frozen sibling holds (its required check goes red) until the Epic is whole again or a human intervenes.
  5. Release. Once whole, the release train releases the Epic's repos from one admin dispatch and records a tag receipt per repo.
  6. Archive. Each repo's release.yaml clears its awaiting-release board items after it ships.

Operator runbook

Every entry is admin-gated at the point of action; the destructive ones (rollback) additionally force a human approval. All the write actions honor AGENT_KILL_SWITCH and automation:paused.

Halt everything (incident brake)

Set the AGENT_KILL_SWITCH org variable (in the affected org) to any value that is not an off-token — e.g. on. Effect, org-wide and immediate on the next event/sweep:

  • merge-gate posts failure on every PR — nothing merges (the halt rests on the gate holding every PR, so it can't fast-path a standalone).
  • Every board writer, auto-merge arm, freeze poster, and rollback no-ops or refuses.

Lift by setting the value back to an off-token — canonically off (a created org variable cannot be empty, so off is the resting value; unset behaves the same but isn't discoverable). The switch is fail-safe: a fat-fingered or garbage value halts, not runs. It is a cooperative in-action flag, not a security boundary — if the incident is a compromised App, revoke the installation / rotate AGENT_BOT_PRIVATE_KEY instead.

Pause one Epic

Add the automation:paused label to the Epic issue. Holds THAT Epic's automation (merge-gate holds its members; the detector and rollback skip it) while other Epics keep moving. Best-effort — a label-read blip fails open (the pause is skipped that pass); it is an operator convenience, not the hard halt. Remove the label to resume.

Recover a partial landing

Usually automatic: the sweep freezes the surviving siblings and rolls a landable stray forward within grace. Intervene only when:

  • A stray can't land (merge conflict, failing checks). The freeze holds the whole surviving set (their required check is red) — fix the stray and let it re-enter the queue, or escalate to a rollback if the landed subset is genuinely broken.
  • The freeze is very wide (blast-cap tripwire fired: a loud red sweep). A freeze spanning more distinct repos than the cap is almost certainly a mis-scoped Epic — it still posts (fail toward freezing), but investigate the Epic's membership first.

Roll back an Epic (INV-1 genuinely violated)

Dispatch epic-rollback for Epic N. It is admin-only, dry_run-default, and typed-confirm (revert-epic-<N>). Always dry-run first — it prints the reconstructed ledger + the planned per-repo reverts and does no writes. Then run live:

  • It reconstructs the merged epic:<N> ledger from GitHub (REST-authoritative squash SHAs), git-reverts each squash newest-first, opens one revert PR per repo, and groups them under a fresh rollback-Epic epic:<M> through the unchanged merge-gate.
  • The App authored the reverts, and GitHub 422s a self-approval, so the wave PARKS pending a human approval — forced four-eyes on a destructive op. Review + approve each revert PR; the gate then greens and the wave lands in reverse, atomically.
  • A revert conflict opens a HELD draft placeholder that holds the whole wave (no partial rollback); finish it by hand. The blast-cap aborts loud above the cap (a rollback that wide is opt-in — fail toward NOT reverting).

A rollback does not un-merge history — each revert is a new forward commit, so a mistaken rollback is itself revertible.

Release an Epic (the release train)

Dispatch release-train for Epic N (admin-gated, dry_run-default). Rehearse first (the dry run dispatches each repo's release.yaml with dry_run=true and cuts nothing), then run live:

  • It reconstructs the Epic's landed repos, derives each repo's semver bump from its conventional-commit range (fix→patch, feat→minor, !/BREAKING CHANGE→major), and drives each repo's own release.yaml — any order (INV-1 makes them independent).
  • It records a receipt (the cut tag) per repo on the Epic. Idempotent resume: a re-dispatch skips every repo already shipped for this Epic (derived from live tags, never the receipt), so a partial train re-cuts only the unshipped.
  • A repo parked at the protected release environment approval gate is pending, not failed — approve the run, then re-dispatch to record the receipt. A botched cut (tag pushed, Release missing) is flagged loudly for manual repair, never silently skipped.

Cross-repo hotfix

A bug on a released line that spans several repos is not a special "hotfix train" — it is a standard Epic, run fast: label the fix PRs epic:<N>, let the merge-gate land them atomically, and release with the release train. There is deliberately no cross-repo hotfix orchestrator — the Epic machinery already gives atomicity plus a coordinated release. The single-repo hotfix path (hotfix.yaml: baseline → ephemeral → cut → reconcile → cleanup Task) and its vocabulary live in manage-versions.


Version coordination — the two rules manage-versions owns

The saga lands and releases a set of repos; keeping them version-compatible across that release is manage-versions' domain. Two rules from there govern how a saga is shaped:

  • Publish-before-rollout. When one repo's change is needed by another, the dependency PR merges and releases FIRST, and the consumer re-pins to the released tag before its PR merges. Within a single Epic the merge-gate lands the set together, but the release order across repos still honors this: release the dependency, re-pin, then release the consumer (the release train derives bumps per repo; cross-repo re-pins are manage-versions').
  • Expand→contract (staged breaking change). A breaking change never ships in one step: expand — add the new path alongside the old, non-breaking, and release; migrate every consumer; contract — remove the old path in a later release. Drafted ahead as blocked-by sub-issues. This is how a cross-repo breaking change stays landable atomically at every step: no single merge breaks a consumer, so INV-1 holds throughout.

See manage-versions for the mechanics (exact go.mod pins, pseudo-version development against an unreleased dep, the tag-push release, Renovate re-bumps) and its landing-failed runbook for the version-recovery decision tree.


How this composes

plan-feature (creates the Epic + Task sub-issues)
   └─ implement-feature (per-repo branches, one PR per Task, each labelled epic:<N>)
        └─ THE SAGA (this skill): merge-gate lands the set atomically (INV-1)
             ├─ partial landing → epic-freeze recovers, or epic-rollback undoes (human-approved)
             ├─ manage-versions: publish-before-rollout · expand→contract · landing-failed runbook
             └─ release-train releases the Epic's repos → archive
  • Enforcer skills: none — read the a-novel-kit/workflows actions for the ground truth.
  • resolve-pr-feedback for the conversation on a held or frozen PR — explain why it is held (waiting on its Epic set / frozen by a partial landing), not just that it is.
  • manage-versions for anything version-shaped in the rollout.

Principles

The mechanisms are described above; these are the judgment calls they encode.

  • INV-1 is the north star. Every mechanism enforces "land together or not at all," or compensates for a violation.
  • Recover forward before you roll back. Rollback is the last resort, and it is human-approved.
  • The sweep is the floor, level-triggered. Add no stateful shortcut.
  • Halt is fail-safe and cooperative. Revoke the App for a real compromise.
  • Freeze the vocabulary, use it consistently.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.