Gcp live gke rollout guard
Skill Raishin/vanguard-frontier-agentic/skills/gcp/gcp-live-gke-rollout-guard
Curated marketplace of AI skills, agents, and rules for cloud, zero-trust, and compliance-aware engineering - works with Claude Code, Codex, Cursor, Copilot, and more.
npx -y skills add Raishin/vanguard-frontier-agentic --skill gcp-live-gke-rollout-guardAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 18 stars18 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Gate GKE deployment mutations, node pool upgrades, and cluster control-plane version changes against rollback posture and PDB audit before any production change. Prevents irreversible node pool upgrades from proceeding without PodDisruptionBudget verification, surge settings review, and explicit operator approval.
SKILL.md
5.9 KB, ~1.2k tokens by cl100k_base, as published. Nobody here has run it
GCP Live GKE Rollout Guard
Purpose
Act as the guarded live GCP operator for gcp-live-gke-rollout-guard work. Gate GKE deployment mutations, node pool upgrades, and cluster control-plane version changes. Insist on PDB audit and rollback posture evidence before execution, and treat any ambiguous approval or target as a stop condition.
When to Use
Use this skill when:
- A GKE node pool upgrade is requested (Kubernetes minor or patch version bump)
- A cluster control-plane version change is planned
- A Deployment or DaemonSet rollout is being executed against a production workload
- Surge upgrade settings or max-unavailable values need to be changed
- An operator needs to audit PodDisruptionBudgets before a disruptive node pool operation
- An emergency rollback of a broken rollout is required
When NOT to Use
Do not use this skill when:
- The target is a non-production cluster with no PDB requirements and no live traffic
- The task is creating a brand-new cluster (no existing workloads at risk)
- The task is purely read-only cluster inspection with no mutation intent
- The task involves Cloud Run, App Engine, or other non-GKE compute
Pre-Flight Checklist
Before executing any GKE mutation, verify all of the following:
- Cluster identity confirmed — run
gcloud container clusters describe <CLUSTER> --region <REGION> --project <PROJECT>and confirm the cluster name, version, and region match the intended target. - Active principal confirmed — run
gcloud auth listandgcloud config get-value accountto confirm the active identity has the required role. - Current node pool version and target version captured — document both before proceeding; confirm the target version is available in the release channel.
- PDB audit complete — run
kubectl get pdb --all-namespacesand confirm no PDB hasDISRUPTIONS ALLOWED: 0for workloads running on the affected node pool. - Surge upgrade settings reviewed — confirm
maxSurgeandmaxUnavailablesettings on the node pool are appropriate for the workload disruption tolerance. - Rollback posture acknowledged — node pool upgrades cannot be downgraded; operator must explicitly acknowledge this is one-way.
- Maintenance window and change window confirmed — confirm the upgrade is within the approved maintenance window and any required change tickets are approved.
- Rollout history captured — run
kubectl rollout history deployment/<NAME> -n <NAMESPACE>to document the pre-change state for Deployment rollouts.
Required Confirmation
The operator must explicitly state all of the following before any mutation is executed:
- "I confirm the cluster is
<CLUSTER_NAME>in project<PROJECT_ID>, region<REGION>." - "I confirm the target version is
<TARGET_VERSION>and I understand node pool upgrades cannot be downgraded." - "I have reviewed PDB status for all workloads on this node pool and no disruption-blocking PDB is present."
- "I approve this rollout action."
Execution Steps
- Capture pre-change state: cluster version, node pool version, all PDB states, Deployment rollout history.
- Confirm active principal and IAM role (
roles/container.clusterAdminfor mutation). - Present the planned change and its blast radius to the operator for explicit approval.
- Execute the mutation:
- Node pool upgrade:
gcloud container node-pools upgrade <POOL> --cluster <CLUSTER> --region <REGION> --project <PROJECT> - Cluster control-plane upgrade:
gcloud container clusters upgrade <CLUSTER> --master --cluster-version <VERSION> --region <REGION> --project <PROJECT> - Deployment rollout:
kubectl set image deployment/<NAME> <CONTAINER>=<IMAGE> -n <NAMESPACE>
- Node pool upgrade:
- Monitor rollout progress:
kubectl rollout status deployment/<NAME> -n <NAMESPACE>orgcloud container operations describe <OPERATION_ID>. - Verify all nodes reach
Readystatus and all workloads are running post-upgrade.
Rollback Procedure
- Deployment rollback (reversible):
kubectl rollout undo deployment/<NAME> -n <NAMESPACE> - Node pool upgrade (NOT reversible): A completed node pool upgrade cannot be downgraded. If the upgrade fails mid-way, manual node recreation or a new node pool at the previous version is required.
- Control-plane upgrade (NOT reversible): Control-plane version cannot be rolled back. If the upgrade causes issues, you must address them on the upgraded version or redeploy the cluster.
- Document the incident and open a GCP Support case if node pool corruption is suspected.
Post-Change Verification
- Run
gcloud container clusters describe <CLUSTER>— confirmcurrentMasterVersionmatches target. - Run
gcloud container node-pools describe <POOL> --cluster <CLUSTER>— confirmversionmatches target. - Run
kubectl get nodes— confirm all nodes showReadywith the new version. - Run
kubectl get pods --all-namespaces— confirm no pods inCrashLoopBackOfforPendingstate. - Run
kubectl get pdb --all-namespaces— confirm all PDBs still show healthy disruption budgets. - Verify application health via service-level health checks, error rate, and latency metrics in Cloud Monitoring.
Response Shape
- Cluster and node pool identity confirmation
- Current cluster/node pool version vs. target
- PDB audit for affected workloads
- Rollout strategy and surge settings
- Approval status
- Proposed or executed rollout action
- Post-rollout verification steps