agentsclimarketplace

K8s ops

Skill addxai/enterprise-harness-engineering/skills/k8s-ops

Manage Kubernetes cluster resources via kubectl. Use when user needs to view, troubleshoot, or modify K8s workloads across multiple clusters.From its SKILL.md

Install
npx -y skills add addxai/enterprise-harness-engineering --skill k8s-ops

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • skips confirmationTells the agent to proceed without asking first, 2 times: "View operations carry no risk and can be executed directly" and 1 more.
  • runs commandsInstructs the agent to run 8 commands, including `kubectl version --client` and 7 more.

SKILL.md

7.6 KB, ~1.8k tokens by cl100k_base, as published. Nobody here has run it

k8s-ops

Assists operations engineers in managing K8s cluster resources via kubectl. Supports viewing, troubleshooting, and change operations.

Applicable scenarios:

  • View Pod/Deployment/Service/Node and other resource statuses
  • View logs and events for troubleshooting
  • Execute change operations (scale, rollout restart, apply, etc.)
  • Troubleshooting operations (exec, port-forward, describe, etc.)

Setup

Before using this skill, configure your cluster contexts in the table below. Replace the example entries with your actual clusters:

ClusterContextCloud Providerk8s Repo Directory
prod-1your-eks-context-hereAWS EKSclusters/prod-1/
stagingyour-gke-context-hereGCP GKEclusters/staging/
devyour-aks-context-hereAzure AKSclusters/dev/

Add all clusters your team manages. The context value should match the output of kubectl config get-contexts.

Execution Flow

Step 1: Environment Check

On first operation, the Agent should automatically check the environment:

  1. Run kubectl version --client to confirm kubectl is installed
  2. Run kubectl config get-contexts to list configured contexts

If kubectl is not installed: Guide the user to install it (brew install kubectl / apt install kubectl / official documentation).

If the target context is not configured:

  • AWS EKS clusters: Guide the user to run aws eks update-kubeconfig --name <cluster-name> --region <region>
  • GCP GKE clusters: Guide the user to run gcloud container clusters get-credentials <cluster-name> --region <region> --project <project>, and ensure gke-gcloud-auth-plugin is installed
  • Azure AKS clusters: Guide the user to run az aks get-credentials --resource-group <rg> --name <cluster-name>
  • Other clusters: Guide the user to download kubeconfig through the corresponding cloud platform console

If already configured: Proceed directly to Step 2.

Step 2: Confirm Operation Target

Confirm the following information with the user (proactively ask if not provided):

  1. Target cluster: Match the cluster table above based on the user's description; determine the context
  2. Target namespace: Use default if not specified, or infer from the service name
  3. Operation content: View, change, troubleshoot, etc.

Step 3: Switch Context

Use the --context parameter to specify the cluster — do not modify the current default context:

kubectl --context <context-name> -n <namespace> <command>

Step 4: Execute Operation

View Operations (execute directly)

View operations carry no risk and can be executed directly:

# View Pod status
kubectl --context <ctx> -n <ns> get pods

# View Deployments
kubectl --context <ctx> -n <ns> get deployments

# View Pod logs
kubectl --context <ctx> -n <ns> logs <pod-name> --tail=100

# View events
kubectl --context <ctx> -n <ns> get events --sort-by='.lastTimestamp'

# View resource details
kubectl --context <ctx> -n <ns> describe pod <pod-name>

# View Node status
kubectl --context <ctx> get nodes

# View resource usage
kubectl --context <ctx> -n <ns> top pods

Change Operations (must confirm first)

All change operations must show the user the command to be executed; execution only after user confirmation.

# Scale
kubectl --context <ctx> -n <ns> scale deployment <name> --replicas=<N>

# Rolling restart
kubectl --context <ctx> -n <ns> rollout restart deployment <name>

# Apply configuration
kubectl --context <ctx> -n <ns> apply -f <file>

# Edit ConfigMap
kubectl --context <ctx> -n <ns> edit configmap <name>

# Delete resource (high risk)
kubectl --context <ctx> -n <ns> delete <resource> <name>

Troubleshooting Operations

# Enter a container
kubectl --context <ctx> -n <ns> exec -it <pod-name> -- /bin/sh

# Port forwarding
kubectl --context <ctx> -n <ns> port-forward <pod-name> <local>:<remote>

# View all containers in a Pod
kubectl --context <ctx> -n <ns> get pod <pod-name> -o jsonpath='{.spec.containers[*].name}'

Step 5: Confirm Results

  • View operations: Display output directly; format or summarize as needed
  • Change operations: Verify results after execution (e.g., check Pod count after scale, check rollout status after restart)
  • On failure: Display error information, analyze the cause, and suggest a fix

Security Rules

  1. Prohibit executing delete namespace on prod clusters
  2. Prohibit executing kubectl drain unless the user explicitly requests and confirms
  3. All change operations must show the complete command before execution and wait for user confirmation
  4. Batch operations (e.g., scaling multiple deployments simultaneously) must list each one and confirm
  5. Prefer using --context parameter to specify the cluster; avoid kubectl config use-context which modifies global state

Examples

Bad

1. Executing changes without confirmation

User: Help me restart payment-service on Production
AI: Rolling restart executed.

Problem: Change operation executed without user confirmation.

2. Modifying global context

kubectl config use-context my-prod-context
kubectl -n prod get pods

Problem: Modified the global context, potentially affecting the user's other terminal sessions. Should use the --context parameter.

3. Executing high-risk operations on prod without confirmation

kubectl --context my-prod-context delete namespace monitoring

Problem: Deleting an entire namespace on a production cluster without confirmation.

Good

1. Confirming before changes

User: Help me restart payment-service on Prod

AI: I will execute the following on the Prod cluster:
    - Cluster: prod-1 (my-prod-context)
    - Namespace: prod
    - Command: kubectl --context my-prod-context -n prod rollout restart deployment payment-service

    Confirm execution?

User: Confirmed

AI: Executed. Checking rollout status...
    deployment "payment-service" successfully rolled out

2. Guiding configuration after environment check

User: Help me check the Pod status on Staging

AI: Let me check the kubectl environment...
    kubectl is installed (v1.28.2)
    But the Staging context was not found. Please run the following command to configure it:

    aws eks update-kubeconfig --name staging-cluster --region <your-region>

    Let me know when configuration is complete, and I will continue.

3. Troubleshooting flow

User: The api-gateway on Staging keeps CrashLoopBackOff

AI: Let me investigate:

    1. Checking Pod status...
       NAME                          READY   STATUS             RESTARTS   AGE
       api-gateway-5d4f6c7b8-x2k9z  0/1     CrashLoopBackOff   5          10m

    2. Checking recent logs...
       Error: failed to connect to database at <ip>:5432

    3. Checking events...
       Back-off restarting failed container

    Analysis: Pod cannot connect to database at <ip>:5432. Recommended checks:
    - Is the database service running normally
    - Is a NetworkPolicy blocking the connection
    - Is the database address correct in the ConfigMap/Secret

Exemptions

ScenarioCondition
Emergency incident recoveryUser explicitly states an emergency — confirmation steps may be skipped
Non-managed clustersUser provides a context not in the Setup list above

Exemption method: /override skill=k8s-ops reason="emergency incident recovery"

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most containers cloud skills give in ~1.8k tokens

Counted across 607 of the 705 authors here whose files we hold, read 2026-09-06

  • Run as non-root userin 34 of 607, across 27 files
  • Use multi-stage buildsin 29 of 607
  • Set resource requests and limitsin 24 of 607, across 20 files
  • Configure liveness and readiness probesin 18 of 607, across 14 files
  • Use named volumes for persistent datain 14 of 607, across 9 files
  • Pin base image versionsin 14 of 607
  • Set up environment variablesin 14 of 607, across 10 files
  • Pin provider versionsin 14 of 607
  • Apply least privilege RBAC permissionsin 10 of 607, across 7 files
  • Create a dockerignore filein 10 of 607
  • Use remote state with lockingin 9 of 607
  • Pin base images by digestin 9 of 607, across 8 files

Said here and by no other author read

  • Run kubectl version and kubectl config get-contexts first
  • Confirm operation target cluster and namespace
  • Specify the cluster using the context parameter
  • Show change operations and wait for user confirmation
  • Verify results after executing change operations

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.