Infra kit.domain.k8s doctor
Skill huyngopt1994/infras-kit/infras-kit-plugin/skills/infra-kit.domain.k8s-doctor
Infras Kit Provide a Kit For All Infras Working
npx -y skills add huyngopt1994/infras-kit --skill infra-kit.domain.k8s-doctorAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
- 4 stars4 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Debug Kubernetes workload, networking, routing, and rollout issues using read-only kubectl flows across pods, services, endpoints, HTTPRoute, ingress, and gateway resources.
SKILL.md
7.4 KB, as published. Nobody here has run it
Kubernetes Debug
Use this skill when the user is troubleshooting Kubernetes runtime behavior such as failing Pods, broken Service routing, missing endpoints, HTTPRoute or Ingress issues, rollout failures, or namespace-scoped application reachability problems.
This skill is read-only by default. Prefer kubectl get, kubectl describe, kubectl logs, and kubectl events style inspection first.
This skill does not perform write actions. It is limited to read-only troubleshooting.
Never run mutating commands from this skill, including:
kubectl deletekubectl patchkubectl scalekubectl rollout restartkubectl apply- changing context or namespace defaults
For interactive troubleshooting commands that may still be useful during investigation, ask for approval step by step before each individual command:
kubectl execkubectl debugkubectl cpkubectl port-forward
When checking namespaced resources, prefer explicit namespace scoping on every command:
kubectl get pod -n <namespace>
kubectl describe svc -n <namespace> <name>
kubectl logs -n <namespace> <pod>
Do not rely on the current namespace implicitly when debugging user workloads.
Outcomes
- Isolate whether the failure is in the workload, Service wiring, or traffic layer above it
- Explain the concrete broken hop in the request path
- Keep investigation safe by default through read-only inspection
- Leave the user with the exact evidence, command trail, and likely next fix
Where This Fits In The Flow
- Use during intake/triage when the ticket is “something is broken in k8s” and you need a read-only diagnosis.
- Feed the findings back into
infra-kit.workflowartifacts (spec.md/plan.md/tasks.md) before making changes.
Workflow
- Confirm the namespace, workload name, traffic entrypoint, and observed symptom before digging deeper.
- Start at the backend and work outward, bottom to top:
- Pod
- controller (
Deployment,StatefulSet,DaemonSet,Job) - Service
- EndpointSlice
- HTTPRoute or Ingress
- Gateway or ingress controller exposure when relevant
- Inspect Pods first:
- readiness and liveness failures
- restart counts
- image pull issues
- scheduling failures
- recent logs
- namespace events
- if the Pod is in
Error,CrashLoopBackOff,ImagePullBackOff, or another unhealthy state, capture bothdescribe podoutput and error logs into text files for later analysis
- Check the owning controller next for rollout health, replica mismatch, and selector correctness.
- Check the Service:
- selector matches Pod labels
- target port maps to a real container port
- type and annotations fit the intended exposure model
- Check
EndpointSliceobjects, not just Service existence. A healthy Service with zero or wrong endpoints is a common break point. - If Gateway API is used, inspect:
HTTPRoute- parent refs
- backend refs
- route status conditions
- Gateway listener attachment
- If Ingress is used, inspect:
- rules and path matching
- backend Service and port references
- ingress class
- controller events/status
- Check namespace guardrails that commonly block healthy workloads:
NetworkPolicyfor denied east-west or ingress-controller trafficResourceQuotafor failed scheduling or rejected createsLimitRangefor implicit resource defaults or invalid workload sizing
- Only after the path is mapped should you suggest interactive checks such as
execorport-forward. - Do not run those commands automatically. Present the exact next command, explain why it is needed, and wait for explicit user approval before each step.
- If a likely fix requires patching, deleting, restarting, scaling, or applying resources, stop at the diagnosis and tell the user what change is recommended rather than performing it.
Hallucination Guardrails
- Only report health or routing conclusions that are backed by concrete
kubectloutput; cite the exact command (and captured file when applicable) so the user can trace every statement to evidence. - If a resource, namespace, or controller cannot be found, say so explicitly instead of assuming its state; ask the user for corrected names when needed.
- When permissions, kubeconfig access, or tooling limitations block a command, document the blocker and keep the analysis scoped to the data that was actually retrievable.
- Separate read-only evidence from recommended write actions clearly so users understand no mutation occurred and can decide whether to run the fix themselves.
Command Pattern
Prefer a read-only sequence like:
kubectl get pods -n <namespace> -o wide
kubectl describe pod -n <namespace> <pod>
kubectl logs -n <namespace> <pod> --container <container> --tail=200
kubectl get deploy -n <namespace>
kubectl describe deploy -n <namespace> <deploy>
kubectl get svc -n <namespace>
kubectl describe svc -n <namespace> <service>
kubectl get endpointslice -n <namespace>
kubectl describe endpointslice -n <namespace> <slice>
kubectl get networkpolicy -n <namespace>
kubectl describe networkpolicy -n <namespace> <policy>
kubectl get resourcequota -n <namespace>
kubectl describe resourcequota -n <namespace>
kubectl get limitrange -n <namespace>
kubectl describe limitrange -n <namespace>
kubectl get httproute -n <namespace>
kubectl describe httproute -n <namespace> <route>
kubectl get ingress -n <namespace>
kubectl describe ingress -n <namespace> <ingress>
kubectl get events -n <namespace> --sort-by=.lastTimestamp
When the Pod is unhealthy, prefer capturing artifacts as files:
kubectl describe pod -n <namespace> <pod> > describe_pod.txt
kubectl logs -n <namespace> <pod> --container <container> --previous --tail=200 > log_error.txt
If --previous is not applicable, capture the current container logs instead.
If the issue is cross-namespace or controller-level, widen scope deliberately and say why.
Review Priorities
When debugging Kubernetes, check in this order:
- Pod health and recent events
- Controller rollout status and selector consistency
- Service selector and port wiring
- EndpointSlice population and backend IP/port correctness
- NetworkPolicy, ResourceQuota, and LimitRange side effects
- HTTPRoute or Ingress routing rules and status conditions
- Gateway, ingress controller, or external exposure layer
Read-Only Rules
- Default to inspection commands only.
- Never treat
exec,debug,cp, orport-forwardas implicitly allowed. - Ask before entering containers or creating debug containers.
- Ask before port-forwarding, even if it seems harmless.
- Do not perform write, restart, scaling, delete, patch, or apply actions from this skill.
- Make the exact command explicit before requesting approval for each non-read-only step.
- Approval is per step, not blanket approval for a whole debugging session.
Bundled helper:
bash scripts/collect_pod_debug.sh <namespace> <pod> [container]
Delivery Standard
Always leave the user with:
- the exact hop where traffic or workload health breaks
- the commands used to prove it
- the key evidence from Pods, Services, endpoints, routes, and namespace guardrails
- the least invasive recommended fix, clearly separated from read-only findings