agentsclimarketplace

Containerization

Skill MARUCIE/openclaw-foundry/web/public/packs/spellbook-security-auditor/skills/containerization

Use when writing Dockerfiles, setting up docker-compose for local dev, configuring Kubernetes resources (Deployment, Service, Ingress, HPA), sizing pod resource limits, or packaging a service with Helm.From its SKILL.md

Install
npx -y skills add MARUCIE/openclaw-foundry --skill containerization

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

17.6 KB, ~4.7k tokens by cl100k_base, as published. Nobody here has run it

是什么

这是一份容器化安全规范,从审计视角覆盖 Dockerfile 安全基线、Kubernetes(K8s,容器编排平台)RBAC(基于角色的访问控制)、镜像签名、运行时隔离、网络策略等关键控制点,让团队上容器时不留下 root 运行、特权挂载、未签名镜像等高危漏洞。

怎么用

  1. 审计 Dockerfile 时按非 root 用户、最小基础镜像、构建参数清理三条基线逐项检查,发现违规直接打回。
  2. 审计 K8s 部署清单时重点看 securityContext(安全上下文)、resource limits、ReadOnlyRootFilesystem 三项配置是否合规。
  3. 镜像入仓前必须过 Trivy 或 Snyk 等漏洞扫描,按文档约定的 CVE(公共漏洞披露)阈值卡严重等级。
  4. 集群层面按 RBAC 最小权限模板配 Role 与 RoleBinding,避免 default ServiceAccount 拿到全局权限。
  5. 网络层启用 NetworkPolicy(网络策略)默认拒绝,按白名单逐步放开服务间通信,缩小横向移动面。

架构图

flowchart LR
    A[Dockerfile 审计] --> B[镜像漏洞扫描]
    B --> C[签名入仓]
    C --> D[K8s RBAC 校验]
    D --> E[NetworkPolicy 隔离]
    E --> F[运行时监控]

Containerization

Package and orchestrate services with Docker and Kubernetes using production-ready patterns for image builds, local development, cluster deployments, and autoscaling.

When to Activate

  • Writing or reviewing a Dockerfile for a service
  • Setting up docker-compose for local development
  • Writing Kubernetes manifests for a new service
  • Deploying to a Kubernetes cluster
  • Sizing CPU/memory requests and limits for a pod
  • Setting up a Helm chart
  • Optimizing Docker image build time or image size

Dockerfile Best Practices

Multi-Stage Builds

Multi-stage builds keep build-time tools out of the final image, reducing attack surface and image size.

Python

# Stage 1: build dependencies
FROM python:3.12-slim AS builder
WORKDIR /app
COPY pyproject.toml uv.lock ./
RUN pip install uv && uv sync --frozen --no-dev

# Stage 2: runtime image
FROM python:3.12-slim AS runtime
WORKDIR /app
COPY --from=builder /app/.venv /app/.venv
COPY src/ ./src/
ENV PATH="/app/.venv/bin:$PATH"
USER 1000:1000
EXPOSE 8000
HEALTHCHECK --interval=30s --timeout=5s CMD curl -f http://localhost:8000/health || exit 1
ENTRYPOINT ["python", "-m", "uvicorn", "src.main:app", "--host", "0.0.0.0", "--port", "8000"]

Node.js / TypeScript

FROM node:20-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci --include=dev
COPY . .
RUN npm run build

FROM node:20-alpine AS runtime
WORKDIR /app
COPY package*.json ./
RUN npm ci --omit=dev
COPY --from=builder /app/dist ./dist
USER node
EXPOSE 3000
HEALTHCHECK --interval=30s --timeout=5s CMD wget -qO- http://localhost:3000/health || exit 1
CMD ["node", "dist/index.js"]

Go (smallest possible image)

FROM golang:1.22-alpine AS builder
WORKDIR /app
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 GOOS=linux go build -ldflags="-w -s" -o /app/server ./cmd/server

FROM gcr.io/distroless/static-debian12 AS runtime
COPY --from=builder /app/server /server
EXPOSE 8080
USER nonroot:nonroot
ENTRYPOINT ["/server"]

Layer Caching Order

Docker caches each layer. If a layer's input changes, all subsequent layers are invalidated. Copy dependency manifests before source code so that the expensive dependency-install step is only re-run when dependencies actually change, not on every source edit.

# BAD — source copy before dependency install; every code change busts the npm ci cache
COPY . .
RUN npm ci

# GOOD — copy only the manifest first; npm ci cache survives source-only changes
COPY package*.json ./
RUN npm ci
COPY . .

The same principle applies to all runtimes:

RuntimeCopy firstThen install
Python (uv)pyproject.toml uv.lockuv sync --frozen
Nodepackage*.jsonnpm ci
Gogo.mod go.sumgo mod download

.dockerignore

Always create a .dockerignore at the repo root. Files excluded here are never sent to the Docker build context, speeding up builds and preventing accidental secret leaks.

.git
.gitignore
.env
*.env
__pycache__
*.pyc
node_modules
dist
build
.pytest_cache
.coverage
*.log
README.md

Base Image Selection

ImageSizeVulnerability surfaceBest for
ubuntu:22.04~80 MBHighDev/debug only
debian:bookworm-slim~75 MBMediumGeneral purpose
python:3.12-slim~150 MBMediumPython apps
node:20-alpine~170 MBLowNode apps
alpine:3.19~7 MBVery lowCustom builds
gcr.io/distroless/static-debian12~2 MBMinimalGo static binaries
gcr.io/distroless/python3-debian12~80 MBMinimalPython (no shell!)

Pin to digest for reproducibility in production:

FROM python:3.12-slim@sha256:abc123...

docker-compose for Local Development

Use docker-compose for wiring together the application and its backing services locally. Keep secrets in .env (gitignored) and load them via env_file.

version: '3.9'
services:
  app:
    build:
      context: .
      target: runtime  # use multi-stage target
    ports:
      - "8000:8000"
    env_file:
      - .env
    environment:
      DATABASE_URL: postgresql://user:pass@db:5432/appdb
    depends_on:
      db:
        condition: service_healthy
      redis:
        condition: service_healthy
    volumes:
      - ./src:/app/src  # hot-reload in dev
    profiles:
      - dev

  db:
    image: postgres:16-alpine
    environment:
      POSTGRES_USER: user
      POSTGRES_PASSWORD: pass
      POSTGRES_DB: appdb
    volumes:
      - postgres_data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U user -d appdb"]
      interval: 5s
      timeout: 3s
      retries: 5

  redis:
    image: redis:7-alpine
    healthcheck:
      test: ["CMD", "redis-cli", "ping"]
      interval: 5s
      timeout: 3s
      retries: 5

volumes:
  postgres_data:

Profiles. Use profiles to mark optional services. Start only what you need:

docker compose --profile dev up

Override file. Create docker-compose.override.yml for local-only tweaks (e.g., mounting a local SDK, exposing extra ports). Add it to .gitignore so it never ships.

# docker-compose.override.yml (gitignored)
services:
  app:
    environment:
      DEBUG: "true"
    volumes:
      - ../my-local-sdk:/app/vendor/sdk

Kubernetes Core Resources

Deployment

apiVersion: apps/v1
kind: Deployment
metadata:
  name: payment-service
  namespace: production
spec:
  replicas: 3
  selector:
    matchLabels:
      app: payment-service
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0  # zero-downtime: always have full capacity during rollout
  template:
    metadata:
      labels:
        app: payment-service
    spec:
      containers:
        - name: payment-service
          image: ghcr.io/org/payment-service:abc123
          ports:
            - containerPort: 8000
          resources:
            requests:
              cpu: 100m
              memory: 128Mi
            limits:
              cpu: 500m
              memory: 256Mi
          envFrom:
            - configMapRef:
                name: payment-service-config
          env:
            - name: DATABASE_URL
              valueFrom:
                secretKeyRef:
                  name: payment-service-secrets
                  key: database-url
          readinessProbe:
            httpGet:
              path: /health/ready
              port: 8000
            initialDelaySeconds: 5
            periodSeconds: 10
          livenessProbe:
            httpGet:
              path: /health/live
              port: 8000
            initialDelaySeconds: 30
            periodSeconds: 30
            failureThreshold: 3

Service

apiVersion: v1
kind: Service
metadata:
  name: payment-service
  namespace: production
spec:
  selector:
    app: payment-service
  ports:
    - port: 80
      targetPort: 8000
  type: ClusterIP  # internal only; use LoadBalancer for external

Service type reference:

TypeAccessibilityUse case
ClusterIPCluster-internal onlyInternal services
NodePortExternal via node IP:portDev/testing
LoadBalancerExternal via cloud LBProd external services
ExternalNameDNS aliasOff-cluster services

Ingress

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: payment-service
  annotations:
    cert-manager.io/cluster-issuer: letsencrypt-prod
spec:
  tls:
    - hosts:
        - api.example.com
      secretName: api-tls-cert
  rules:
    - host: api.example.com
      http:
        paths:
          - path: /payments
            pathType: Prefix
            backend:
              service:
                name: payment-service
                port:
                  number: 80

ConfigMap and Secret

apiVersion: v1
kind: ConfigMap
metadata:
  name: payment-service-config
data:
  LOG_LEVEL: "INFO"
  APP_ENV: "production"
---
apiVersion: v1
kind: Secret
metadata:
  name: payment-service-secrets
type: Opaque
data:
  database-url: <base64-encoded-value>  # echo -n "postgresql://..." | base64

Never commit Secrets to git. Use one of:

  • sealed-secrets — encrypts Secret with a cluster key; safe to commit
  • external-secrets-operator — syncs from AWS Secrets Manager / GCP Secret Manager / Vault
  • Vault agent injection — sidecar writes secrets to an in-memory volume

Resource Management

Requests vs Limits

FieldPurposeWhat happens when exceeded
requests.cpuGuaranteed CPU; used for schedulingN/A — node selection only
limits.cpuMaximum CPU the container may useThrottled (not killed)
requests.memoryGuaranteed memory; used for schedulingN/A — node selection only
limits.memoryMaximum memory the container may useOOMKilled (pod restarted)
  • Requests are what the scheduler uses to decide which node a pod lands on.
  • Limits are enforced at runtime by the kernel cgroup.
  • Setting limits.memory without headroom above requests.memory invites spurious OOMKills under GC pressure.

QoS Classes

Kubernetes assigns a QoS class based on how requests and limits are configured. Higher QoS = last to be evicted under node memory pressure.

ClassConditionEviction priority
Guaranteedrequests == limits for every resourceLast to be evicted
Burstablerequests < limits for at least one resourceMiddle
BestEffortNo requests or limits set at allFirst to be evicted

For critical services, set requests == limits to achieve Guaranteed QoS. For batch jobs or low-priority workers, Burstable is acceptable.

Health Probes

ProbeWhat it checksFailure action
readinessProbeIs pod ready to receive traffic?Remove from Service endpoints
livenessProbeIs pod alive?Restart pod
startupProbeHas pod finished starting? (slow-starting apps)Replaces liveness until started

Common mistake: setting livenessProbe.failureThreshold too low (e.g., 2 with periodSeconds: 10) causes restart loops during slow GC pauses or transient DB query spikes. For most services, failureThreshold: 3 with periodSeconds: 30 is a safer baseline.

# BAD — aggressive liveness; 20 seconds of GC pause triggers restart
livenessProbe:
  httpGet:
    path: /health/live
    port: 8000
  periodSeconds: 10
  failureThreshold: 2

# GOOD — tolerates 90 seconds of unresponsiveness before restarting
livenessProbe:
  httpGet:
    path: /health/live
    port: 8000
  initialDelaySeconds: 30
  periodSeconds: 30
  failureThreshold: 3

HorizontalPodAutoscaler

HPA scales the replica count based on observed metrics. Requires the metrics-server addon (or custom metrics adapter for non-CPU metrics).

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: payment-service
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: payment-service
  minReplicas: 2
  maxReplicas: 20
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70

HPA only works correctly when requests.cpu is set — it calculates utilization as actual / request. Without a CPU request, HPA cannot compute a meaningful ratio and will not scale.

Helm Basics

Helm packages Kubernetes manifests into versioned, parameterised charts.

Chart structure:

mychart/
├── Chart.yaml          # metadata (name, version, appVersion)
├── values.yaml         # default values
├── templates/
│   ├── deployment.yaml
│   ├── service.yaml
│   └── _helpers.tpl    # reusable template snippets

Key commands:

helm install my-release ./mychart --values prod-values.yaml
helm upgrade my-release ./mychart --values prod-values.yaml
helm rollback my-release 1
helm diff upgrade my-release ./mychart  # requires helm-diff plugin

values.yaml pattern — expose only what varies per environment:

# values.yaml
image:
  repository: ghcr.io/org/payment-service
  tag: latest  # override per environment with --set image.tag=abc123

replicaCount: 3

resources:
  requests:
    cpu: 100m
    memory: 128Mi
  limits:
    cpu: 500m
    memory: 256Mi

Override at release time without editing the chart:

helm upgrade my-release ./mychart \
  --values prod-values.yaml \
  --set image.tag=abc123

Security Context

Apply securityContext at both the pod level and the container level. The settings below satisfy most CIS Kubernetes Benchmark requirements.

# pod-level: applies to all containers in the pod
spec:
  securityContext:
    runAsNonRoot: true
    runAsUser: 1000
    runAsGroup: 1000
    seccompProfile:
      type: RuntimeDefault

  containers:
    - name: app
      # container-level: overrides/extends pod-level settings
      securityContext:
        readOnlyRootFilesystem: true
        allowPrivilegeEscalation: false
        capabilities:
          drop: ["ALL"]

If readOnlyRootFilesystem: true causes write errors, mount an explicit emptyDir volume for the specific writable path rather than disabling the restriction:

# BAD — disabling readOnly to allow one directory to be writable
securityContext:
  readOnlyRootFilesystem: false

# GOOD — keep readOnly, mount emptyDir for the specific path
securityContext:
  readOnlyRootFilesystem: true
volumeMounts:
  - name: tmp
    mountPath: /tmp
volumes:
  - name: tmp
    emptyDir: {}

See also: ci-cd, deployment-strategies, security

Red Flags

  • Running the container process as root — a process breakout inside the container inherits root on the host; always set USER 1000:1000 (or USER node) in the final Dockerfile stage
  • Copying the entire build context before installing dependenciesCOPY . . before RUN npm ci busts the layer cache on every source change, making every build a full cold install
  • Using latest tag in Kubernetes manifestsimagePullPolicy: Always with latest means different nodes may pull different images across a rolling deploy; pin to a commit SHA or versioned tag
  • Setting limits.memory equal to requests.memory with no headroom — a JVM or Python GC spike briefly exceeds the request value; without headroom the pod is OOMKilled and restarted during normal operation
  • Liveness probe with low failureThreshold (1–2) and short periodSeconds (5–10) — a slow GC pause or cold DB query triggers an unnecessary pod restart loop; use failureThreshold: 3 and periodSeconds: 30 as a baseline
  • Storing Kubernetes Secrets as plain base64 in git — base64 is not encryption; use Sealed Secrets or external-secrets-operator so plaintext values never enter version control
  • No .dockerignore — the full build context (including .git, node_modules, .env) is sent to the Docker daemon on every build, leaking secrets and adding seconds of unnecessary transfer
  • HPA configured without CPU requests set — HPA calculates utilization as actual / request; a missing request means the denominator is undefined and the autoscaler cannot make scaling decisions

Checklist

  • Multi-stage Dockerfile used — build tools not in final image
  • Source files copied after dependency files (layer cache optimization)
  • Non-root user set in Dockerfile (USER 1000:1000 or USER node)
  • .dockerignore excludes .git, .env, node_modules, __pycache__
  • HEALTHCHECK instruction defined in Dockerfile
  • All pods have readinessProbe and livenessProbe configured
  • CPU and memory requests and limits set on every container
  • Secrets stored in Kubernetes Secrets (or external secrets manager), not ConfigMaps
  • securityContext sets runAsNonRoot: true and allowPrivilegeEscalation: false
  • HPA configured for services with variable load
  • maxUnavailable: 0 in rolling update strategy for zero-downtime deployments
  • Image tagged with commit SHA, not latest

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,851. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.