Claude Skill

infra-containers-kubernetes

Kubernetes manifests, Helm charts, Kustomize overlays, and resource patterns

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_infra-containers-kubernetes_skills_infra-containers-kubernetes-3a51ef5.zip · 18 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/infra-containers-kubernetes/skills/infra-containers-kubernetes
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Kubernetes Patterns

Quick Guide: Declarative YAML manifests for Kubernetes workloads. Use apps/v1 Deployments with resource requests/limits, health probes, and Pod Security Standards (restricted profile). Helm for templated multi-environment releases. Kustomize for patch-based overlays without templating. Always set securityContext (runAsNonRoot, drop ALL capabilities, readOnlyRootFilesystem), resource requests/limits on every container, and liveness/readiness probes on every pod.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST set resource requests AND limits on every container -- pods without requests are unschedulable under resource pressure and pods without limits can OOM-kill neighbors)

(You MUST set securityContext with runAsNonRoot: true, allowPrivilegeEscalation: false, drop ALL capabilities, and readOnlyRootFilesystem: true on every container)

(You MUST define both liveness and readiness probes -- without readiness probes, traffic reaches unready pods; without liveness probes, hung processes are never restarted)

(You MUST use the current stable apiVersion for each resource -- apps/v1 for Deployments, networking.k8s.io/v1 for Ingress, autoscaling/v2 for HPA, policy/v1 for PDB)

</critical_requirements>


Examples

  • Core Manifests - Deployments, Services, Ingress, ConfigMaps, Secrets, Namespaces
  • Helm Charts - Chart structure, values.yaml, templates, helpers, dependencies
  • Operations - HPA scaling, RBAC, health checks, resource limits, PDB, NetworkPolicy
  • Quick Reference - API versions, kubectl commands, label conventions, decision frameworks

Auto-detection: Kubernetes, kubectl, k8s, Deployment, Service, Ingress, ConfigMap, Secret, HPA, HorizontalPodAutoscaler, Helm, helm chart, Kustomize, kustomization, RBAC, Role, ClusterRole, PodDisruptionBudget, NetworkPolicy, Pod, StatefulSet, DaemonSet, CronJob, Job, PersistentVolumeClaim, apiVersion, kind, metadata, spec

When to use:

  • Writing Kubernetes Deployment, Service, Ingress, or other resource manifests
  • Creating Helm charts for templated multi-environment releases
  • Building Kustomize overlays for environment-specific patches
  • Configuring RBAC roles and bindings for least-privilege access
  • Setting up HPA autoscaling, PDB, or resource limits
  • Defining health checks (liveness, readiness, startup probes)
  • Managing ConfigMaps, Secrets, and environment configuration
  • Writing NetworkPolicy for pod-to-pod traffic control

When NOT to use:

  • Container image building (use a containerization skill)
  • CI/CD pipeline definitions (use a CI/CD skill)
  • Infrastructure provisioning (use an IaC tool)
  • Service mesh configuration beyond basic Kubernetes resources
  • Managed Kubernetes cluster setup (cloud provider control plane configuration)

Key patterns covered:

  • Deployment with security context, resource limits, and probes
  • Service types (ClusterIP, NodePort, LoadBalancer, Headless)
  • Ingress with TLS and path routing (networking.k8s.io/v1)
  • ConfigMap and Secret management (envFrom, volume mounts)
  • HPA autoscaling (autoscaling/v2 metrics array)
  • RBAC (Role, ClusterRole, RoleBinding, ServiceAccount)
  • Pod Security Standards (restricted profile)
  • Helm chart structure, values, templates, and helpers
  • Kustomize base/overlay pattern with strategic merge patches
  • PodDisruptionBudget for safe rollouts
  • NetworkPolicy for pod traffic isolation



<decision_framework>

Decision Framework

Helm vs Kustomize

Do you distribute charts to external teams?
  +-- YES --> Helm (package management, versioned releases)
  +-- NO  --> Do you need Go templating / conditionals?
      +-- YES --> Helm (parameterized templates)
      +-- NO  --> Do you prefer plain YAML with patches?
          +-- YES --> Kustomize (built into kubectl, no templating)
          +-- NO  --> Either works. Choose team familiarity.

Workload Type Selection

What is the workload?
  +-- Stateless HTTP/API server? --> Deployment
  +-- Needs stable network identity / ordered startup? --> StatefulSet
  +-- Must run on every node (logging, monitoring)? --> DaemonSet
  +-- One-time batch job? --> Job
  +-- Scheduled recurring job? --> CronJob

Service Type Selection

Who needs to reach this service?
  +-- Other pods in the cluster? --> ClusterIP (default)
  +-- External traffic via HTTP/HTTPS? --> ClusterIP + Ingress
  +-- External TCP/UDP without Ingress? --> LoadBalancer
  +-- Direct node access (dev/test)? --> NodePort
  +-- StatefulSet pod discovery? --> Headless (clusterIP: None)

Resource Requests vs Limits

What QoS class do you need?
  +-- Guaranteed (critical workloads) --> requests == limits
  +-- Burstable (typical workloads) --> requests < limits
  +-- BestEffort (batch, non-critical) --> no requests/limits (NOT recommended for production)

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Missing resources.requests and resources.limits on containers -- unschedulable under pressure, can OOM-kill neighbors
  • Running as root (no securityContext.runAsNonRoot: true) -- container escape gives host root access
  • Using :latest image tag -- non-deterministic deployments, impossible to roll back to known version
  • Missing liveness/readiness probes -- hung processes never restart, traffic hits unready pods
  • Using deprecated apiVersions (extensions/v1beta1, autoscaling/v2beta2) -- will fail on modern clusters
  • Wildcard RBAC rules (verbs: ["*"], resources: ["*"]) -- violates least privilege, security risk
  • Storing actual secrets in manifests committed to Git -- use sealed secrets or external secret operators

Medium Priority Issues:

  • automountServiceAccountToken not set to false -- every pod gets a token that can access the API server
  • No PodDisruptionBudget -- cluster upgrades or node drains can terminate all replicas simultaneously
  • No NetworkPolicy -- all pods can communicate with all other pods by default
  • Using spec.targetCPUUtilizationPercentage in HPA -- deprecated in autoscaling/v2, use metrics array
  • Missing revisionHistoryLimit on Deployments -- unlimited old ReplicaSets consume etcd storage
  • ConfigMap/Secret changes not triggering rollout -- pods keep stale config until manually restarted

Gotchas & Edge Cases:

  • readOnlyRootFilesystem: true breaks apps that write to /tmp -- add an emptyDir volume mount for /tmp
  • requests.memory too low causes OOMKill; limits.cpu too low causes CPU throttling (latency spikes, not kills)
  • Ingress pathType: Prefix matches /api AND /api-docs -- use pathType: Exact for precise matching or add trailing /
  • Secrets are base64-encoded, NOT encrypted -- anyone with RBAC read access to secrets can decode them
  • HPA and manual replicas in Deployment conflict -- remove spec.replicas from manifests when HPA is active
  • kubectl apply vs kubectl create -- apply is declarative and idempotent; create fails if resource exists
  • Kustomize commonLabels adds labels to selectors too -- changing them on existing Deployments breaks selector immutability
  • Helm lookup function doesn't work during helm template (no cluster access) -- only works during helm install/upgrade
  • Pod Security Admission enforces at namespace level via labels -- pods in unlabeled namespaces get the default (usually privileged)

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST set resource requests AND limits on every container -- pods without requests are unschedulable under resource pressure and pods without limits can OOM-kill neighbors)

(You MUST set securityContext with runAsNonRoot: true, allowPrivilegeEscalation: false, drop ALL capabilities, and readOnlyRootFilesystem: true on every container)

(You MUST define both liveness and readiness probes -- without readiness probes, traffic reaches unready pods; without liveness probes, hung processes are never restarted)

(You MUST use the current stable apiVersion for each resource -- apps/v1 for Deployments, networking.k8s.io/v1 for Ingress, autoscaling/v2 for HPA, policy/v1 for PDB)

Failure to follow these rules will produce insecure pods with root access, unschedulable workloads, unmonitored health, and manifests that fail on modern clusters.

</critical_reminders>

Files (skills)
  • examples
    • core.md 10.4 KB
      # Kubernetes -- Core Manifest Examples
      
      > Deployments, Services, Ingress, ConfigMaps, Secrets, Namespaces, and Kustomize overlays. Reference from [SKILL.md](../SKILL.md).
      
      **Related examples:**
      
      - [helm.md](helm.md) - Helm chart structure, values, templates, helpers
      - [operations.md](operations.md) - HPA, RBAC, health checks, resource limits, PDB, NetworkPolicy
      
      ---
      
      ## Example 1: Complete Application Stack
      
      A production Deployment with matching Service, Ingress, ConfigMap, and Namespace.
      
      ### Namespace with Pod Security
      
      ```yaml
      apiVersion: v1
      kind: Namespace
      metadata:
        name: app
        labels:
          pod-security.kubernetes.io/enforce: restricted
          pod-security.kubernetes.io/audit: restricted
          pod-security.kubernetes.io/warn: restricted
      ```
      
      ### Deployment
      
      ```yaml
      apiVersion: apps/v1
      kind: Deployment
      metadata:
        name: api-server
        namespace: app
        labels:
          app.kubernetes.io/name: api-server
          app.kubernetes.io/component: backend
          app.kubernetes.io/part-of: my-platform
      spec:
        replicas: 3
        revisionHistoryLimit: 5
        strategy:
          type: RollingUpdate
          rollingUpdate:
            maxUnavailable: 1
            maxSurge: 1
        selector:
          matchLabels:
            app.kubernetes.io/name: api-server
        template:
          metadata:
            labels:
              app.kubernetes.io/name: api-server
              app.kubernetes.io/component: backend
            annotations:
              # Force rollout when ConfigMap changes
              checksum/config: '{{ include (print $.Template.BasePath "/configmap.yaml") . | sha256sum }}'
          spec:
            serviceAccountName: api-server
            automountServiceAccountToken: false
            terminationGracePeriodSeconds: 30
            securityContext:
              runAsNonRoot: true
              runAsUser: 1000
              runAsGroup: 1000
              fsGroup: 2000
              seccompProfile:
                type: RuntimeDefault
            containers:
              - name: api-server
                image: registry.example.com/api-server:v1.2.3
                imagePullPolicy: IfNotPresent
                ports:
                  - name: http
                    containerPort: 3000
                    protocol: TCP
                envFrom:
                  - configMapRef:
                      name: api-config
                env:
                  - name: DATABASE_URL
                    valueFrom:
                      secretKeyRef:
                        name: api-secrets
                        key: DATABASE_URL
                securityContext:
                  allowPrivilegeEscalation: false
                  readOnlyRootFilesystem: true
                  runAsNonRoot: true
                  capabilities:
                    drop: ["ALL"]
                resources:
                  requests:
                    cpu: 250m
                    memory: 256Mi
                  limits:
                    cpu: "1"
                    memory: 512Mi
                livenessProbe:
                  httpGet:
                    path: /healthz
                    port: http
                  initialDelaySeconds: 15
                  periodSeconds: 20
                  timeoutSeconds: 5
                  failureThreshold: 3
                readinessProbe:
                  httpGet:
                    path: /ready
                    port: http
                  initialDelaySeconds: 5
                  periodSeconds: 10
                  timeoutSeconds: 3
                  failureThreshold: 3
                startupProbe:
                  httpGet:
                    path: /healthz
                    port: http
                  failureThreshold: 30
                  periodSeconds: 10
                volumeMounts:
                  - name: tmp
                    mountPath: /tmp
            volumes:
              - name: tmp
                emptyDir: {}
      ```
      
      **Why good:** Restricted pod security context, dedicated service account with token disabled, three probe types, `emptyDir` for `/tmp` (needed because `readOnlyRootFilesystem: true`), named port for probe references, ConfigMap checksum annotation triggers rollout on config changes, rolling update strategy with explicit bounds
      
      ### Service
      
      ```yaml
      apiVersion: v1
      kind: Service
      metadata:
        name: api-server
        namespace: app
        labels:
          app.kubernetes.io/name: api-server
      spec:
        type: ClusterIP
        selector:
          app.kubernetes.io/name: api-server
        ports:
          - name: http
            port: 80
            targetPort: http
            protocol: TCP
      ```
      
      ### ConfigMap and Secret
      
      ```yaml
      apiVersion: v1
      kind: ConfigMap
      metadata:
        name: api-config
        namespace: app
      data:
        LOG_LEVEL: "info"
        NODE_ENV: "production"
        MAX_CONNECTIONS: "100"
      ---
      apiVersion: v1
      kind: Secret
      metadata:
        name: api-secrets
        namespace: app
      type: Opaque
      stringData:
        DATABASE_URL: "postgresql://user:pass@db-svc:5432/app"
        API_KEY: "sk-live-abc123"
      ```
      
      ### Ingress with TLS
      
      ```yaml
      apiVersion: networking.k8s.io/v1
      kind: Ingress
      metadata:
        name: api-ingress
        namespace: app
        annotations:
          nginx.ingress.kubernetes.io/ssl-redirect: "true"
          nginx.ingress.kubernetes.io/proxy-body-size: "10m"
      spec:
        ingressClassName: nginx
        tls:
          - hosts:
              - api.example.com
            secretName: api-tls
        rules:
          - host: api.example.com
            http:
              paths:
                - path: /api
                  pathType: Prefix
                  backend:
                    service:
                      name: api-server
                      port:
                        number: 80
                - path: /health
                  pathType: Exact
                  backend:
                    service:
                      name: api-server
                      port:
                        number: 80
      ```
      
      **Why good:** `ingressClassName` field (not deprecated `kubernetes.io/ingress.class` annotation), TLS with secret reference, mixed `pathType` (Prefix for API routes, Exact for health)
      
      ---
      
      ## Example 2: Service Types
      
      ### NodePort (Development / Testing)
      
      ```yaml
      apiVersion: v1
      kind: Service
      metadata:
        name: api-nodeport
        namespace: dev
      spec:
        type: NodePort
        selector:
          app.kubernetes.io/name: api-server
        ports:
          - port: 80
            targetPort: 3000
            nodePort: 30080
      ```
      
      ### LoadBalancer (External Traffic)
      
      ```yaml
      apiVersion: v1
      kind: Service
      metadata:
        name: api-lb
        namespace: app
        annotations:
          # Cloud-provider-specific annotations go here
          service.beta.kubernetes.io/aws-load-balancer-scheme: internet-facing
      spec:
        type: LoadBalancer
        selector:
          app.kubernetes.io/name: api-server
        ports:
          - port: 443
            targetPort: 3000
            protocol: TCP
      ```
      
      ### Headless Service (StatefulSet Discovery)
      
      ```yaml
      apiVersion: v1
      kind: Service
      metadata:
        name: db-headless
        namespace: app
      spec:
        type: ClusterIP
        clusterIP: None
        selector:
          app.kubernetes.io/name: database
        ports:
          - port: 5432
            targetPort: 5432
      ```
      
      **Why headless:** Each pod gets a DNS record like `db-0.db-headless.app.svc.cluster.local`, enabling direct pod addressing for StatefulSets.
      
      ---
      
      ## Example 3: ConfigMap Injection Patterns
      
      ### envFrom (All Keys)
      
      ```yaml
      containers:
        - name: app
          envFrom:
            - configMapRef:
                name: api-config
            - secretRef:
                name: api-secrets
      ```
      
      ### Selective env (Specific Keys)
      
      ```yaml
      containers:
        - name: app
          env:
            - name: DB_HOST
              valueFrom:
                configMapKeyRef:
                  name: api-config
                  key: DB_HOST
            - name: DB_PASSWORD
              valueFrom:
                secretKeyRef:
                  name: api-secrets
                  key: DB_PASSWORD
      ```
      
      ### Volume Mount (Config Files)
      
      ```yaml
      containers:
        - name: app
          volumeMounts:
            - name: config-volume
              mountPath: /etc/app/config
              readOnly: true
      volumes:
        - name: config-volume
          configMap:
            name: app-file-config
            items:
              - key: app.conf
                path: app.conf
              - key: logging.conf
                path: logging.conf
      ```
      
      ---
      
      ## Example 4: Kustomize Base/Overlay
      
      ### Base
      
      ```yaml
      # k8s/base/kustomization.yaml
      apiVersion: kustomize.config.k8s.io/v1beta1
      kind: Kustomization
      resources:
        - deployment.yaml
        - service.yaml
        - configmap.yaml
      commonLabels:
        app.kubernetes.io/name: api-server
      ```
      
      ### Staging Overlay
      
      ```yaml
      # k8s/overlays/staging/kustomization.yaml
      apiVersion: kustomize.config.k8s.io/v1beta1
      kind: Kustomization
      resources:
        - ../../base
      namespace: staging
      patches:
        - path: replica-patch.yaml
      configMapGenerator:
        - name: api-config
          behavior: merge
          literals:
            - LOG_LEVEL=debug
            - NODE_ENV=staging
      ```
      
      ```yaml
      # k8s/overlays/staging/replica-patch.yaml
      apiVersion: apps/v1
      kind: Deployment
      metadata:
        name: api-server
      spec:
        replicas: 1
      ```
      
      ### Production Overlay
      
      ```yaml
      # k8s/overlays/production/kustomization.yaml
      apiVersion: kustomize.config.k8s.io/v1beta1
      kind: Kustomization
      resources:
        - ../../base
        - hpa.yaml
        - pdb.yaml
      namespace: production
      patches:
        - path: resource-patch.yaml
      configMapGenerator:
        - name: api-config
          behavior: merge
          literals:
            - LOG_LEVEL=warn
            - NODE_ENV=production
      secretGenerator:
        - name: api-secrets
          envs:
            - secrets.env
      ```
      
      ```yaml
      # k8s/overlays/production/resource-patch.yaml
      apiVersion: apps/v1
      kind: Deployment
      metadata:
        name: api-server
      spec:
        replicas: 3
        template:
          spec:
            containers:
              - name: api-server
                resources:
                  requests:
                    cpu: 500m
                    memory: 512Mi
                  limits:
                    cpu: "2"
                    memory: 1Gi
      ```
      
      **Why good:** Base manifests are valid YAML (no template syntax), overlays contain only diffs, `configMapGenerator` appends content hash to name (auto-triggers rollouts), `secretGenerator` keeps secrets out of base manifests
      
      **Apply:** `kubectl apply -k k8s/overlays/production/`
      
      ---
      
      ## Anti-Patterns
      
      ### Bad: Missing Security Context
      
      ```yaml
      # BAD: No security context -- runs as root by default
      apiVersion: apps/v1
      kind: Deployment
      metadata:
        name: api-server
      spec:
        template:
          spec:
            containers:
              - name: api-server
                image: api-server:latest
                ports:
                  - containerPort: 3000
      ```
      
      **Why bad:** Runs as root (default), no resource limits (can consume all node resources), `:latest` tag (non-deterministic), no probes (unmonitored health), no service account (uses default with API access)
      
      ### Bad: Overly Permissive RBAC
      
      ```yaml
      # BAD: Wildcard permissions
      apiVersion: rbac.authorization.k8s.io/v1
      kind: ClusterRole
      metadata:
        name: app-role
      rules:
        - apiGroups: ["*"]
          resources: ["*"]
          verbs: ["*"]
      ```
      
      **Why bad:** Grants cluster-admin equivalent access, violates least privilege, any compromised pod has full cluster control
      
      ### Bad: Deprecated API Version
      
      ```yaml
      # BAD: extensions/v1beta1 removed in Kubernetes 1.22
      apiVersion: extensions/v1beta1
      kind: Ingress
      metadata:
        name: api-ingress
        annotations:
          kubernetes.io/ingress.class: nginx
      ```
      
      **Why bad:** `extensions/v1beta1` removed since v1.22, `kubernetes.io/ingress.class` annotation deprecated in favor of `ingressClassName` field, missing `pathType` (required in v1)
      
    • helm.md 10.5 KB
      # Kubernetes -- Helm Chart Examples
      
      > Helm chart structure, values.yaml, templates, helpers, dependencies, and multi-environment releases. Reference from [SKILL.md](../SKILL.md).
      
      **Related examples:**
      
      - [core.md](core.md) - Deployments, Services, Ingress, ConfigMaps, Secrets
      - [operations.md](operations.md) - HPA, RBAC, health checks, resource limits, PDB
      
      ---
      
      ## Example 1: Complete Helm Chart
      
      ### Chart.yaml
      
      ```yaml
      apiVersion: v2
      name: my-app
      description: A Helm chart for my application
      type: application
      version: 0.1.0
      appVersion: "1.0.0"
      maintainers:
        - name: team-platform
      dependencies:
        - name: postgresql
          version: "~15.0"
          repository: "https://charts.bitnami.com/bitnami"
          condition: postgresql.enabled
      ```
      
      **Why good:** `apiVersion: v2` (Helm 3), dependencies inline (not separate `requirements.yaml`), `condition` enables toggling sub-charts, SemVer for both chart version and appVersion
      
      ### values.yaml
      
      ```yaml
      # -- Number of replicas (ignored when HPA enabled)
      replicaCount: 3
      
      image:
        repository: registry.example.com/my-app
        tag: "" # Defaults to chart appVersion
        pullPolicy: IfNotPresent
      
      service:
        type: ClusterIP
        port: 80
      
      ingress:
        enabled: false
        className: nginx
        hosts:
          - host: app.example.com
            paths:
              - path: /
                pathType: Prefix
        tls: []
      
      resources:
        requests:
          cpu: 250m
          memory: 256Mi
        limits:
          cpu: "1"
          memory: 512Mi
      
      autoscaling:
        enabled: false
        minReplicas: 2
        maxReplicas: 10
        targetCPUUtilization: 70
      
      securityContext:
        runAsNonRoot: true
        runAsUser: 1000
        allowPrivilegeEscalation: false
        readOnlyRootFilesystem: true
        capabilities:
          drop: ["ALL"]
      
      podSecurityContext:
        fsGroup: 2000
        seccompProfile:
          type: RuntimeDefault
      
      probes:
        liveness:
          path: /healthz
          initialDelaySeconds: 15
          periodSeconds: 20
        readiness:
          path: /ready
          initialDelaySeconds: 5
          periodSeconds: 10
      
      serviceAccount:
        create: true
        automountServiceAccountToken: false
        annotations: {}
      
      postgresql:
        enabled: true
      ```
      
      **Why good:** Grouped by concern, comments explain non-obvious defaults, security values mirror restricted Pod Security Standards, HPA and Ingress conditionally enabled, image tag defaults to appVersion
      
      ### templates/\_helpers.tpl
      
      ```yaml
      {{/*
      Expand the name of the chart.
      */}}
      {{- define "my-app.name" -}}
      {{- default .Chart.Name .Values.nameOverride | trunc 63 | trimSuffix "-" }}
      {{- end }}
      
      {{/*
      Create a fully qualified app name.
      */}}
      {{- define "my-app.fullname" -}}
      {{- if .Values.fullnameOverride }}
      {{- .Values.fullnameOverride | trunc 63 | trimSuffix "-" }}
      {{- else }}
      {{- $name := default .Chart.Name .Values.nameOverride }}
      {{- if contains $name .Release.Name }}
      {{- .Release.Name | trunc 63 | trimSuffix "-" }}
      {{- else }}
      {{- printf "%s-%s" .Release.Name $name | trunc 63 | trimSuffix "-" }}
      {{- end }}
      {{- end }}
      {{- end }}
      
      {{/*
      Common labels
      */}}
      {{- define "my-app.labels" -}}
      helm.sh/chart: {{ include "my-app.chart" . }}
      {{ include "my-app.selectorLabels" . }}
      {{- if .Chart.AppVersion }}
      app.kubernetes.io/version: {{ .Chart.AppVersion | quote }}
      {{- end }}
      app.kubernetes.io/managed-by: {{ .Release.Service }}
      {{- end }}
      
      {{/*
      Selector labels -- immutable after first deployment
      */}}
      {{- define "my-app.selectorLabels" -}}
      app.kubernetes.io/name: {{ include "my-app.name" . }}
      app.kubernetes.io/instance: {{ .Release.Name }}
      {{- end }}
      
      {{/*
      Chart label
      */}}
      {{- define "my-app.chart" -}}
      {{- printf "%s-%s" .Chart.Name .Chart.Version | replace "+" "_" | trunc 63 | trimSuffix "-" }}
      {{- end }}
      
      {{/*
      Service account name
      */}}
      {{- define "my-app.serviceAccountName" -}}
      {{- if .Values.serviceAccount.create }}
      {{- default (include "my-app.fullname" .) .Values.serviceAccount.name }}
      {{- else }}
      {{- default "default" .Values.serviceAccount.name }}
      {{- end }}
      {{- end }}
      ```
      
      **Why good:** Selector labels separated from common labels (selectors are immutable), names truncated to 63 chars (Kubernetes limit), `fullname` avoids redundant chart name in release name
      
      ### templates/deployment.yaml
      
      ```yaml
      apiVersion: apps/v1
      kind: Deployment
      metadata:
        name: {{ include "my-app.fullname" . }}
        labels:
          {{- include "my-app.labels" . | nindent 4 }}
      spec:
        {{- if not .Values.autoscaling.enabled }}
        replicas: {{ .Values.replicaCount }}
        {{- end }}
        revisionHistoryLimit: 5
        selector:
          matchLabels:
            {{- include "my-app.selectorLabels" . | nindent 6 }}
        template:
          metadata:
            labels:
              {{- include "my-app.selectorLabels" . | nindent 8 }}
          spec:
            serviceAccountName: {{ include "my-app.serviceAccountName" . }}
            automountServiceAccountToken: {{ .Values.serviceAccount.automountServiceAccountToken }}
            securityContext:
              {{- toYaml .Values.podSecurityContext | nindent 8 }}
            containers:
              - name: {{ .Chart.Name }}
                image: "{{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}"
                imagePullPolicy: {{ .Values.image.pullPolicy }}
                ports:
                  - name: http
                    containerPort: 3000
                    protocol: TCP
                securityContext:
                  {{- toYaml .Values.securityContext | nindent 12 }}
                resources:
                  {{- toYaml .Values.resources | nindent 12 }}
                livenessProbe:
                  httpGet:
                    path: {{ .Values.probes.liveness.path }}
                    port: http
                  initialDelaySeconds: {{ .Values.probes.liveness.initialDelaySeconds }}
                  periodSeconds: {{ .Values.probes.liveness.periodSeconds }}
                readinessProbe:
                  httpGet:
                    path: {{ .Values.probes.readiness.path }}
                    port: http
                  initialDelaySeconds: {{ .Values.probes.readiness.initialDelaySeconds }}
                  periodSeconds: {{ .Values.probes.readiness.periodSeconds }}
                volumeMounts:
                  - name: tmp
                    mountPath: /tmp
            volumes:
              - name: tmp
                emptyDir: {}
      ```
      
      **Why good:** Replicas conditionally omitted when HPA enabled, image tag falls back to `appVersion`, security contexts injected from values via `toYaml`, `emptyDir` for tmp (readOnlyRootFilesystem), `nindent` handles YAML indentation correctly
      
      ### templates/service.yaml
      
      ```yaml
      apiVersion: v1
      kind: Service
      metadata:
        name: { { include "my-app.fullname" . } }
        labels: { { - include "my-app.labels" . | nindent 4 } }
      spec:
        type: { { .Values.service.type } }
        selector: { { - include "my-app.selectorLabels" . | nindent 4 } }
        ports:
          - name: http
            port: { { .Values.service.port } }
            targetPort: http
            protocol: TCP
      ```
      
      ### templates/ingress.yaml (Conditional)
      
      ```yaml
      {{- if .Values.ingress.enabled -}}
      apiVersion: networking.k8s.io/v1
      kind: Ingress
      metadata:
        name: {{ include "my-app.fullname" . }}
        labels:
          {{- include "my-app.labels" . | nindent 4 }}
        {{- with .Values.ingress.annotations }}
        annotations:
          {{- toYaml . | nindent 4 }}
        {{- end }}
      spec:
        ingressClassName: {{ .Values.ingress.className }}
        {{- if .Values.ingress.tls }}
        tls:
          {{- range .Values.ingress.tls }}
          - hosts:
              {{- range .hosts }}
              - {{ . | quote }}
              {{- end }}
            secretName: {{ .secretName }}
          {{- end }}
        {{- end }}
        rules:
          {{- range .Values.ingress.hosts }}
          - host: {{ .host | quote }}
            http:
              paths:
                {{- range .paths }}
                - path: {{ .path }}
                  pathType: {{ .pathType }}
                  backend:
                    service:
                      name: {{ include "my-app.fullname" $ }}
                      port:
                        number: {{ $.Values.service.port }}
                {{- end }}
          {{- end }}
      {{- end }}
      ```
      
      **Why good:** Entire resource conditional on `ingress.enabled`, `ingressClassName` field (not deprecated annotation), `$` used for root context inside range loops, hosts quoted for safety
      
      ### templates/hpa.yaml (Conditional)
      
      ```yaml
      {{- if .Values.autoscaling.enabled }}
      apiVersion: autoscaling/v2
      kind: HorizontalPodAutoscaler
      metadata:
        name: {{ include "my-app.fullname" . }}
        labels:
          {{- include "my-app.labels" . | nindent 4 }}
      spec:
        scaleTargetRef:
          apiVersion: apps/v1
          kind: Deployment
          name: {{ include "my-app.fullname" . }}
        minReplicas: {{ .Values.autoscaling.minReplicas }}
        maxReplicas: {{ .Values.autoscaling.maxReplicas }}
        metrics:
          - type: Resource
            resource:
              name: cpu
              target:
                type: Utilization
                averageUtilization: {{ .Values.autoscaling.targetCPUUtilization }}
      {{- end }}
      ```
      
      ---
      
      ## Example 2: Multi-Environment Values
      
      ### values-staging.yaml
      
      ```yaml
      replicaCount: 1
      
      image:
        tag: "staging-latest"
      
      ingress:
        enabled: true
        hosts:
          - host: staging.example.com
            paths:
              - path: /
                pathType: Prefix
      
      resources:
        requests:
          cpu: 100m
          memory: 128Mi
        limits:
          cpu: 500m
          memory: 256Mi
      
      postgresql:
        enabled: true
      ```
      
      ### values-production.yaml
      
      ```yaml
      replicaCount: 3
      
      ingress:
        enabled: true
        hosts:
          - host: app.example.com
            paths:
              - path: /
                pathType: Prefix
        tls:
          - hosts:
              - app.example.com
            secretName: app-tls
      
      autoscaling:
        enabled: true
        minReplicas: 3
        maxReplicas: 20
        targetCPUUtilization: 60
      
      resources:
        requests:
          cpu: 500m
          memory: 512Mi
        limits:
          cpu: "2"
          memory: 1Gi
      ```
      
      **Deploy:** `helm upgrade --install my-app ./my-app -f values-production.yaml -n production`
      
      ---
      
      ## Anti-Patterns
      
      ### Bad: Complex Logic in Templates
      
      ```yaml
      # BAD: Business logic in template
      {{- if and .Values.autoscaling.enabled (gt (int .Values.autoscaling.maxReplicas) 5) (not .Values.maintenance.enabled) }}
      {{- if or (eq .Values.environment "production") (eq .Values.environment "staging") }}
      replicas: {{ mul .Values.replicaCount 2 | add 1 }}
      {{- end }}
      {{- end }}
      ```
      
      **Why bad:** Complex conditionals belong in `_helpers.tpl` as named templates or in application logic, not inlined in resource templates
      
      ### Bad: Hardcoded Values in Templates
      
      ```yaml
      # BAD: Values that change between environments hardcoded
      containers:
        - name: app
          image: "myregistry.io/app:v1.0.0"
          resources:
            requests:
              cpu: 500m
              memory: 512Mi
      ```
      
      **Why bad:** Image and resource values hardcoded in template instead of parameterized through `values.yaml` -- every environment gets the same config
      
      ### Bad: Missing Conditional for HPA + Replicas
      
      ```yaml
      # BAD: replicas always set even when HPA manages scaling
      spec:
        replicas: { { .Values.replicaCount } }
      ```
      
      **Why bad:** When HPA is active, Helm resets replicas to the manifest value on every upgrade, overriding HPA's scaling decisions. Wrap in `{{- if not .Values.autoscaling.enabled }}`
      
    • operations.md 11 KB
      # Kubernetes -- Operations Examples
      
      > HPA autoscaling, RBAC, health checks, resource limits, PodDisruptionBudget, NetworkPolicy, and Pod Security. Reference from [SKILL.md](../SKILL.md).
      
      **Related examples:**
      
      - [core.md](core.md) - Deployments, Services, Ingress, ConfigMaps, Secrets
      - [helm.md](helm.md) - Helm chart structure, values, templates, helpers
      
      ---
      
      ## Example 1: HPA with Multiple Metrics and Scaling Behavior
      
      ```yaml
      apiVersion: autoscaling/v2
      kind: HorizontalPodAutoscaler
      metadata:
        name: api-server
        namespace: app
      spec:
        scaleTargetRef:
          apiVersion: apps/v1
          kind: Deployment
          name: api-server
        minReplicas: 2
        maxReplicas: 20
        metrics:
          - type: Resource
            resource:
              name: cpu
              target:
                type: Utilization
                averageUtilization: 70
          - type: Resource
            resource:
              name: memory
              target:
                type: Utilization
                averageUtilization: 80
        behavior:
          scaleUp:
            stabilizationWindowSeconds: 60
            policies:
              - type: Pods
                value: 4
                periodSeconds: 60
              - type: Percent
                value: 100
                periodSeconds: 60
            selectPolicy: Max
          scaleDown:
            stabilizationWindowSeconds: 300
            policies:
              - type: Pods
                value: 1
                periodSeconds: 120
            selectPolicy: Min
      ```
      
      **Why good:** Uses stable `autoscaling/v2` with `metrics` array, both CPU and memory targets, scale-up allows aggressive burst (max of 4 pods or 100% increase per minute), scale-down is conservative (1 pod every 2 minutes, 5-minute stabilization) to prevent flapping
      
      **Gotcha:** Remove `spec.replicas` from the Deployment when HPA is active -- otherwise `helm upgrade` or `kubectl apply` resets the replica count on every deployment.
      
      ---
      
      ## Example 2: RBAC Patterns
      
      ### Namespaced Role (Preferred)
      
      ```yaml
      apiVersion: v1
      kind: ServiceAccount
      metadata:
        name: api-server
        namespace: app
      automountServiceAccountToken: false
      ---
      apiVersion: rbac.authorization.k8s.io/v1
      kind: Role
      metadata:
        name: api-server-role
        namespace: app
      rules:
        - apiGroups: [""]
          resources: ["configmaps"]
          verbs: ["get", "list", "watch"]
        - apiGroups: [""]
          resources: ["secrets"]
          verbs: ["get"]
          resourceNames: ["api-secrets"] # Restrict to specific secret
      ---
      apiVersion: rbac.authorization.k8s.io/v1
      kind: RoleBinding
      metadata:
        name: api-server-binding
        namespace: app
      subjects:
        - kind: ServiceAccount
          name: api-server
          namespace: app
      roleRef:
        kind: Role
        name: api-server-role
        apiGroup: rbac.authorization.k8s.io
      ```
      
      **Why good:** Namespaced Role (not ClusterRole), specific resources and verbs (not wildcards), `resourceNames` further restricts secret access to named resources, `automountServiceAccountToken: false` prevents unnecessary API access
      
      ### ClusterRole (When Required)
      
      Use ClusterRole only when a workload genuinely needs cluster-wide access (e.g., operators, controllers, monitoring agents).
      
      ```yaml
      apiVersion: rbac.authorization.k8s.io/v1
      kind: ClusterRole
      metadata:
        name: namespace-reader
      rules:
        - apiGroups: [""]
          resources: ["namespaces"]
          verbs: ["get", "list", "watch"]
      ---
      apiVersion: rbac.authorization.k8s.io/v1
      kind: ClusterRoleBinding
      metadata:
        name: monitoring-namespace-reader
      subjects:
        - kind: ServiceAccount
          name: monitoring
          namespace: monitoring
      roleRef:
        kind: ClusterRole
        name: namespace-reader
        apiGroup: rbac.authorization.k8s.io
      ```
      
      **When to use ClusterRole:**
      
      - Reading namespaces or nodes (cluster-scoped resources)
      - Custom operators that watch resources across namespaces
      - Monitoring/logging agents that collect from all namespaces
      
      ---
      
      ## Example 3: Health Check Patterns
      
      ### HTTP Probes (Most Common)
      
      ```yaml
      containers:
        - name: api-server
          livenessProbe:
            httpGet:
              path: /healthz
              port: 3000
            initialDelaySeconds: 15
            periodSeconds: 20
            timeoutSeconds: 5
            failureThreshold: 3
          readinessProbe:
            httpGet:
              path: /ready
              port: 3000
            initialDelaySeconds: 5
            periodSeconds: 10
            timeoutSeconds: 3
            failureThreshold: 3
          startupProbe:
            httpGet:
              path: /healthz
              port: 3000
            failureThreshold: 30
            periodSeconds: 10
      ```
      
      **Probe design:**
      
      - `/healthz` (liveness) -- Is the process alive? Check basic responsiveness only. Do NOT check external dependencies (database, Redis). A false-negative restarts a healthy pod.
      - `/ready` (readiness) -- Can this pod serve traffic? Check that connections to required dependencies are established.
      - Startup probe -- Gives slow-starting apps time to initialize. Blocks liveness/readiness until successful. `failureThreshold * periodSeconds` = max startup time (300s here).
      
      ### TCP Probe (Non-HTTP Services)
      
      ```yaml
      containers:
        - name: database
          livenessProbe:
            tcpSocket:
              port: 5432
            periodSeconds: 20
          readinessProbe:
            tcpSocket:
              port: 5432
            periodSeconds: 10
      ```
      
      ### Exec Probe (Custom Check)
      
      ```yaml
      containers:
        - name: worker
          livenessProbe:
            exec:
              command:
                - /bin/sh
                - -c
                - "test -f /tmp/healthy"
            periodSeconds: 30
      ```
      
      ### gRPC Probe (gRPC Services)
      
      ```yaml
      containers:
        - name: grpc-service
          livenessProbe:
            grpc:
              port: 50051
            periodSeconds: 20
          readinessProbe:
            grpc:
              port: 50051
            periodSeconds: 10
      ```
      
      ---
      
      ## Example 4: Resource Limits and QoS
      
      ### Guaranteed QoS (Critical Services)
      
      ```yaml
      # requests == limits = Guaranteed QoS class
      # Highest priority, never evicted under memory pressure
      containers:
        - name: payment-service
          resources:
            requests:
              cpu: 500m
              memory: 512Mi
            limits:
              cpu: 500m
              memory: 512Mi
      ```
      
      ### Burstable QoS (Typical Workloads)
      
      ```yaml
      # requests < limits = Burstable QoS class
      # Can burst above requests when resources available
      containers:
        - name: api-server
          resources:
            requests:
              cpu: 250m
              memory: 256Mi
            limits:
              cpu: "1"
              memory: 512Mi
      ```
      
      **Sizing guidance:**
      
      - Set memory requests at 90th percentile of observed usage
      - Set memory limits at 150-200% of requests to allow spikes without OOMKill
      - Set CPU requests based on sustained usage, limits based on peak
      - CPU throttling (hitting limit) causes latency spikes but NOT kills
      - Memory exceeding limits causes OOMKill -- the container is terminated and restarted
      
      ---
      
      ## Example 5: PodDisruptionBudget
      
      Ensures minimum availability during voluntary disruptions (node drain, cluster upgrade, rolling update).
      
      ```yaml
      apiVersion: policy/v1
      kind: PodDisruptionBudget
      metadata:
        name: api-server
        namespace: app
      spec:
        minAvailable: 2
        selector:
          matchLabels:
            app.kubernetes.io/name: api-server
      ```
      
      **Alternative: maxUnavailable**
      
      ```yaml
      spec:
        maxUnavailable: 1
        selector:
          matchLabels:
            app.kubernetes.io/name: api-server
      ```
      
      **When to use which:**
      
      - `minAvailable: N` -- Use when you know the minimum pod count for service health
      - `maxUnavailable: N` -- Use when you want to limit disruption rate (better for larger deployments)
      - Cannot set both `minAvailable` and `maxUnavailable`
      - PDB only protects against **voluntary** disruptions (drain, upgrade) -- not OOMKill or node failure
      
      ---
      
      ## Example 6: NetworkPolicy
      
      Restrict pod-to-pod traffic. By default, all pods can communicate with all other pods.
      
      ### Allow Only Specific Ingress
      
      ```yaml
      apiVersion: networking.k8s.io/v1
      kind: NetworkPolicy
      metadata:
        name: api-server-policy
        namespace: app
      spec:
        podSelector:
          matchLabels:
            app.kubernetes.io/name: api-server
        policyTypes:
          - Ingress
          - Egress
        ingress:
          # Allow traffic from ingress controller
          - from:
              - namespaceSelector:
                  matchLabels:
                    kubernetes.io/metadata.name: ingress-nginx
            ports:
              - port: 3000
                protocol: TCP
          # Allow traffic from frontend pods
          - from:
              - podSelector:
                  matchLabels:
                    app.kubernetes.io/name: frontend
            ports:
              - port: 3000
                protocol: TCP
        egress:
          # Allow DNS
          - to:
              - namespaceSelector: {}
            ports:
              - port: 53
                protocol: UDP
              - port: 53
                protocol: TCP
          # Allow database
          - to:
              - podSelector:
                  matchLabels:
                    app.kubernetes.io/name: database
            ports:
              - port: 5432
                protocol: TCP
      ```
      
      **Why good:** Explicit ingress sources (ingress controller and frontend only), explicit egress (DNS and database only), all other traffic implicitly denied once a NetworkPolicy selects a pod
      
      **Gotcha:** NetworkPolicy requires a CNI plugin that supports it (Calico, Cilium, Weave). The default kubenet CNI ignores NetworkPolicy.
      
      ### Default Deny All (Namespace-Level)
      
      ```yaml
      apiVersion: networking.k8s.io/v1
      kind: NetworkPolicy
      metadata:
        name: default-deny-all
        namespace: app
      spec:
        podSelector: {}
        policyTypes:
          - Ingress
          - Egress
      ```
      
      **Why useful:** Apply as baseline, then add allow policies for specific communication paths. Empty `podSelector` selects all pods in the namespace.
      
      ---
      
      ## Example 7: Pod Security Context (Restricted Profile)
      
      The complete restricted security context that satisfies the Pod Security Standards restricted level.
      
      ### Pod Level
      
      ```yaml
      spec:
        securityContext:
          runAsNonRoot: true
          runAsUser: 1000
          runAsGroup: 1000
          fsGroup: 2000
          seccompProfile:
            type: RuntimeDefault
      ```
      
      ### Container Level
      
      ```yaml
      containers:
        - name: app
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            runAsNonRoot: true
            capabilities:
              drop: ["ALL"]
      ```
      
      ### Handling readOnlyRootFilesystem
      
      Apps that write to `/tmp`, `/var/cache`, or similar directories need `emptyDir` volume mounts:
      
      ```yaml
      containers:
        - name: app
          securityContext:
            readOnlyRootFilesystem: true
          volumeMounts:
            - name: tmp
              mountPath: /tmp
            - name: cache
              mountPath: /var/cache
      volumes:
        - name: tmp
          emptyDir: {}
        - name: cache
          emptyDir:
            sizeLimit: 100Mi
      ```
      
      ---
      
      ## Anti-Patterns
      
      ### Bad: Liveness Probe Checks External Dependencies
      
      ```yaml
      # BAD: Liveness probe checks database connectivity
      livenessProbe:
        httpGet:
          path: /health # Returns 500 if database is down
          port: 3000
      ```
      
      **Why bad:** If the database goes down, ALL pods fail their liveness probe and get restarted simultaneously. The pods are healthy -- they just can't reach the database. This turns a dependency outage into a cascading failure. Check external dependencies in readiness probe only.
      
      ### Bad: No PDB with Multiple Replicas
      
      ```yaml
      # BAD: 3 replicas but no PDB -- node drain can take all 3 down
      apiVersion: apps/v1
      kind: Deployment
      metadata:
        name: api-server
      spec:
        replicas: 3
        # No PodDisruptionBudget defined
      ```
      
      **Why bad:** During a cluster upgrade or node drain, all 3 pods could be evicted simultaneously if they happen to be on the same node or on nodes being drained in sequence. Always pair multi-replica Deployments with a PDB.
      
  • reference.md 7.4 KB
    # Kubernetes Quick Reference
    
    > Stable API versions, kubectl commands, label conventions, and decision tables. Reference from [SKILL.md](SKILL.md).
    
    ---
    
    ## Stable API Versions (Kubernetes 1.30+)
    
    | Resource                         | apiVersion                     | Stable Since |
    | -------------------------------- | ------------------------------ | ------------ |
    | Deployment                       | `apps/v1`                      | v1.9         |
    | StatefulSet                      | `apps/v1`                      | v1.9         |
    | DaemonSet                        | `apps/v1`                      | v1.9         |
    | ReplicaSet                       | `apps/v1`                      | v1.9         |
    | Service                          | `v1`                           | v1.0         |
    | ConfigMap                        | `v1`                           | v1.0         |
    | Secret                           | `v1`                           | v1.0         |
    | Namespace                        | `v1`                           | v1.0         |
    | ServiceAccount                   | `v1`                           | v1.0         |
    | PersistentVolumeClaim            | `v1`                           | v1.0         |
    | Ingress                          | `networking.k8s.io/v1`         | v1.19        |
    | IngressClass                     | `networking.k8s.io/v1`         | v1.19        |
    | NetworkPolicy                    | `networking.k8s.io/v1`         | v1.7         |
    | HorizontalPodAutoscaler          | `autoscaling/v2`               | v1.23        |
    | PodDisruptionBudget              | `policy/v1`                    | v1.21        |
    | Job                              | `batch/v1`                     | v1.0         |
    | CronJob                          | `batch/v1`                     | v1.21        |
    | Role / ClusterRole               | `rbac.authorization.k8s.io/v1` | v1.8         |
    | RoleBinding / ClusterRoleBinding | `rbac.authorization.k8s.io/v1` | v1.8         |
    
    ---
    
    ## Removed API Versions (Do NOT Use)
    
    | Resource            | Removed apiVersion          | Removed In | Use Instead            |
    | ------------------- | --------------------------- | ---------- | ---------------------- |
    | Ingress             | `extensions/v1beta1`        | v1.22      | `networking.k8s.io/v1` |
    | Ingress             | `networking.k8s.io/v1beta1` | v1.22      | `networking.k8s.io/v1` |
    | CronJob             | `batch/v1beta1`             | v1.25      | `batch/v1`             |
    | HPA                 | `autoscaling/v2beta1`       | v1.25      | `autoscaling/v2`       |
    | HPA                 | `autoscaling/v2beta2`       | v1.26      | `autoscaling/v2`       |
    | PodDisruptionBudget | `policy/v1beta1`            | v1.25      | `policy/v1`            |
    | PodSecurityPolicy   | `policy/v1beta1`            | v1.25      | Pod Security Admission |
    
    ---
    
    ## Standard Labels (app.kubernetes.io)
    
    | Label                          | Purpose                       | Example           |
    | ------------------------------ | ----------------------------- | ----------------- |
    | `app.kubernetes.io/name`       | Application name              | `api-server`      |
    | `app.kubernetes.io/instance`   | Unique instance of the app    | `api-server-prod` |
    | `app.kubernetes.io/version`    | Application version           | `1.2.3`           |
    | `app.kubernetes.io/component`  | Component within architecture | `backend`         |
    | `app.kubernetes.io/part-of`    | Higher-level application      | `my-platform`     |
    | `app.kubernetes.io/managed-by` | Tool managing this resource   | `helm`            |
    
    ---
    
    ## kubectl Essential Commands
    
    ```bash
    # Apply manifests (declarative, idempotent)
    kubectl apply -f manifest.yaml
    kubectl apply -k overlays/production/     # Kustomize
    
    # Inspect resources
    kubectl get pods -n app -o wide
    kubectl describe deployment api-server -n app
    kubectl logs -f deployment/api-server -n app --tail=100
    
    # Debug
    kubectl exec -it pod/api-server-abc123 -n app -- sh
    kubectl port-forward svc/api-server 3000:80 -n app
    kubectl top pods -n app                   # Requires metrics-server
    
    # Rollout management
    kubectl rollout status deployment/api-server -n app
    kubectl rollout history deployment/api-server -n app
    kubectl rollout undo deployment/api-server -n app
    kubectl rollout restart deployment/api-server -n app  # Force rolling restart
    
    # Dry run and diff
    kubectl apply -f manifest.yaml --dry-run=server  # Server-side validation
    kubectl diff -f manifest.yaml                     # Show what would change
    
    # Resource cleanup
    kubectl delete -f manifest.yaml
    kubectl delete pod api-server-abc123 -n app --grace-period=30
    ```
    
    ---
    
    ## Probe Types
    
    | Probe Type       | Purpose                        | Failure Action                         |
    | ---------------- | ------------------------------ | -------------------------------------- |
    | `livenessProbe`  | Is the process healthy?        | Restart the container                  |
    | `readinessProbe` | Can the pod serve traffic?     | Remove from Service endpoints          |
    | `startupProbe`   | Has the app finished starting? | Block liveness/readiness until success |
    
    **Probe mechanisms:** `httpGet`, `tcpSocket`, `exec`, `grpc`
    
    **Timing fields:** `initialDelaySeconds`, `periodSeconds`, `timeoutSeconds`, `failureThreshold`, `successThreshold`
    
    ---
    
    ## Resource Units
    
    | Resource | Unit        | Examples                         |
    | -------- | ----------- | -------------------------------- |
    | CPU      | Millicores  | `100m` = 0.1 CPU, `"1"` = 1 CPU  |
    | Memory   | Bytes (IEC) | `128Mi` = 128 MiB, `1Gi` = 1 GiB |
    
    **QoS classes:**
    
    - **Guaranteed:** requests == limits (highest priority, no eviction under pressure)
    - **Burstable:** requests < limits (typical, evicted after BestEffort)
    - **BestEffort:** no requests/limits (first to be evicted)
    
    ---
    
    ## Pod Security Standards
    
    | Level      | Enforcement               | Key Controls                                                                |
    | ---------- | ------------------------- | --------------------------------------------------------------------------- |
    | Privileged | No restrictions           | System workloads only                                                       |
    | Baseline   | Prevents known escalation | No privileged, no hostNetwork/PID/IPC, no hostPath                          |
    | Restricted | Full hardening            | runAsNonRoot, drop ALL caps, readOnlyRootFilesystem, seccomp RuntimeDefault |
    
    **Enforce via namespace labels:**
    
    ```yaml
    apiVersion: v1
    kind: Namespace
    metadata:
      name: app
      labels:
        pod-security.kubernetes.io/enforce: restricted
        pod-security.kubernetes.io/audit: restricted
        pod-security.kubernetes.io/warn: restricted
    ```
    
    ---
    
    ## Helm Commands
    
    ```bash
    # Create / Install / Upgrade
    helm create my-app                              # Scaffold new chart
    helm install my-release ./my-app                # Install from local chart
    helm upgrade my-release ./my-app -f values-prod.yaml  # Upgrade with values
    helm upgrade --install my-release ./my-app      # Install or upgrade
    
    # Inspect
    helm list -n app                                # List releases
    helm status my-release -n app                   # Release status
    helm get values my-release -n app               # Current values
    helm template my-release ./my-app               # Render templates locally
    
    # Rollback
    helm rollback my-release 1 -n app               # Rollback to revision 1
    helm history my-release -n app                   # Release history
    
    # Diff (requires helm-diff plugin)
    helm diff upgrade my-release ./my-app -f values-prod.yaml
    ```
    
  • SKILL.md 18.4 KB
    ---
    name: infra-containers-kubernetes
    description: Kubernetes manifests, Helm charts, Kustomize overlays, and resource patterns
    ---
    
    # Kubernetes Patterns
    
    > **Quick Guide:** Declarative YAML manifests for Kubernetes workloads. Use `apps/v1` Deployments with resource requests/limits, health probes, and Pod Security Standards (restricted profile). Helm for templated multi-environment releases. Kustomize for patch-based overlays without templating. Always set `securityContext` (runAsNonRoot, drop ALL capabilities, readOnlyRootFilesystem), resource requests/limits on every container, and liveness/readiness probes on every pod.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST set resource requests AND limits on every container -- pods without requests are unschedulable under resource pressure and pods without limits can OOM-kill neighbors)**
    
    **(You MUST set securityContext with runAsNonRoot: true, allowPrivilegeEscalation: false, drop ALL capabilities, and readOnlyRootFilesystem: true on every container)**
    
    **(You MUST define both liveness and readiness probes -- without readiness probes, traffic reaches unready pods; without liveness probes, hung processes are never restarted)**
    
    **(You MUST use the current stable apiVersion for each resource -- apps/v1 for Deployments, networking.k8s.io/v1 for Ingress, autoscaling/v2 for HPA, policy/v1 for PDB)**
    
    </critical_requirements>
    
    ---
    
    ## Examples
    
    - [Core Manifests](examples/core.md) - Deployments, Services, Ingress, ConfigMaps, Secrets, Namespaces
    - [Helm Charts](examples/helm.md) - Chart structure, values.yaml, templates, helpers, dependencies
    - [Operations](examples/operations.md) - HPA scaling, RBAC, health checks, resource limits, PDB, NetworkPolicy
    - [Quick Reference](reference.md) - API versions, kubectl commands, label conventions, decision frameworks
    
    ---
    
    **Auto-detection:** Kubernetes, kubectl, k8s, Deployment, Service, Ingress, ConfigMap, Secret, HPA, HorizontalPodAutoscaler, Helm, helm chart, Kustomize, kustomization, RBAC, Role, ClusterRole, PodDisruptionBudget, NetworkPolicy, Pod, StatefulSet, DaemonSet, CronJob, Job, PersistentVolumeClaim, apiVersion, kind, metadata, spec
    
    **When to use:**
    
    - Writing Kubernetes Deployment, Service, Ingress, or other resource manifests
    - Creating Helm charts for templated multi-environment releases
    - Building Kustomize overlays for environment-specific patches
    - Configuring RBAC roles and bindings for least-privilege access
    - Setting up HPA autoscaling, PDB, or resource limits
    - Defining health checks (liveness, readiness, startup probes)
    - Managing ConfigMaps, Secrets, and environment configuration
    - Writing NetworkPolicy for pod-to-pod traffic control
    
    **When NOT to use:**
    
    - Container image building (use a containerization skill)
    - CI/CD pipeline definitions (use a CI/CD skill)
    - Infrastructure provisioning (use an IaC tool)
    - Service mesh configuration beyond basic Kubernetes resources
    - Managed Kubernetes cluster setup (cloud provider control plane configuration)
    
    **Key patterns covered:**
    
    - Deployment with security context, resource limits, and probes
    - Service types (ClusterIP, NodePort, LoadBalancer, Headless)
    - Ingress with TLS and path routing (networking.k8s.io/v1)
    - ConfigMap and Secret management (envFrom, volume mounts)
    - HPA autoscaling (autoscaling/v2 metrics array)
    - RBAC (Role, ClusterRole, RoleBinding, ServiceAccount)
    - Pod Security Standards (restricted profile)
    - Helm chart structure, values, templates, and helpers
    - Kustomize base/overlay pattern with strategic merge patches
    - PodDisruptionBudget for safe rollouts
    - NetworkPolicy for pod traffic isolation
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    Kubernetes is a declarative container orchestration platform. You describe the **desired state** in YAML manifests and Kubernetes continuously reconciles actual state to match. Every resource should be version-controlled, reproducible, and deployable via `kubectl apply` or a GitOps pipeline.
    
    **Core principles:**
    
    1. **Declarative over imperative** -- Use `kubectl apply -f` with manifests, not `kubectl run` or `kubectl create`
    2. **Security by default** -- Every pod runs as non-root, drops all capabilities, uses read-only root filesystem
    3. **Resource-aware** -- Every container declares requests (scheduling guarantee) and limits (ceiling)
    4. **Observable** -- Every pod has health probes so the platform can detect and recover from failures
    5. **Least privilege** -- RBAC grants only the permissions each workload needs, scoped to namespace when possible
    
    **Helm vs Kustomize:**
    
    - **Helm** -- Templating engine with package management. Use when you need parameterized releases, dependency management, or you distribute charts to others
    - **Kustomize** -- Patch-based overlays built into kubectl. Use when you want to keep base manifests as valid YAML and apply environment-specific patches without templating
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Production Deployment
    
    A production-ready Deployment includes security context, resource limits, health probes, and standard labels.
    
    ```yaml
    apiVersion: apps/v1
    kind: Deployment
    metadata:
      name: api-server
      labels:
        app.kubernetes.io/name: api-server
        app.kubernetes.io/component: backend
    spec:
      replicas: 3
      revisionHistoryLimit: 5
      selector:
        matchLabels:
          app.kubernetes.io/name: api-server
      template:
        metadata:
          labels:
            app.kubernetes.io/name: api-server
            app.kubernetes.io/component: backend
        spec:
          serviceAccountName: api-server
          automountServiceAccountToken: false
          securityContext:
            runAsNonRoot: true
            runAsUser: 1000
            fsGroup: 2000
            seccompProfile:
              type: RuntimeDefault
          containers:
            - name: api-server
              image: registry.example.com/api-server:v1.2.3
              ports:
                - containerPort: 3000
                  protocol: TCP
              securityContext:
                allowPrivilegeEscalation: false
                readOnlyRootFilesystem: true
                capabilities:
                  drop: ["ALL"]
              resources:
                requests:
                  cpu: 250m
                  memory: 256Mi
                limits:
                  cpu: "1"
                  memory: 512Mi
              livenessProbe:
                httpGet:
                  path: /healthz
                  port: 3000
                initialDelaySeconds: 15
                periodSeconds: 20
              readinessProbe:
                httpGet:
                  path: /ready
                  port: 3000
                initialDelaySeconds: 5
                periodSeconds: 10
    ```
    
    **Why good:** Restricted security context (non-root, drop ALL, read-only FS, seccomp), resource requests AND limits, both liveness and readiness probes, standard `app.kubernetes.io` labels, pinned image tag (not `:latest`), dedicated service account with token auto-mount disabled
    
    See [examples/core.md](examples/core.md) for the full Deployment with matching Service, Ingress, and ConfigMap.
    
    ---
    
    ### Pattern 2: Service Types
    
    Services expose pods to network traffic. Choose the type based on who needs access.
    
    ```yaml
    # ClusterIP (default) -- internal traffic only
    apiVersion: v1
    kind: Service
    metadata:
      name: api-server
    spec:
      type: ClusterIP
      selector:
        app.kubernetes.io/name: api-server
      ports:
        - port: 80
          targetPort: 3000
          protocol: TCP
    ```
    
    **When to use each type:**
    
    - **ClusterIP** (default) -- Internal services, microservice-to-microservice communication
    - **NodePort** -- Development/testing, direct node access (ports 30000-32767)
    - **LoadBalancer** -- External traffic via cloud provider load balancer
    - **Headless** (`clusterIP: None`) -- StatefulSet DNS discovery, direct pod addressing
    
    See [examples/core.md](examples/core.md) for all service type examples.
    
    ---
    
    ### Pattern 3: Ingress with TLS
    
    Ingress routes external HTTP/HTTPS traffic to Services. Uses `networking.k8s.io/v1` (stable since 1.19).
    
    ```yaml
    apiVersion: networking.k8s.io/v1
    kind: Ingress
    metadata:
      name: api-ingress
      annotations:
        nginx.ingress.kubernetes.io/ssl-redirect: "true"
    spec:
      ingressClassName: nginx
      tls:
        - hosts:
            - api.example.com
          secretName: api-tls
      rules:
        - host: api.example.com
          http:
            paths:
              - path: /
                pathType: Prefix
                backend:
                  service:
                    name: api-server
                    port:
                      number: 80
    ```
    
    **Why good:** Uses stable `networking.k8s.io/v1`, `ingressClassName` field (not deprecated annotation), TLS termination with Secret reference, explicit `pathType`
    
    See [examples/core.md](examples/core.md) for multi-path Ingress and path type comparison.
    
    ---
    
    ### Pattern 4: ConfigMap and Secret Management
    
    ConfigMaps hold non-sensitive configuration; Secrets hold sensitive data (base64-encoded, not encrypted at rest by default).
    
    ```yaml
    apiVersion: v1
    kind: ConfigMap
    metadata:
      name: api-config
    data:
      LOG_LEVEL: "info"
      MAX_CONNECTIONS: "100"
    ---
    apiVersion: v1
    kind: Secret
    metadata:
      name: api-secrets
    type: Opaque
    stringData:
      DATABASE_URL: "postgresql://user:pass@db:5432/app"
    ```
    
    **Injection patterns:**
    
    - `envFrom` -- Load all keys as environment variables (simple, flat config)
    - `env.valueFrom` -- Select individual keys (when you need only specific values)
    - Volume mount -- Mount as files (when apps read config from filesystem)
    
    See [examples/core.md](examples/core.md) for envFrom, selective env, and volume mount examples.
    
    ---
    
    ### Pattern 5: Helm Chart Structure
    
    Helm charts package Kubernetes manifests as reusable, parameterized releases.
    
    ```
    my-app/
      Chart.yaml          # apiVersion: v2, name, version, appVersion
      values.yaml         # Default configuration values
      templates/
        deployment.yaml   # Templated Deployment
        service.yaml      # Templated Service
        ingress.yaml      # Templated Ingress (conditional)
        _helpers.tpl      # Named template definitions
        NOTES.txt         # Post-install message
    ```
    
    ```yaml
    # Chart.yaml
    apiVersion: v2
    name: my-app
    description: A Helm chart for my application
    type: application
    version: 0.1.0
    appVersion: "1.0.0"
    ```
    
    **Key principles:** Values that change between environments go in `values.yaml`, not in templates. Keep template logic minimal -- complex conditionals belong in `_helpers.tpl` as named templates.
    
    See [examples/helm.md](examples/helm.md) for complete chart with templates, helpers, values, and multi-environment overrides.
    
    ---
    
    ### Pattern 6: Kustomize Base/Overlay
    
    Kustomize patches valid YAML manifests without templating. Base manifests stay deployable as-is.
    
    ```
    k8s/
      base/
        kustomization.yaml
        deployment.yaml
        service.yaml
      overlays/
        staging/
          kustomization.yaml
          replica-patch.yaml
        production/
          kustomization.yaml
          replica-patch.yaml
          resource-patch.yaml
    ```
    
    ```yaml
    # base/kustomization.yaml
    apiVersion: kustomize.config.k8s.io/v1beta1
    kind: Kustomization
    resources:
      - deployment.yaml
      - service.yaml
    
    # overlays/production/kustomization.yaml
    apiVersion: kustomize.config.k8s.io/v1beta1
    kind: Kustomization
    resources:
      - ../../base
    patches:
      - path: replica-patch.yaml
      - path: resource-patch.yaml
    namespace: production
    commonLabels:
      env: production
    ```
    
    **Why good:** Base manifests are valid `kubectl apply` targets, overlays only contain diffs, no template syntax to learn, built into `kubectl apply -k`
    
    See [examples/helm.md](examples/helm.md) for Kustomize patches and secretGenerator.
    
    ---
    
    ### Pattern 7: RBAC Least Privilege
    
    Grant only the permissions each workload needs. Prefer namespaced Role over cluster-wide ClusterRole.
    
    ```yaml
    apiVersion: v1
    kind: ServiceAccount
    metadata:
      name: api-server
      namespace: app
    automountServiceAccountToken: false
    ---
    apiVersion: rbac.authorization.k8s.io/v1
    kind: Role
    metadata:
      name: api-server-role
      namespace: app
    rules:
      - apiGroups: [""]
        resources: ["configmaps", "secrets"]
        verbs: ["get", "list", "watch"]
    ---
    apiVersion: rbac.authorization.k8s.io/v1
    kind: RoleBinding
    metadata:
      name: api-server-binding
      namespace: app
    subjects:
      - kind: ServiceAccount
        name: api-server
        namespace: app
    roleRef:
      kind: Role
      name: api-server-role
      apiGroup: rbac.authorization.k8s.io
    ```
    
    **Why good:** Dedicated service account (not default), auto-mount disabled, namespaced Role (not ClusterRole), specific resources and verbs (not wildcard `*`)
    
    See [examples/operations.md](examples/operations.md) for ClusterRole patterns and when to use them.
    
    ---
    
    ### Pattern 8: HPA Autoscaling
    
    Use `autoscaling/v2` (stable since 1.23). The `metrics` array replaces the old `targetCPUUtilizationPercentage` field.
    
    ```yaml
    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    metadata:
      name: api-server
    spec:
      scaleTargetRef:
        apiVersion: apps/v1
        kind: Deployment
        name: api-server
      minReplicas: 2
      maxReplicas: 10
      metrics:
        - type: Resource
          resource:
            name: cpu
            target:
              type: Utilization
              averageUtilization: 70
        - type: Resource
          resource:
            name: memory
            target:
              type: Utilization
              averageUtilization: 80
      behavior:
        scaleDown:
          stabilizationWindowSeconds: 300
    ```
    
    **Why good:** Uses stable `autoscaling/v2` with `metrics` array (not deprecated `targetCPUUtilizationPercentage`), both CPU and memory metrics, scale-down stabilization prevents flapping
    
    See [examples/operations.md](examples/operations.md) for custom metrics and scaling behavior configuration.
    
    </patterns>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Helm vs Kustomize
    
    ```
    Do you distribute charts to external teams?
      +-- YES --> Helm (package management, versioned releases)
      +-- NO  --> Do you need Go templating / conditionals?
          +-- YES --> Helm (parameterized templates)
          +-- NO  --> Do you prefer plain YAML with patches?
              +-- YES --> Kustomize (built into kubectl, no templating)
              +-- NO  --> Either works. Choose team familiarity.
    ```
    
    ### Workload Type Selection
    
    ```
    What is the workload?
      +-- Stateless HTTP/API server? --> Deployment
      +-- Needs stable network identity / ordered startup? --> StatefulSet
      +-- Must run on every node (logging, monitoring)? --> DaemonSet
      +-- One-time batch job? --> Job
      +-- Scheduled recurring job? --> CronJob
    ```
    
    ### Service Type Selection
    
    ```
    Who needs to reach this service?
      +-- Other pods in the cluster? --> ClusterIP (default)
      +-- External traffic via HTTP/HTTPS? --> ClusterIP + Ingress
      +-- External TCP/UDP without Ingress? --> LoadBalancer
      +-- Direct node access (dev/test)? --> NodePort
      +-- StatefulSet pod discovery? --> Headless (clusterIP: None)
    ```
    
    ### Resource Requests vs Limits
    
    ```
    What QoS class do you need?
      +-- Guaranteed (critical workloads) --> requests == limits
      +-- Burstable (typical workloads) --> requests < limits
      +-- BestEffort (batch, non-critical) --> no requests/limits (NOT recommended for production)
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Missing `resources.requests` and `resources.limits` on containers -- unschedulable under pressure, can OOM-kill neighbors
    - Running as root (no `securityContext.runAsNonRoot: true`) -- container escape gives host root access
    - Using `:latest` image tag -- non-deterministic deployments, impossible to roll back to known version
    - Missing liveness/readiness probes -- hung processes never restart, traffic hits unready pods
    - Using deprecated apiVersions (`extensions/v1beta1`, `autoscaling/v2beta2`) -- will fail on modern clusters
    - Wildcard RBAC rules (`verbs: ["*"]`, `resources: ["*"]`) -- violates least privilege, security risk
    - Storing actual secrets in manifests committed to Git -- use sealed secrets or external secret operators
    
    **Medium Priority Issues:**
    
    - `automountServiceAccountToken` not set to false -- every pod gets a token that can access the API server
    - No `PodDisruptionBudget` -- cluster upgrades or node drains can terminate all replicas simultaneously
    - No `NetworkPolicy` -- all pods can communicate with all other pods by default
    - Using `spec.targetCPUUtilizationPercentage` in HPA -- deprecated in `autoscaling/v2`, use `metrics` array
    - Missing `revisionHistoryLimit` on Deployments -- unlimited old ReplicaSets consume etcd storage
    - ConfigMap/Secret changes not triggering rollout -- pods keep stale config until manually restarted
    
    **Gotchas & Edge Cases:**
    
    - `readOnlyRootFilesystem: true` breaks apps that write to `/tmp` -- add an `emptyDir` volume mount for `/tmp`
    - `requests.memory` too low causes OOMKill; `limits.cpu` too low causes CPU throttling (latency spikes, not kills)
    - Ingress `pathType: Prefix` matches `/api` AND `/api-docs` -- use `pathType: Exact` for precise matching or add trailing `/`
    - Secrets are base64-encoded, NOT encrypted -- anyone with RBAC read access to secrets can decode them
    - HPA and manual `replicas` in Deployment conflict -- remove `spec.replicas` from manifests when HPA is active
    - `kubectl apply` vs `kubectl create` -- apply is declarative and idempotent; create fails if resource exists
    - Kustomize `commonLabels` adds labels to selectors too -- changing them on existing Deployments breaks selector immutability
    - Helm `lookup` function doesn't work during `helm template` (no cluster access) -- only works during `helm install/upgrade`
    - Pod Security Admission enforces at namespace level via labels -- pods in unlabeled namespaces get the default (usually privileged)
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST set resource requests AND limits on every container -- pods without requests are unschedulable under resource pressure and pods without limits can OOM-kill neighbors)**
    
    **(You MUST set securityContext with runAsNonRoot: true, allowPrivilegeEscalation: false, drop ALL capabilities, and readOnlyRootFilesystem: true on every container)**
    
    **(You MUST define both liveness and readiness probes -- without readiness probes, traffic reaches unready pods; without liveness probes, hung processes are never restarted)**
    
    **(You MUST use the current stable apiVersion for each resource -- apps/v1 for Deployments, networking.k8s.io/v1 for Ingress, autoscaling/v2 for HPA, policy/v1 for PDB)**
    
    **Failure to follow these rules will produce insecure pods with root access, unschedulable workloads, unmonitored health, and manifests that fail on modern clusters.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related