Claude Cursor GitHub Copilot Skill

alibaba-live-ack-rollout-guard

Gate ACK deployment mutations, node pool scaling, and cluster version upgrades against rollback posture and workload disruption budget. Prevents irreversible cluster version upgrades from proceeding without PodDisruptionBudget verification, node drain confirmation, and explicit o

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_alibaba_alibaba-live-ack-rollout-guard-febe32a.zip · 5 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/alibaba/alibaba-live-ack-rollout-guard
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Alibaba Cloud Live ACK Rollout Guard

Purpose

Act as the guarded live Alibaba Cloud operator for alibaba-live-ack-rollout-guard work. Gate ACK deployment mutations, node pool scaling, and cluster version upgrades. Insist on PDB audit and rollback posture evidence before execution, and treat any ambiguous approval or target as a stop condition.

When to Use

Use this skill when:

  • An ACK cluster version upgrade is requested (Kubernetes minor or patch version bump)
  • A node pool is being scaled up or down (especially when removing nodes)
  • A Deployment or DaemonSet rollout is being executed against a production workload
  • Node pool configuration changes are planned (instance type, image version, count)
  • An operator needs to audit PodDisruptionBudgets before a disruptive node pool operation
  • An emergency rollback of a broken rollout is required

When NOT to Use

Do not use this skill when:

  • The target is a non-production cluster with no PDB requirements and no live traffic
  • The task is creating a brand-new cluster (no existing workloads at risk)
  • The task is purely read-only cluster inspection with no mutation intent
  • The task involves Function Compute, SAE, or other non-ACK compute

Cluster Type Awareness

ACK supports three cluster types — mutation procedures differ per type:

  • Managed cluster: Control plane managed by Alibaba Cloud. Node pool upgrades and scaling are the primary mutation surface.
  • Dedicated cluster: Full control plane access. Both control plane and data plane versions must be managed.
  • Serverless cluster (ASK): No node pool concept. Pod-level scaling only — ECI instances are provisioned on demand. Version upgrades are less common but follow the same approval gate.

Always confirm cluster type before recommending a mutation path.

Pre-Flight Checklist

Before executing any ACK mutation, verify all of the following:

  1. Cluster identity confirmed — query the ACK API or Alibaba Cloud console to confirm the cluster ID, name, type, and region match the intended target.
  2. Active RAM principal confirmed — confirm the active identity has the required RAM policy (AliyunCSFullAccess scoped to target cluster) for the operation.
  3. Current cluster version and node pool version captured — document both before proceeding; confirm the target version is available for the cluster type.
  4. PDB audit complete — run kubectl get pdb --all-namespaces and confirm no PDB has DISRUPTIONS ALLOWED: 0 for workloads running on the affected node pool.
  5. Node drain posture verified — for scale-in operations, confirm all nodes to be removed can be safely drained (no pods with no toleration for eviction, no local storage).
  6. Rollback posture acknowledged — cluster version upgrades cannot be downgraded; operator must explicitly acknowledge this is one-way.
  7. Maintenance window confirmed — confirm the upgrade is within the approved change window.
  8. Rollout history captured — run kubectl rollout history deployment/<NAME> -n <NAMESPACE> to document the pre-change state for Deployment rollouts.

Required Confirmation

The operator must explicitly state all of the following before any mutation is executed:

  • "I confirm the cluster is <CLUSTER_ID> (<CLUSTER_NAME>) of type <managed/dedicated/serverless> in region <REGION>."
  • "I confirm the target version is <TARGET_VERSION> and I understand cluster version upgrades cannot be downgraded."
  • "I have reviewed PDB status for all workloads on this node pool and no disruption-blocking PDB is present."
  • "I approve this rollout action."

Execution Steps

  1. Capture pre-change state: cluster version, node pool version, all PDB states, Deployment rollout history.
  2. Confirm active RAM principal and policy scope.
  3. Present the planned change and its blast radius to the operator for explicit approval.
  4. Execute the mutation via the ACK console, Alibaba Cloud CLI (aliyun cs), or kubectl as appropriate.
  5. Monitor rollout progress: kubectl rollout status deployment/<NAME> -n <NAMESPACE> or poll the ACK task status via API.
  6. Verify all nodes reach Ready status and all workloads are running post-upgrade.

Rollback Procedure

  • Deployment rollback (reversible): kubectl rollout undo deployment/<NAME> -n <NAMESPACE>
  • Node pool scaling scale-in (partially reversible): New nodes can be added back, but drained workloads need to be rescheduled.
  • Cluster version upgrade (NOT reversible): A completed cluster version upgrade cannot be downgraded. If the upgrade causes issues, address them on the upgraded version or contact Alibaba Cloud support.
  • Document the incident and open an Alibaba Cloud support ticket if cluster corruption is suspected.

Post-Change Verification

  1. Confirm cluster version matches target via ACK console or API.
  2. Run kubectl get nodes — confirm all nodes show Ready with the new version.
  3. Run kubectl get pods --all-namespaces — confirm no pods in CrashLoopBackOff or Pending state.
  4. Run kubectl get pdb --all-namespaces — confirm all PDBs still show healthy disruption budgets.
  5. Verify application health via CloudMonitor metrics: error rate and latency for affected workloads.

Response Shape

  1. Cluster type and version confirmed
  2. Node pool inventory and version status
  3. PDB audit for affected workloads
  4. Rollout strategy
  5. Approval status
  6. Executed action
  7. Post-rollout verification
Files (vanguard-frontier-agentic)
  • references
    • official-sources.md 1.8 KB
      # Official Sources — Alibaba Cloud Live ACK Rollout Guard
      
      Authoritative Alibaba Cloud documentation for ACK rollout, node pool, and upgrade operations.
      
      ## Core References
      
      - **ACK Overview** — https://www.alibabacloud.com/help/en/ack
        Overview of Alibaba Cloud Container Service for Kubernetes (ACK), cluster types, and feature comparison.
      
      - **Upgrade a Cluster** — https://www.alibabacloud.com/help/en/ack/ack-managed-and-dedicated/user-guide/upgrade-a-cluster
        Step-by-step guide for upgrading ACK managed and dedicated cluster versions, including pre-upgrade checks.
      
      - **Managed Node Pools** — https://www.alibabacloud.com/help/en/ack/ack-managed-and-dedicated/user-guide/node-pool-overview
        Node pool architecture, managed upgrade settings, and scaling configuration.
      
      - **Scale a Node Pool** — https://www.alibabacloud.com/help/en/ack/ack-managed-and-dedicated/user-guide/scale-a-node-pool
        How to scale node pools up and down, including drain behavior and auto-scaling policies.
      
      - **Serverless Kubernetes (ASK)** — https://www.alibabacloud.com/help/en/ack/serverless-kubernetes
        Serverless cluster architecture, ECI-based workload scheduling, and version management.
      
      - **PodDisruptionBudget** — https://kubernetes.io/docs/tasks/run-application/configure-pdb/
        Kubernetes upstream documentation for configuring and auditing PDBs before disruptive operations.
      
      - **ACK Audit Logging** — https://www.alibabacloud.com/help/en/ack/ack-managed-and-dedicated/user-guide/audit-log-overview
        Enabling and querying ACK audit logs for cluster and node pool operations.
      
      - **ACK Release Notes** — https://www.alibabacloud.com/help/en/ack/ack-managed-and-dedicated/product-overview/release-notes
        Available Kubernetes versions for ACK clusters, deprecation timelines, and upgrade paths.
      
    • workflow-and-output.md 3.4 KB
      # Workflow and Output — Alibaba Cloud Live ACK Rollout Guard
      
      ## Step-by-Step Workflow
      
      ### Phase 1: Identity and Scope Confirmation
      
      1. Confirm the active RAM principal and its policy scope.
      2. Describe the target cluster to confirm cluster ID, type, region, and current version via the ACK console or API.
      3. List all node pools and their current versions:
         ```
         aliyun cs GET /clusters/<CLUSTER_ID>/nodepools
         ```
      
      ### Phase 2: PDB Audit
      
      4. List all PodDisruptionBudgets across all namespaces:
         ```
         kubectl get pdb --all-namespaces -o wide
         ```
      5. Identify any PDB with `DISRUPTIONS ALLOWED: 0` — these are blocking conditions.
      6. For each blocking PDB, identify the affected workload and its node pool placement.
      
      ### Phase 3: Rollout Strategy Review
      
      7. For node pool scale-in: confirm drain posture — identify any pods with local storage or no eviction toleration.
      8. For Deployment rollouts, review rollout history:
         ```
         kubectl rollout history deployment/<NAME> -n <NAMESPACE>
         ```
      9. For cluster version upgrades: confirm the target version is available and document that the operation is irreversible.
      
      ### Phase 4: Approval Gate
      
      10. Present all evidence to the operator: cluster identity, type, current vs. target version, PDB findings, drain posture.
      11. Require explicit written approval including acknowledgment that cluster version upgrades are irreversible.
      12. Do not proceed until approval is received.
      
      ### Phase 5: Execution
      
      13. Execute the approved operation:
          - Node pool scale via console or:
            ```
            aliyun cs POST /clusters/<CLUSTER_ID>/nodepools/<NODEPOOL_ID>/scale
            ```
          - Cluster version upgrade via console or ACK API.
          - Deployment rollout:
            ```
            kubectl set image deployment/<NAME> <CONTAINER>=<IMAGE> -n <NAMESPACE>
            ```
      14. Monitor the operation:
          ```
          kubectl rollout status deployment/<NAME> -n <NAMESPACE>
          aliyun cs GET /clusters/<CLUSTER_ID>/tasks/<TASK_ID>
          ```
      
      ### Phase 6: Post-Change Verification
      
      15. Confirm all nodes are Ready at the new version:
          ```
          kubectl get nodes -o wide
          ```
      16. Confirm no pods are in error states:
          ```
          kubectl get pods --all-namespaces | grep -v Running | grep -v Completed
          ```
      17. Re-audit PDBs to confirm disruption budgets are healthy.
      18. Check CloudMonitor for elevated error rates or latency anomalies in the 15 minutes following the change.
      
      ## Expected Output Format
      
      The agent response for an ACK rollout operation must include:
      
      ```
      CLUSTER IDENTITY
        Cluster ID:     <cluster-id>
        Cluster Name:   <cluster-name>
        Cluster Type:   <managed / dedicated / serverless>
        Region:         <region>
        Current Version: <x.y.z>
        Target Version:  <x.y.z>
      
      NODE POOL
        Pool ID:        <nodepool-id>
        Pool Name:      <nodepool-name>
        Current Version: <x.y.z>
        Node Count:     <N>
      
      PDB AUDIT
        Blocking PDBs:  [NONE | <list of PDB names with ALLOWED=0>]
        Non-blocking:   <count>
      
      APPROVAL STATUS
        Operator:       <RAM principal>
        Approved:       [YES / NO / PENDING]
        Irreversibility acknowledged: [YES / NO]
      
      ACTION
        [BLOCKED — reason] OR [EXECUTING — task ID] OR [COMPLETE]
      
      ROLLBACK POSTURE
        [NOT POSSIBLE — cluster version upgrade is one-way]
        OR
        [AVAILABLE — kubectl rollout undo deployment/<NAME>]
      
      VERIFICATION
        Nodes Ready:    <count>/<total>
        Pods Healthy:   <count>/<total>
        PDB Status:     OK
        Error Rate:     <value or "not yet checked">
      ```
      
  • metadata.json 1.1 KB
    {
      "id": "alibaba-live-ack-rollout-guard",
      "name": "Alibaba Cloud Live ACK Rollout Guard",
      "version": "0.1.0",
      "type": "skill",
      "provider": "alibaba",
      "harnesses": [
        "codex",
        "copilot",
        "claude-code",
        "cursor",
        "gemini",
        "kiro"
      ],
      "summary": "Gate ACK deployment mutations, node pool scaling, and cluster version upgrades against rollback posture and workload disruption budget before any production change.",
      "source_type": "original",
      "official_docs": [
        "https://www.alibabacloud.com/help/en/ack",
        "https://www.alibabacloud.com/help/en/ack/ack-managed-and-dedicated/user-guide/upgrade-a-cluster"
      ],
      "last_verified": "2026-05-08",
      "path": "skills/alibaba/alibaba-live-ack-rollout-guard",
      "author": "github: VincentChuWaiChow",
      "security_notes": "ACK cluster version downgrade is not supported. Node pool scale-in evicts workloads without PDB compliance check. Addon upgrades can break workloads if version incompatible. Never approve a cluster mutation without explicit rollback posture and disruption budget audit."
    }
    
  • SKILL.md 5.9 KB
    ---
    name: alibaba-live-ack-rollout-guard
    description: Gate ACK deployment mutations, node pool scaling, and cluster version upgrades against rollback posture and workload disruption budget. Prevents irreversible cluster version upgrades from proceeding without PodDisruptionBudget verification, node drain confirmation, and explicit operator approval.
    allowed-tools: Read Grep Glob
    metadata:
      author: "github: VincentChuWaiChow"
      version: "0.1.0"
      updated: "2026-05-08"
      category: delivery
    ---
    
    # Alibaba Cloud Live ACK Rollout Guard
    
    ## Purpose
    
    Act as the guarded live Alibaba Cloud operator for alibaba-live-ack-rollout-guard work. Gate ACK deployment mutations, node pool scaling, and cluster version upgrades. Insist on PDB audit and rollback posture evidence before execution, and treat any ambiguous approval or target as a stop condition.
    
    ## When to Use
    
    Use this skill when:
    
    - An ACK cluster version upgrade is requested (Kubernetes minor or patch version bump)
    - A node pool is being scaled up or down (especially when removing nodes)
    - A Deployment or DaemonSet rollout is being executed against a production workload
    - Node pool configuration changes are planned (instance type, image version, count)
    - An operator needs to audit PodDisruptionBudgets before a disruptive node pool operation
    - An emergency rollback of a broken rollout is required
    
    ## When NOT to Use
    
    Do not use this skill when:
    
    - The target is a non-production cluster with no PDB requirements and no live traffic
    - The task is creating a brand-new cluster (no existing workloads at risk)
    - The task is purely read-only cluster inspection with no mutation intent
    - The task involves Function Compute, SAE, or other non-ACK compute
    
    ## Cluster Type Awareness
    
    ACK supports three cluster types — mutation procedures differ per type:
    
    - **Managed cluster**: Control plane managed by Alibaba Cloud. Node pool upgrades and scaling are the primary mutation surface.
    - **Dedicated cluster**: Full control plane access. Both control plane and data plane versions must be managed.
    - **Serverless cluster (ASK)**: No node pool concept. Pod-level scaling only — ECI instances are provisioned on demand. Version upgrades are less common but follow the same approval gate.
    
    Always confirm cluster type before recommending a mutation path.
    
    ## Pre-Flight Checklist
    
    Before executing any ACK mutation, verify all of the following:
    
    1. **Cluster identity confirmed** — query the ACK API or Alibaba Cloud console to confirm the cluster ID, name, type, and region match the intended target.
    2. **Active RAM principal confirmed** — confirm the active identity has the required RAM policy (`AliyunCSFullAccess` scoped to target cluster) for the operation.
    3. **Current cluster version and node pool version captured** — document both before proceeding; confirm the target version is available for the cluster type.
    4. **PDB audit complete** — run `kubectl get pdb --all-namespaces` and confirm no PDB has `DISRUPTIONS ALLOWED: 0` for workloads running on the affected node pool.
    5. **Node drain posture verified** — for scale-in operations, confirm all nodes to be removed can be safely drained (no pods with no toleration for eviction, no local storage).
    6. **Rollback posture acknowledged** — cluster version upgrades cannot be downgraded; operator must explicitly acknowledge this is one-way.
    7. **Maintenance window confirmed** — confirm the upgrade is within the approved change window.
    8. **Rollout history captured** — run `kubectl rollout history deployment/<NAME> -n <NAMESPACE>` to document the pre-change state for Deployment rollouts.
    
    ## Required Confirmation
    
    The operator must explicitly state all of the following before any mutation is executed:
    
    - "I confirm the cluster is `<CLUSTER_ID>` (`<CLUSTER_NAME>`) of type `<managed/dedicated/serverless>` in region `<REGION>`."
    - "I confirm the target version is `<TARGET_VERSION>` and I understand cluster version upgrades cannot be downgraded."
    - "I have reviewed PDB status for all workloads on this node pool and no disruption-blocking PDB is present."
    - "I approve this rollout action."
    
    ## Execution Steps
    
    1. Capture pre-change state: cluster version, node pool version, all PDB states, Deployment rollout history.
    2. Confirm active RAM principal and policy scope.
    3. Present the planned change and its blast radius to the operator for explicit approval.
    4. Execute the mutation via the ACK console, Alibaba Cloud CLI (`aliyun cs`), or kubectl as appropriate.
    5. Monitor rollout progress: `kubectl rollout status deployment/<NAME> -n <NAMESPACE>` or poll the ACK task status via API.
    6. Verify all nodes reach `Ready` status and all workloads are running post-upgrade.
    
    ## Rollback Procedure
    
    - **Deployment rollback** (reversible): `kubectl rollout undo deployment/<NAME> -n <NAMESPACE>`
    - **Node pool scaling scale-in** (partially reversible): New nodes can be added back, but drained workloads need to be rescheduled.
    - **Cluster version upgrade** (NOT reversible): A completed cluster version upgrade cannot be downgraded. If the upgrade causes issues, address them on the upgraded version or contact Alibaba Cloud support.
    - Document the incident and open an Alibaba Cloud support ticket if cluster corruption is suspected.
    
    ## Post-Change Verification
    
    1. Confirm cluster version matches target via ACK console or API.
    2. Run `kubectl get nodes` — confirm all nodes show `Ready` with the new version.
    3. Run `kubectl get pods --all-namespaces` — confirm no pods in `CrashLoopBackOff` or `Pending` state.
    4. Run `kubectl get pdb --all-namespaces` — confirm all PDBs still show healthy disruption budgets.
    5. Verify application health via CloudMonitor metrics: error rate and latency for affected workloads.
    
    ## Response Shape
    
    1. Cluster type and version confirmed
    2. Node pool inventory and version status
    3. PDB audit for affected workloads
    4. Rollout strategy
    5. Approval status
    6. Executed action
    7. Post-rollout verification
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related