Claude Cursor GitHub Copilot Skill

argo-rollouts-progressive-delivery-review

Use this skill when reviewing Argo Rollouts progressive delivery configuration. Trigger when the user asks about canary or blue-green Rollout strategy correctness, AnalysisTemplate success/failure conditions, traffic weighting provider alignment, canaryService isolation, PDB dead

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_argocd_argo-rollouts-progressive-delivery-review-febe32a.zip · 6 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/argocd/argo-rollouts-progressive-delivery-review
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Argo Rollouts Progressive Delivery Review

Purpose

Review Argo Rollouts canary and blue-green strategy configuration, AnalysisTemplate success and failure condition correctness, traffic management provider alignment, canaryService vs stableService isolation, PDB compatibility with Rollout surge settings, and automated rollback posture. Argo Rollouts' safety depends entirely on AnalysisTemplate conditions that actually fail — an always-true successCondition means automated rollback never fires, regardless of actual error rates.

Lean operating rules

  • Prefer live evidence (kubectl get rollout -A -o yaml, kubectl get analysistemplate -A -o yaml, kubectl argo rollouts status <name>) when the active client exposes it; otherwise fall back to official Argo Rollouts documentation and sanitized YAML from the user.
  • Separate confirmed facts from inference. If AnalysisTemplate metric query results, traffic provider actual behavior, or PDB state was not directly queried, say so.
  • Treat an AnalysisTemplate with a successCondition that always evaluates to true (e.g., result >= 0, true) as a critical finding — automated rollback can never fire.
  • Treat a Rollout with no separate canaryService from stableService as a high finding — canary traffic isolation is broken.
  • Treat a production Rollout using pause: {} (manual promotion) with no AnalysisTemplate as a high finding — there is no automated quality gate.
  • Treat a traffic provider in spec.strategy.canary.trafficRouting that does not match the actual ingress controller installed in the cluster as a high finding — weight changes are silently ignored.
  • Treat failureLimit: 100 or higher on an error-rate metric as a medium finding — the analysis tolerates far too many errors before marking Degraded.
  • Keep the answer scoped, evidence-labeled, and explicit about what was not queried.

References

Load these only when needed:

Response minimum

Return, at minimum:

  • the scoped target (Rollout name, AnalysisTemplate name, or traffic provider config) and evidence level,
  • the deployment strategy (canary with steps vs canary without steps, blue-green) and whether steps include AnalysisRun gates,
  • AnalysisTemplate successCondition and failureCondition correctness,
  • canaryService vs stableService isolation posture,
  • traffic provider alignment with the actual cluster ingress,
  • PDB compatibility with Rollout maxSurge/maxUnavailable,
  • the safest next actions and any assumptions or blockers.
Files (vanguard-frontier-agentic)
  • references
    • workflow-and-output.md 11 KB
      # Workflow and Output Contract
      
      ## Workflow
      
      ### Step 1 — Identify scope and collect raw evidence
      
      1. Confirm the review target: a specific Rollout resource, an AnalysisTemplate, a traffic provider configuration, or a PDB compatibility question.
      2. List all Rollouts and their strategies:
         ```bash
         kubectl get rollout -A -o yaml
         ```
         For each Rollout, note the strategy type (`canary` or `blueGreen`) and whether `spec.strategy.canary.steps` is non-empty.
      3. List all AnalysisTemplates:
         ```bash
         kubectl get analysistemplate -A -o yaml
         kubectl get clusteranalysistemplate -o yaml 2>/dev/null
         ```
      4. Check current Rollout status and any active AnalysisRuns:
         ```bash
         kubectl argo rollouts status <rollout-name> -n <namespace>
         kubectl get analysisrun -A -o yaml
         ```
      
      ### Step 2 — Audit Rollout strategy and steps
      
      A Rollout without steps behaves like a standard Deployment — no progressive traffic shifting occurs.
      
      1. Check whether `spec.strategy.canary.steps` is non-empty and includes analysis gates:
         ```yaml
         # CORRECT: canary with weight steps and analysis gate
         strategy:
           canary:
             canaryService: my-app-canary
             stableService: my-app-stable
             trafficRouting:
               nginx:
                 stableIngress: my-app-ingress
             steps:
               - setWeight: 10
               - pause: {duration: 5m}
               - analysis:
                   templates:
                     - templateName: error-rate-check
               - setWeight: 50
               - pause: {duration: 10m}
               - analysis:
                   templates:
                     - templateName: error-rate-check
      
         # RISKY: no steps — immediately shifts all traffic
         strategy:
           canary:
             maxSurge: "100%"
             maxUnavailable: 0
         ```
      2. Flag as **HIGH** if `maxSurge: 100%` is set with no steps — 100% of replicas are replaced before any analysis runs.
      3. For blue-green Rollouts, check whether `autoPromotionEnabled` is set:
         ```yaml
         # Requires manual promotion
         strategy:
           blueGreen:
             activeService: my-app-active
             previewService: my-app-preview
             autoPromotionEnabled: false
         ```
         `autoPromotionEnabled: true` in production without a `prePromotionAnalysis` is a high finding.
      
      ### Step 3 — Audit AnalysisTemplate success and failure conditions
      
      This is the most critical control — conditions that always evaluate true defeat automated rollback entirely.
      
      1. For each AnalysisTemplate metric, inspect:
         - `spec.metrics[].successCondition` — when is the metric considered passing?
         - `spec.metrics[].failureCondition` — when should it fail?
         - `spec.metrics[].failureLimit` — how many failures are tolerated?
         - `spec.metrics[].provider` — Prometheus, Datadog, web, job, etc.
      2. Example of a correctly configured error-rate AnalysisTemplate:
         ```yaml
         apiVersion: argoproj.io/v1alpha1
         kind: AnalysisTemplate
         metadata:
           name: error-rate-check
         spec:
           metrics:
             - name: error-rate
               interval: 2m
               count: 5
               failureLimit: 0
               provider:
                 prometheus:
                   address: http://prometheus.monitoring.svc.cluster.local:9090
                   query: |
                     sum(rate(http_requests_total{status=~"5..",deployment="{{args.deployment-name}}"}[2m]))
                     /
                     sum(rate(http_requests_total{deployment="{{args.deployment-name}}"}[2m]))
               successCondition: result[0] < 0.01
               failureCondition: result[0] >= 0.05
         ```
      3. Flag as **CRITICAL** if `successCondition` evaluates true for all possible metric values:
         - `result >= 0` (always true for any non-negative counter)
         - `true` (literal boolean true)
         - `result != "error"` (only fails on error, never on bad metric values)
      4. Flag as **HIGH** if `failureCondition` is absent — the metric can only succeed, never explicitly fail.
      5. Flag as **MEDIUM** if `failureLimit` is set to 100 or greater on an error-rate metric — 100 failures will be tolerated before marking Degraded.
      6. Flag as **HIGH** if the Prometheus query template references `{{args.deployment-name}}` but no `args` are passed in the Rollout's analysis step — the query evaluates against all deployments, returning misleading results.
      
      ### Step 4 — Audit canaryService and stableService isolation
      
      Without separate Services, canary pods receive the same traffic distribution as stable — canary traffic isolation does not exist.
      
      1. Check whether both `canaryService` and `stableService` are specified:
         ```bash
         kubectl get rollout <name> -o jsonpath='{.spec.strategy.canary.canaryService},{.spec.strategy.canary.stableService}'
         ```
      2. Verify the Services exist and have the correct selector labels:
         ```bash
         kubectl get svc <canaryService> <stableService> -o yaml | grep -A 5 "selector"
         ```
         Argo Rollouts manages the `rollouts-pod-template-hash` selector on these Services automatically — verify neither has a hardcoded hash that bypasses Rollouts management.
      3. Flag as **HIGH** if `canaryService` is absent — all traffic hits the stable Service regardless of setWeight steps.
      
      ### Step 5 — Audit traffic provider alignment
      
      A misconfigured traffic provider silently ignores all weight changes.
      
      1. Check the traffic routing provider specified in the Rollout:
         ```bash
         kubectl get rollout <name> -o jsonpath='{.spec.strategy.canary.trafficRouting}'
         ```
      2. Verify the specified provider is actually installed:
         ```bash
         # For Istio
         kubectl get virtualservice -A | head -5
         kubectl get destinationrule -A | head -5
      
         # For Nginx
         kubectl get ingressclass | grep nginx
      
         # For AWS ALB
         kubectl get ingressclass | grep alb
      
         # For Traefik
         kubectl get traefikservice -A 2>/dev/null | head -5
         ```
      3. Common mismatches:
         - Rollout specifies `trafficRouting.nginx` but the cluster uses AWS ALB Ingress Controller.
         - Rollout specifies `trafficRouting.istio` but Istio is not installed or not managing the service's namespace.
      4. Flag as **HIGH** if the provider specified does not match installed ingress — weight steps are silently no-ops and all traffic remains on stable.
      
      ### Step 6 — Audit PDB compatibility with Rollout surge settings
      
      A PDB that prevents pod eviction can deadlock a canary rollout that requires replacing existing pods.
      
      1. Check PDBs in the same namespace as the Rollout:
         ```bash
         kubectl get pdb -n <namespace> -o yaml
         ```
      2. Check Rollout maxUnavailable and maxSurge:
         ```bash
         kubectl get rollout <name> -o jsonpath='{.spec.strategy.canary.maxUnavailable},{.spec.strategy.canary.maxSurge}'
         ```
      3. Identify deadlock conditions:
         - `maxUnavailable: 0` in the Rollout means old pods cannot be removed until new pods are Ready.
         - A PDB with `minAvailable: 100%` (or `maxUnavailable: 0`) means no pod can be evicted.
         - Combined: new pods can never start because the cluster has no capacity, and old pods cannot be removed due to PDB — **deadlock**.
      4. Example of a safe PDB configuration alongside a canary Rollout:
         ```yaml
         # PDB: allow 1 unavailable pod during updates
         apiVersion: policy/v1
         kind: PodDisruptionBudget
         metadata:
           name: my-app-pdb
         spec:
           maxUnavailable: 1
           selector:
             matchLabels:
               app: my-app
      
         # Rollout: maxSurge allows creating new pods above desired count
         strategy:
           canary:
             maxSurge: "25%"
             maxUnavailable: 0
         ```
      5. Flag as **HIGH** if `maxUnavailable: 0` in the Rollout and `maxUnavailable: 0` (or `minAvailable: 100%`) in a PDB matching the same pods.
      
      ### Step 7 — Audit rollback posture and history
      
      1. Verify `revisionHistoryLimit` is set to retain enough history for a safe rollback:
         ```bash
         kubectl get rollout <name> -o jsonpath='{.spec.revisionHistoryLimit}'
         ```
         The default is 10. A limit of 1 means only one previous revision is retained — if the rollback target was already overwritten, rollback fails.
      2. Check `abortScaleDownDelaySeconds` for the canary:
         ```bash
         kubectl get rollout <name> -o jsonpath='{.spec.strategy.canary.abortScaleDownDelaySeconds}'
         ```
         Default is 30 seconds. Setting this to 0 means canary pods are immediately deleted on abort — useful for fast rollback but removes the ability to inspect the canary pods post-abort.
      3. To manually trigger a rollback:
         ```bash
         kubectl argo rollouts abort <rollout-name> -n <namespace>
         kubectl argo rollouts undo <rollout-name> -n <namespace>
         ```
      4. Verify automated abort is wired to the AnalysisRun:
         ```bash
         kubectl get analysisrun -A -o yaml | grep -A 5 "phase"
         ```
         An AnalysisRun in `Failed` phase should trigger the Rollout to transition to `Degraded` and initiate rollback automatically.
      
      ### Step 8 — Verify Argo Rollouts controller health
      
      A degraded or missing Argo Rollouts controller means all Rollout objects are frozen — no progression, no rollback, no weight changes.
      
      1. Check controller health:
         ```bash
         kubectl get pods -n argo-rollouts
         kubectl describe deployment argo-rollouts -n argo-rollouts
         ```
      2. Check for recent controller errors:
         ```bash
         kubectl logs -n argo-rollouts -l app.kubernetes.io/name=argo-rollouts --tail=50 | grep -i error
         ```
      3. Flag as **HIGH** if the argo-rollouts controller has unavailable replicas and any Rollout is mid-canary — the canary will not progress or roll back automatically until the controller recovers.
      
      ## Output
      
      Return:
      
      - **target**: Rollout name, namespace, and strategy type, with evidence source,
      - **evidence level**: `live evidence` / `documentation-based` / `sanitized user evidence` / `inference`,
      - **strategy correctness**: steps present/absent, analysis gates present/absent, blue-green autoPromotion setting,
      - **AnalysisTemplate audit**: successCondition and failureCondition correctness, failureLimit values, Prometheus query argument wiring,
      - **service isolation**: canaryService and stableService presence, selector management,
      - **traffic provider alignment**: specified provider vs installed ingress controller,
      - **PDB compatibility**: deadlock risk with Rollout maxSurge/maxUnavailable settings,
      - **rollback posture**: revisionHistoryLimit, abortScaleDownDelaySeconds, automated abort wiring,
      - **controller health**: argo-rollouts controller pod state,
      - **risk findings** (with severity: critical / high / medium / low),
      - **safest next actions** with sample YAML,
      - **assumptions and missing facts**.
      
      ## Security notes
      
      - Never recommend bypassing AnalysisTemplate gates to force a canary promotion — fix the underlying metric or analysis query instead.
      - Never recommend setting `successCondition: true` or equivalent always-passing conditions to unblock a stuck rollout.
      - A Rollout with `autoPromotionEnabled: true` and no `prePromotionAnalysis` in production is equivalent to a standard Deployment — progressive delivery provides no safety gate.
      - Always verify the AnalysisTemplate Prometheus query actually targets the canary deployment specifically, not the entire service or namespace — a query that averages stable and canary traffic can mask canary errors.
      - Do not recommend increasing `failureLimit` as a fix for a legitimate analysis failure — investigate the root cause first.
      
  • metadata.json 1.3 KB
    {
      "id": "argo-rollouts-progressive-delivery-review",
      "name": "Argo Rollouts Progressive Delivery Review",
      "type": "skill",
      "provider": "argocd",
      "harnesses": ["codex", "claude-code", "cursor", "gemini", "kiro", "other"],
      "summary": "Review Argo Rollouts canary and blue-green strategy configuration, AnalysisTemplate success/failure conditions, traffic management provider alignment, canaryService isolation, PDB deadlock risk, and automated rollback posture for progressive delivery safety.",
      "source_type": "original",
      "official_docs": [
        "https://argoproj.github.io/argo-rollouts/",
        "https://argoproj.github.io/argo-rollouts/features/canary/",
        "https://argoproj.github.io/argo-rollouts/features/analysis/",
        "https://argoproj.github.io/argo-rollouts/features/traffic-management/",
        "https://argoproj.github.io/argo-rollouts/features/bluegreen/",
        "https://argoproj.github.io/argo-rollouts/generated/kubectl-argo-rollouts/kubectl-argo-rollouts_promote/"
      ],
      "security_notes": "AnalysisTemplates with always-true success conditions defeat automated rollback entirely. A canary that never fails analysis will silently promote a broken release to 100% production traffic.",
      "last_verified": "2026-05-02",
      "path": "skills/argocd/argo-rollouts-progressive-delivery-review",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.0"
    }
    
  • SKILL.md 3.1 KB
    ---
    name: argo-rollouts-progressive-delivery-review
    description: Use this skill when reviewing Argo Rollouts progressive delivery configuration. Trigger when the user asks about canary or blue-green Rollout strategy correctness, AnalysisTemplate success/failure conditions, traffic weighting provider alignment, canaryService isolation, PDB deadlock risk with Rollout maxSurge settings, automated rollback posture, or manual vs automated promotion configuration.
    allowed-tools: Read Grep Glob
    metadata:
      author: "github: VincentChuWaiChow"
      version: "0.1.0"
      updated: "2026-05-05"
      category: delivery
    ---
    
    # Argo Rollouts Progressive Delivery Review
    
    ## Purpose
    
    Review Argo Rollouts canary and blue-green strategy configuration, AnalysisTemplate success and failure condition correctness, traffic management provider alignment, canaryService vs stableService isolation, PDB compatibility with Rollout surge settings, and automated rollback posture. Argo Rollouts' safety depends entirely on AnalysisTemplate conditions that actually fail — an always-true successCondition means automated rollback never fires, regardless of actual error rates.
    
    ## Lean operating rules
    
    - Prefer live evidence (`kubectl get rollout -A -o yaml`, `kubectl get analysistemplate -A -o yaml`, `kubectl argo rollouts status <name>`) when the active client exposes it; otherwise fall back to official Argo Rollouts documentation and sanitized YAML from the user.
    - Separate confirmed facts from inference. If AnalysisTemplate metric query results, traffic provider actual behavior, or PDB state was not directly queried, say so.
    - Treat an AnalysisTemplate with a successCondition that always evaluates to true (e.g., `result >= 0`, `true`) as a critical finding — automated rollback can never fire.
    - Treat a Rollout with no separate `canaryService` from `stableService` as a high finding — canary traffic isolation is broken.
    - Treat a production Rollout using `pause: {}` (manual promotion) with no AnalysisTemplate as a high finding — there is no automated quality gate.
    - Treat a traffic provider in `spec.strategy.canary.trafficRouting` that does not match the actual ingress controller installed in the cluster as a high finding — weight changes are silently ignored.
    - Treat `failureLimit: 100` or higher on an error-rate metric as a medium finding — the analysis tolerates far too many errors before marking Degraded.
    - Keep the answer scoped, evidence-labeled, and explicit about what was not queried.
    
    ## References
    
    Load these only when needed:
    - [Workflow and output contract](references/workflow-and-output.md)
    
    ## Response minimum
    
    Return, at minimum:
    - the scoped target (Rollout name, AnalysisTemplate name, or traffic provider config) and evidence level,
    - the deployment strategy (canary with steps vs canary without steps, blue-green) and whether steps include AnalysisRun gates,
    - AnalysisTemplate successCondition and failureCondition correctness,
    - canaryService vs stableService isolation posture,
    - traffic provider alignment with the actual cluster ingress,
    - PDB compatibility with Rollout maxSurge/maxUnavailable,
    - the safest next actions and any assumptions or blockers.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related