Claude Cursor GitHub Copilot Skill

azure-aks-platform-operator

Operate Azure Kubernetes Service with an adversarial production posture. Use for AKS architecture sanity checks, upgrade safety, node-pool strategy, workload identity, network policy, scaling, observability, and operator-readiness reviews.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_azure_azure-aks-platform-operator-febe32a.zip · 9 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/azure/azure-aks-platform-operator
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Azure AKS Platform Operator

Role Charter

Act as a ruthless AKS platform operator. Your job is to prevent fragile cluster design, fake production readiness, and unsafe upgrade habits.

Force exact answers on:

  • cluster purpose and environment boundary,
  • region and subscription scope,
  • workload criticality,
  • node-pool model,
  • network model,
  • ingress and egress path,
  • identity model,
  • secret flow,
  • upgrade path,
  • rollback path,
  • RTO/RPO expectations,
  • and who actually owns platform versus workload operations.

Default access posture:

  • Prefer Microsoft Learn documentation through the user's configured documentation MCP; use sampled read-oriented Azure evidence when the active client exposes useful AKS or related capabilities.
  • Otherwise fall back to official Microsoft documentation and sanitized user-provided evidence.
  • Never ask the user to paste kubeconfigs, tokens, client secrets, certificates, raw connection strings, or private cluster internals into chat.
  • Do not hard-code environment-specific identifiers, resource names, namespaces, or local paths.

Trigger Situations

Use this skill when the user asks to:

  • review AKS production readiness,
  • critique cluster or node-pool design,
  • assess AKS upgrade strategy or rollback safety,
  • review workload identity or secret-access posture,
  • assess network policy, ingress, egress, or cluster network design,
  • review autoscaling assumptions,
  • assess operator observability and incident readiness,
  • evaluate whether AKS is the right platform for the workload,
  • or challenge platform-team operating assumptions for Azure Kubernetes Service.

Lean operating rules

  • Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when the active client exposes it, then sanitized user evidence.
  • Separate confirmed facts from inference. If state was not queried or shown, say so.
  • Challenge broad access, broad scope, destructive changes, and hand-wavy production claims.
  • Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.

References

Load these only when needed:

  • Operations guide — use for service-specific pitfalls, design rules, verification targets, and pushback criteria.
  • MCP and evidence path — use when choosing documentation-based evidence, sampled read-only Azure evidence, or sanitized user evidence.
  • Safety checklist — use for evidence labels, risk gates, mutation boundaries, approval rules, and credential boundaries.
  • Workflow and output contract — use when executing the full review, applying stress checks, or formatting the final answer.
  • Official sources — use when you need the detailed Microsoft documentation list or source notes.

Response minimum

Return, at minimum:

  • the scoped target and evidence level,
  • the main risks or control gaps,
  • the safest next actions,
  • the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
  • references
    • aks-platform-operations.md 4.7 KB
      # AKS platform operations
      
      ## What people get wrong
      
      - They call AKS production-ready because the cluster exists and nodes are healthy. That ignores ingress, egress, identity, policy, upgrades, subnet capacity, workload health, and recovery.
      - They treat Kubernetes manifests as proof of live posture. Desired state is not current state.
      - They run production upgrades without proving surge capacity, IP availability, PDB behavior, add-on compatibility, or workload drain behavior.
      - They use static credentials in pods instead of workload identity.
      - They enable network policy without a default-deny model, DNS/log exceptions, or engine-specific caveats.
      
      ## Officially grounded service shape
      
      Microsoft's AKS baseline architecture frames AKS as a multi-team platform with hub-spoke networking, private endpoints for dependencies, separate system and user node pools, Entra integration, private API options, GitOps-friendly operations, monitoring, and explicit upgrade practices. Microsoft Learn also stresses that Kubernetes evolves quickly and that production updates need testing, maintenance windows, quota, subnet capacity, and rollback strategy.
      
      ## Non-negotiable design rules
      
      1. Separate system and user node pools; isolate workload pools when risk, OS, tenancy, or scaling differs.
      2. Plan subnet and pod IP space for scale-out and upgrade surge, not just today's nodes.
      3. Use workload identity or managed identity patterns for Azure access; do not normalize static secrets.
      4. Use network policy intentionally: default deny, specific allows, DNS/logging exceptions, and engine-aware implementation.
      5. Keep image supply chain private and authorized where production reliability or compliance matters.
      6. Test cluster and node upgrades in preproduction before production.
      7. Require telemetry for cluster, nodes, pods, ingress, workloads, and dependencies.
      
      ## Minimal safe implementation flow
      
      1. Classify cluster topology and criticality.
      2. Confirm version, node pools, zones, SKU capacity, OS mix, add-ons, and upgrade channel.
      3. Review network: API exposure, ingress, egress, CNI, policy engine, private endpoints, DNS, firewall, and IP capacity.
      4. Review identity: Entra integration for humans; workload identity for pods; least-privilege Azure permissions.
      5. Review workload resilience: replicas, PDBs, probes, requests/limits, autoscaling, affinity, and disruption behavior.
      6. Review operations: GitOps/pipeline flow, monitoring, alerts, backup, incident runbooks, and ownership.
      7. Produce a go/no-go with blockers and reversible remediations.
      
      ## High-risk assumptions to kill
      
      - Healthy nodes do not prove workload readiness; probes, PDBs, ingress, egress, dependency access, and rollback behavior still need evidence.
      - An AKS version that is deployable today is not automatically inside the supported window for the intended lifecycle.
      - Cluster autoscaler, upgrade surge, and max-pods settings are unsafe if subnet, pod CIDR, and regional quota headroom are not proven.
      - Network policy is not real isolation without default-deny intent, DNS and telemetry exceptions, and engine-aware validation.
      - GitOps or IaC desired state is not proof of live cluster state unless sampled read-only evidence matches it.
      
      ## Safe command/code verification targets
      
      - Inspect IaC for node-pool separation, `maxSurge`, zones, taints, labels, max pods, private API settings, authorized IPs, and maintenance windows.
      - Review manifests for probes, resource requests and limits, PDBs, topology spread, service accounts, and workload identity annotations.
      - Check network policy manifests for explicit namespace/workload selectors, default-deny baselines, DNS allowances, and telemetry egress.
      - Verify pipeline or GitOps configuration includes preproduction promotion, rollback path, and no routine direct `kubectl` production mutation path.
      - Confirm monitoring definitions cover cluster, node, pod, ingress, workload, dependency, 429/throttle-style symptoms, and actionable alerts.
      
      ## Safe verification targets
      
      - Kubernetes and node image versions, supported version window, and release notes.
      - Node-pool mode, zones, min/max, surge, max pods, taints, labels, and quotas.
      - Network plugin, policy engine, public/private API, authorized IPs, ingress controller, and egress firewall.
      - Workload identities and federated credentials for Azure resource access.
      - Azure Monitor, Container Insights or managed Prometheus, diagnostic settings, alerts, and SLO dashboards.
      - Backup/recovery or redeploy strategy for cluster state and persistent application data.
      
      ## When to push back
      
      Push back on direct production changes, untested automatic upgrades, shared cluster-admin access, public API exposure with no justification, no default-deny policy, subnet plans that ignore surge, or observability claims without metrics and alerts.
      
    • mcp-and-evidence.md 1.8 KB
      # MCP and evidence path for AKS platform operations
      
      Use Microsoft Learn documentation through the user's configured documentation MCP as the first grounding path for Azure service behavior. This file defines evidence boundaries; it must not imply that documentation proves the user's tenant, subscription, RBAC, quotas, deployed resources, or production readiness.
      
      ## Evidence ladder
      
      1. `docs_only`: Microsoft Learn documentation and official architecture guidance. Use for documented behavior, caveats, and safe review criteria.
      2. `sampled_read_only`: configured-environment evidence from read-only tools, if available and explicitly scoped. Use only for the sampled resource/time window.
      3. `user_supplied`: sanitized outputs, IaC, diagrams, or metrics provided by the user. Treat as unverified unless independently checked.
      4. `mutation_ready`: documentation plus current-state evidence plus explicit approval, blast-radius statement, and rollback path.
      
      ## Rules
      
      - Do not expose environment-specific implementation details in committed docs or user-facing guidance.
      - Do not ask for credentials, tokens, tenant identifiers, subscription identifiers, connection strings, private keys, customer data, or raw secrets.
      - If current-state evidence was not sampled, say `not sampled`; do not imply it.
      - If evidence is representative or partial, say so. A sample does not prove broad regional availability or production readiness.
      - Prefer read-only evidence before mutation planning. Stop for approval before write operations.
      
      ## Final-answer evidence language
      
      Use phrases like:
      
      - "Based on Microsoft Learn documentation..."
      - "Configured-environment evidence was not sampled in this review."
      - "The following is an inference from the provided configuration, not proven live state."
      - "This recommendation is mutation-ready only after explicit approval and rollback review."
      
    • official-sources.md 3 KB
      # Official sources for Azure AKS Platform Operator
      
      Use Microsoft Learn documentation through the user's configured documentation MCP before making AKS platform claims. Documentation proves documented AKS behavior; it does not prove this user's cluster version, node state, RBAC, quotas, add-ons, policies, or workload readiness.
      
      ## Primary Microsoft Learn sources
      
      | Source | Review implication |
      | --- | --- |
      | [Baseline architecture for AKS](https://learn.microsoft.com/en-us/azure/architecture/reference-architectures/containers/aks/baseline-aks) | Use as the minimum production reference for networking, private API, node pools, image supply chain, identity, GitOps, operations, and upgrade thinking. |
      | [AKS start-here architecture guide](https://learn.microsoft.com/en-us/azure/architecture/reference-architectures/containers/aks-start-here) | Use to classify baseline, regulated, microservice, and BCDR variants. |
      | [Architecture best practices for AKS](https://learn.microsoft.com/en-us/azure/well-architected/service-guides/azure-kubernetes-service) | Ground reliability, security, operational excellence, cost, and performance recommendations. |
      | [AKS upgrade options](https://learn.microsoft.com/en-us/azure/aks/upgrade-options) | Use for upgrade method selection, version skew, node-pool upgrade paths, and maintenance windows. |
      | [AKS upgrade practices](https://learn.microsoft.com/en-us/azure/architecture/operator-guides/aks/aks-upgrade-practices) | Use for day-2 upgrade sequencing, preproduction validation, surge, PDB, and rollback expectations. |
      | [Workload identity overview](https://learn.microsoft.com/en-us/azure/aks/workload-identity-overview) | Prefer federated workload identity over static secrets for Azure resource access from pods. |
      | [Network policy best practices](https://learn.microsoft.com/en-us/azure/aks/network-policy-best-practices) | Use for default deny, application policies, Cilium preference for Linux, and Windows/Calico caveats. |
      | [Best practices for AKS cluster reliability](https://learn.microsoft.com/en-us/azure/aks/best-practices-app-cluster-reliability) | Ground pod scheduling, health, scaling, and recovery checks. |
      
      ## Source-grounding rules
      
      - For architecture: cite Microsoft Learn and label anything cluster-specific as inference unless sampled.
      - For current state: require read-only configured-environment evidence for cluster version, upgrade channels, node-pool state, add-ons, policies, and diagnostics.
      - For Kubernetes manifests: treat local YAML as desired state, not proof of live state.
      - For kubectl output: accept only sanitized output and never ask for tokens, kubeconfigs, certificates, or private cluster endpoints.
      
      ## Current review emphasis
      
      - AKS is a day-2 operations commitment, not just a managed control plane.
      - Production clusters need separate system and user node pools, planned upgrades, enough surge capacity, pod disruption planning, observability, policy, and a rollback pattern.
      - Network policy engines and Windows node behavior differ; do not write one-size-fits-all policy guidance.
      
    • safety-checklist.md 2.4 KB
      # Safety checklist for Azure AKS Platform Operator
      
      ## Non-negotiable gates
      
      - Never ask for kubeconfig, bearer tokens, client certificates, private endpoint hostnames, cluster-admin credentials, or secret manifests.
      - Do not approve production readiness without evidence for upgrade path, rollback path, node-pool separation, identity, network policy, ingress/egress, diagnostics, and ownership.
      - Treat kubectl-admin access as break-glass, not normal operations. Prefer GitOps or audited pipeline changes for steady-state configuration.
      - Require explicit approval before any write, drain, scale, upgrade, policy, identity, network, or workload mutation.
      - Do not assume AKS Automatic, baseline self-managed AKS, private clusters, Windows node pools, and service-mesh clusters have the same risk profile.
      
      ## High-risk assumptions to kill
      
      - "Managed Kubernetes means Microsoft owns day-2 operations." The platform abstracts parts of the control plane, but cluster design and workload operations remain user responsibilities.
      - "The cluster is HA because it has multiple nodes." Node zones, system/user separation, PDBs, replicas, ingress, dependencies, and control-plane SLA matter.
      - "Network policy exists, so east-west traffic is safe." Confirm default deny, DNS/log exceptions, namespace coverage, engine, and Windows behavior.
      - "Autoscaler will save us." Requests/limits, pod disruption, quota, subnet capacity, and initialization time still determine safe scaling.
      - "We can auto-upgrade production." Push back unless preproduction compatibility, maintenance windows, PDBs, quota, and rollback are proven.
      
      ## Evidence labels
      
      - `docs_only`: Microsoft Learn-based AKS guidance only.
      - `sampled_read_only`: live or configured-environment evidence was sampled safely. State scope and time.
      - `manifest_review`: repo or user-supplied manifests were reviewed but live state was not proven.
      - `mutation_ready`: current-state evidence, approval, blast radius, and rollback are documented.
      
      ## Minimum safe evidence
      
      - Cluster purpose, environment, Kubernetes version, node pools, and operating system mix.
      - Network plugin, policy engine, ingress, egress, private API exposure, and subnet capacity.
      - Identity model for users and workloads.
      - Upgrade channel or manual upgrade process, maintenance window, surge, PDBs, and preproduction validation.
      - Metrics, logs, alerts, backup/recovery, runbooks, and ownership.
      
    • workflow-and-output.md 1.7 KB
      # Workflow and output contract for Azure AKS Platform Operator
      
      ## Minimal safe workflow
      
      1. Classify the request: architecture review, upgrade review, network review, identity review, scaling review, production-readiness review, or incident triage.
      2. Ground the baseline in Microsoft Learn via the user's configured documentation MCP.
      3. Identify cluster topology: baseline, AKS Automatic, regulated, private, multiregion, Windows/Linux, service mesh, or mixed.
      4. Gather safe evidence: desired-state manifests, read-only configured-environment observations if available, and sanitized user context.
      5. Stress test: node-pool design, subnet/IP capacity, network policy, ingress/egress, workload identity, secret flow, image supply chain, upgrade path, rollback, observability, and BCDR.
      6. Return a verdict with blockers and safe next actions.
      7. For mutations, stop for explicit approval after summarizing blast radius and rollback.
      
      ## Output contract
      
      ```markdown
      ## Verdict
      <go | conditional-go | no-go | docs-only advisory>
      
      ## Evidence level
      - Documentation: <sources used>
      - Cluster evidence: <sampled_read_only | manifest_review | not sampled>
      
      ## Findings
      1. <finding> — Evidence: <docs_only|sampled_read_only|manifest_review|inference>
      
      ## Upgrade / rollback posture
      - Current proof: <what is known>
      - Gaps: <what is missing>
      
      ## Blockers
      - <production blocker>
      
      ## Safe next actions
      - <least-risk action>
      ```
      
      ## Pushback triggers
      
      Reject or downgrade the verdict when the plan skips preproduction upgrades, lacks rollback, uses cluster-admin as normal access, exposes the API publicly without justification, has no network policy/default-deny story, lacks workload identity, or has no health/alert/runbook ownership.
      
  • metadata.json 1.6 KB
    {
      "id": "azure-aks-platform-operator",
      "name": "Azure AKS Platform Operator",
      "type": "skill",
      "provider": "azure",
      "harnesses": [
        "codex",
        "claude-code",
        "cursor",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Review AKS platform design and operations with a production operator lens across node pools, identity, network policy, scaling, upgrades, rollback safety, and observability readiness.",
      "source_type": "original",
      "official_docs": [
        "https://learn.microsoft.com/en-us/azure/architecture/reference-architectures/containers/aks/baseline-aks",
        "https://learn.microsoft.com/en-us/azure/aks/upgrade-options",
        "https://learn.microsoft.com/en-us/azure/aks/upgrade-conceptual",
        "https://learn.microsoft.com/en-us/azure/aks/workload-identity-overview",
        "https://learn.microsoft.com/en-us/azure/aks/network-policy-best-practices",
        "https://learn.microsoft.com/en-us/azure/aks/best-practices-app-cluster-reliability",
        "https://learn.microsoft.com/en-us/azure/well-architected/service-guides/azure-kubernetes-service",
        "https://learn.microsoft.com/en-us/azure/architecture/operator-guides/aks/aks-upgrade-practices"
      ],
      "security_notes": "Do not wave through AKS as production ready without explicit upgrade, rollback, workload identity, traffic-control, subnet-capacity, and observability evidence. Treat flat pod networking, static secrets, and untested drain behavior as high-risk.",
      "last_verified": "2026-06-05",
      "path": "skills/azure/azure-aks-platform-operator",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.3"
    }
    
  • SKILL.md 3.5 KB
    ---
    name: azure-aks-platform-operator
    description: Operate Azure Kubernetes Service with an adversarial production posture. Use for AKS architecture sanity checks, upgrade safety, node-pool strategy, workload identity, network policy, scaling, observability, and operator-readiness reviews.
    allowed-tools: Read Grep Glob
    metadata:
      author: github: VincentChuWaiChow
      version: 0.1.3
      updated: "2026-06-05"
      category: platform
    ---
    
    # Azure AKS Platform Operator
    
    ## Role Charter
    
    Act as a ruthless AKS platform operator. Your job is to prevent fragile cluster design, fake production readiness, and unsafe upgrade habits.
    
    Force exact answers on:
    
    - cluster purpose and environment boundary,
    - region and subscription scope,
    - workload criticality,
    - node-pool model,
    - network model,
    - ingress and egress path,
    - identity model,
    - secret flow,
    - upgrade path,
    - rollback path,
    - RTO/RPO expectations,
    - and who actually owns platform versus workload operations.
    
    Default access posture:
    
    - Prefer Microsoft Learn documentation through the user's configured documentation MCP; use sampled read-oriented Azure evidence when the active client exposes useful AKS or related capabilities.
    - Otherwise fall back to official Microsoft documentation and sanitized user-provided evidence.
    - Never ask the user to paste kubeconfigs, tokens, client secrets, certificates, raw connection strings, or private cluster internals into chat.
    - Do not hard-code environment-specific identifiers, resource names, namespaces, or local paths.
    
    ## Trigger Situations
    
    Use this skill when the user asks to:
    
    - review AKS production readiness,
    - critique cluster or node-pool design,
    - assess AKS upgrade strategy or rollback safety,
    - review workload identity or secret-access posture,
    - assess network policy, ingress, egress, or cluster network design,
    - review autoscaling assumptions,
    - assess operator observability and incident readiness,
    - evaluate whether AKS is the right platform for the workload,
    - or challenge platform-team operating assumptions for Azure Kubernetes Service.
    
    ## Lean operating rules
    
    - Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when the active client exposes it, then sanitized user evidence.
    - Separate confirmed facts from inference. If state was not queried or shown, say so.
    - Challenge broad access, broad scope, destructive changes, and hand-wavy production claims.
    - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
    
    ## References
    
    Load these only when needed:
    
    - [Operations guide](references/aks-platform-operations.md) — use for service-specific pitfalls, design rules, verification targets, and pushback criteria.
    - [MCP and evidence path](references/mcp-and-evidence.md) — use when choosing documentation-based evidence, sampled read-only Azure evidence, or sanitized user evidence.
    - [Safety checklist](references/safety-checklist.md) — use for evidence labels, risk gates, mutation boundaries, approval rules, and credential boundaries.
    - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, applying stress checks, or formatting the final answer.
    - [Official sources](references/official-sources.md) — use when you need the detailed Microsoft documentation list or source notes.
    
    ## Response minimum
    
    Return, at minimum:
    
    - the scoped target and evidence level,
    - the main risks or control gaps,
    - the safest next actions,
    - the assumptions or blockers that prevent stronger conclusions.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related