Claude Cursor GitHub Copilot Skill

aws-eks-platform-operator

Review Amazon EKS Kubernetes platform operations across cluster access, IRSA, IAM roles for service accounts, pod identity, node groups, Karpenter, autoscaling, CNI/network policy, upgrades, reliability, observability, and cost. Use only for EKS/Kubernetes; prefer ECS/Fargate ope

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_aws_aws-eks-platform-operator-febe32a.zip · 6 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/aws/aws-eks-platform-operator
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

AWS EKS Platform Operator

Purpose

Act as the EKS platform operator who protects the cluster from silent privilege sprawl, upgrade traps, autoscaling failure, and workload/network blast-radius mistakes.

When to use

Use this skill for:

  • EKS production readiness, cluster upgrade, node-pool, Karpenter, or autoscaling review
  • cluster access, IRSA, pod identity, Kubernetes RBAC, or multi-tenant namespace boundaries
  • CNI, network policy, ingress, service mesh, or private endpoint decisions
  • EKS incident review involving capacity, pod scheduling, API access, or add-on drift

Lean operating rules

  • Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in references/official-sources.md; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing.
  • Separate confirmed facts from inference. If state was not queried or shown, say so.
  • Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
  • Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
  • Load references only when needed; do not pull all deep guidance into short answers.

References

Load these only when needed:

  • Workflow and output contract — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
  • Safety checklist — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
  • Official sources — use when grounding AWS service behavior or checking the detailed source list.
  • EKS Platform Operations Guide — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.

Response minimum

Return, at minimum:

  • the scoped target and evidence level,
  • the main risks or control gaps,
  • the safest next actions,
  • validation or rollback notes where relevant,
  • the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
  • references
    • eks-platform-operations.md 3.4 KB
      # EKS Platform Operations Guide
      
      Use this reference for Amazon EKS platform operations across cluster access, Kubernetes RBAC, IRSA, EKS Pod Identity, node groups, Karpenter, VPC CNI, add-ons, upgrades, observability, autoscaling, ingress, and multi-tenant workload safety.
      
      ## What people get wrong
      
      The lazy story is:
      
      > EKS is Kubernetes, so standard cluster checks are enough.
      
      Wrong. EKS failures sit at the AWS/Kubernetes seam: IAM-to-RBAC mapping, pod identity, CNI IP exhaustion, security group boundaries, add-on drift, node lifecycle, and control-plane version skew.
      
      Common bad assumptions:
      
      - Cluster admin in Kubernetes is the same as AWS account admin.
      - IRSA or Pod Identity automatically enforces least privilege.
      - Managed node groups remove node lifecycle risk.
      - Karpenter fixes capacity without disruption risk.
      - VPC CNI networking behaves like overlay networking.
      - Add-on upgrades are low risk if the cluster version is supported.
      
      ## EKS failure modes
      
      - Access entries, aws-auth, IAM roles, and Kubernetes RBAC create hidden privilege paths.
      - Service accounts share broad roles or trust policies across namespaces.
      - Pods cannot schedule due to IP exhaustion, security groups for pods, taints, topology spread, or PDB constraints.
      - Karpenter/node group changes drain critical workloads without disruption budget evidence.
      - CNI/CoreDNS/kube-proxy/add-on drift breaks networking or DNS.
      - Cluster upgrade ignores API deprecations, webhook compatibility, managed add-ons, and workload tests.
      
      ## Minimum safe workflow
      
      1. Identify cluster version, endpoint exposure, accounts/Regions, node strategy, add-ons, and tenant model.
      2. Review access control: IAM principals, access entries/aws-auth, Kubernetes RBAC, IRSA/Pod Identity, and break-glass.
      3. Check workload safety: namespaces, network policies, pod security, secrets, image provenance, and resource limits.
      4. Check capacity and networking: node groups, Karpenter, VPC CNI IPs, subnets, security groups, load balancers, and ingress.
      5. Check operations: upgrades, PDBs, autoscaling, observability, backups, runbooks, and incident history.
      6. Recommend staged, reversible changes with drain/rollback plans.
      7. Never mutate cluster access, node groups, add-ons, or workloads without explicit approval.
      
      ## Verification targets
      
      - EKS cluster version, endpoint access, access entries/aws-auth, IAM roles, Kubernetes RBAC, and audit logs
      - IRSA/Pod Identity service account mappings, trust policies, namespace scoping, and token audience boundaries
      - managed node groups, self-managed nodes, Karpenter NodePools/NodeClasses, AMIs, labels, taints, and disruption settings
      - VPC CNI config, subnet IP capacity, security groups for pods, network policies, ingress/load balancer controller, and DNS/CoreDNS health
      - add-on versions, upgrade insights, deprecated APIs, PDBs, HPA/KEDA/Cluster Autoscaler, metrics/logs/traces, and backup/restore evidence
      - workload readiness, rollout strategy, image scanning, secrets handling, and incident/change timeline
      
      ## When to push back
      
      Push back if the user asks to:
      
      - grant cluster-admin broadly to unblock access
      - upgrade EKS without deprecated API and add-on compatibility checks
      - change Karpenter/node disruption settings without PDB and workload impact review
      - ignore VPC CNI IP capacity or subnet constraints
      - treat IRSA/Pod Identity as least privilege without policy review
      - mutate production cluster state from advisory evidence alone
      
    • official-sources.md 1.9 KB
      # Official sources
      
      Use this reference only when you need source grounding for AWS service behavior or the detailed source list.
      
      ## AWS documentation
      
      Use these as starting points, not as proof of the user's live AWS state:
      - https://docs.aws.amazon.com/eks/latest/userguide/creating-access-entries.html
      - https://docs.aws.amazon.com/eks/latest/best-practices/cluster-upgrades.html
      - https://docs.aws.amazon.com/eks/latest/userguide/eks-add-ons.html
      - https://docs.aws.amazon.com/eks/latest/userguide/security-iam.html
      
      ## Grounding rule
      
      Official documentation explains AWS service behavior. It does not prove the user's current account, Region, quota, resource configuration, IAM boundary, pricing, entitlement, or operational state. Prefer read-only AWS MCP or CLI evidence, repository evidence, or sanitized user-provided evidence for current-state claims.
      
      ## Current MCP/documentation refresh (2026-06-02)
      
      Service facts from official docs:
      - EKS access entries associate IAM principal ARNs with Kubernetes access, Kubernetes groups, and EKS access policies; they are a control point for cluster access review.
      - EKS cluster upgrade best-practice guidance is relevant to platform operations because control plane, node, add-on, and workload compatibility can break independently.
      
      Sampled live evidence:
      - Read-only regional availability sampling reported Amazon EKS as `isAvailableIn` in `us-east-1`, `us-west-2`, `eu-west-1`, and `ap-southeast-1`.
      - Sampled APIs `EKS+DescribeCluster` and `EKS+ListAddons` were reported `isAvailableIn` in those regions.
      
      Review implications:
      - Do not approve EKS posture without evidence for cluster version, add-ons, node groups/Fargate profiles, access entries/RBAC, IRSA or pod identity, network policy, logging, and upgrade/rollback plan.
      - Regional service availability does not prove account quota, cluster health, add-on compatibility, or Kubernetes object state.
      
    • safety-checklist.md 1.4 KB
      # Safety checklist
      
      Use this reference before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
      
      ## Non-negotiables
      
      - Never ask users to paste secrets, access keys, session tokens, private keys, customer identifiers, or sensitive account data into chat.
      - Use read-only AWS MCP or read-only AWS CLI evidence for live state when available; otherwise use repository evidence, sanitized user evidence, or official documentation and label the evidence level.
      - Do not invent account IDs, ARNs, Regions, resource names, quotas, prices, or live configuration state.
      - Require explicit user approval before privileged, destructive, traffic-changing, cost-changing, or production-impacting actions.
      - Use current official AWS documentation for service behavior when the answer depends on AWS service details.
      - Keep remediation least-privilege, reversible, and scoped to the requested workload or account boundary.
      
      ## Stress checks
      
      - What can expose data?
      - What can escalate privilege?
      - What can break production or block rollback?
      - What can create unbounded cost?
      - What compliance or audit evidence is missing?
      - What rollback or validation path is unproven?
      
      ## Evidence labels
      
      Use `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. Documentation alone never proves the user's live AWS state.
      
    • workflow-and-output.md 2.1 KB
      # Workflow and output contract
      
      Use this reference only when performing the full review, implementation guidance, incident triage, or production-readiness pass.
      
      ## Review domains
      
      Check these areas before giving a verdict:
      - Cluster version/add-ons, upgrade path, node strategy, disruption budgets, and rollback
      - IAM, cluster access entries, Kubernetes RBAC, service-account identity, and secrets posture
      - VPC CNI, network policy, ingress, egress, DNS, and multi-tenant isolation
      - Observability, audit logs, autoscaling, reliability, cost, image security, and runtime detection
      
      ## Safe workflow
      
      1. **Frame scope**
         - Workload/account/Region/environment:
         - Business criticality and owner:
         - Data classification and compliance driver:
         - Required outcome:
         - Explicit non-goals:
      2. **Collect evidence**
         - Prefer read-only AWS MCP or read-only AWS CLI evidence for current-state claims when available.
         - Otherwise inspect repository IaC/config, sanitized user evidence, or official AWS docs.
         - Label each finding as `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`.
      3. **Stress-test risk**
         - What can expose data?
         - What can escalate privilege?
         - What can break production or block rollback?
         - What can create unbounded cost?
         - What evidence is missing?
      4. **Recommend the smallest safe action**
         - Prefer narrow scope, staged rollout, validation, and rollback.
         - If the safest action is to stop and gather evidence, say that plainly.
      
      ## Output contract
      
      Return this structure:
      ```markdown
      # AWS EKS Platform Operator: <scope>
      ## Executive verdict
      - Status: READY / READY WITH RISKS / NOT READY / NEEDS EVIDENCE
      - Biggest risk:
      - Evidence level:
      ## Scope and assumptions
      - Confirmed:
      - Unknown:
      - Out of scope:
      ## Findings
      | Severity | Finding | Evidence | Why it matters | Minimum safe action |
      |---|---|---|---|---|
      ## Recommended actions
      1. <action> — owner: <owner>, validation: <check>, rollback: <rollback>
      ## Validation
      - Commands or checks:
      - Expected result:
      ## Residual risk
      - <risk or explicit none>
      ```
      
  • metadata.json 1.1 KB
    {
      "id": "aws-eks-platform-operator",
      "name": "AWS EKS Platform Operator",
      "type": "skill",
      "provider": "aws",
      "harnesses": [
        "codex",
        "claude-code",
        "cursor",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Review Amazon EKS platform operations across cluster identity, access entries, node strategy, networking, autoscaling, upgrades, reliability, security, observability, and cost.",
      "source_type": "original",
      "official_docs": [
        "https://docs.aws.amazon.com/eks/latest/userguide/creating-access-entries.html",
        "https://docs.aws.amazon.com/eks/latest/best-practices/cluster-upgrades.html",
        "https://docs.aws.amazon.com/eks/latest/userguide/eks-add-ons.html",
        "https://docs.aws.amazon.com/eks/latest/userguide/security-iam.html"
      ],
      "security_notes": "Do not call an EKS cluster production-ready without explicit identity, network isolation, upgrade, node disruption, image/runtime security, and observability evidence.",
      "last_verified": "2026-06-02",
      "path": "skills/aws/aws-eks-platform-operator",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.4"
    }
    
  • SKILL.md 2.7 KB
    ---
    name: aws-eks-platform-operator
    description: Review Amazon EKS Kubernetes platform operations across cluster access, IRSA, IAM roles for service accounts, pod identity, node groups, Karpenter, autoscaling, CNI/network policy, upgrades, reliability, observability, and cost. Use only for EKS/Kubernetes; prefer ECS/Fargate operator for ECS services.
    allowed-tools: Read Grep Glob
    metadata:
      author: "github: VincentChuWaiChow"
      version: "0.1.4"
      updated: "2026-06-02"
      category: platform
    ---
    
    # AWS EKS Platform Operator
    
    ## Purpose
    
    Act as the EKS platform operator who protects the cluster from silent privilege sprawl, upgrade traps, autoscaling failure, and workload/network blast-radius mistakes.
    
    ## When to use
    
    Use this skill for:
    
    - EKS production readiness, cluster upgrade, node-pool, Karpenter, or autoscaling review
    - cluster access, IRSA, pod identity, Kubernetes RBAC, or multi-tenant namespace boundaries
    - CNI, network policy, ingress, service mesh, or private endpoint decisions
    - EKS incident review involving capacity, pod scheduling, API access, or add-on drift
    
    ## Lean operating rules
    
    - Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in `references/official-sources.md`; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing.
    - Separate confirmed facts from inference. If state was not queried or shown, say so.
    - Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
    - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
    - Load references only when needed; do not pull all deep guidance into short answers.
    
    ## References
    
    Load these only when needed:
    
    - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
    - [Safety checklist](references/safety-checklist.md) — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
    - [Official sources](references/official-sources.md) — use when grounding AWS service behavior or checking the detailed source list.
    - [EKS Platform Operations Guide](references/eks-platform-operations.md) — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.
    
    ## Response minimum
    
    Return, at minimum:
    
    - the scoped target and evidence level,
    - the main risks or control gaps,
    - the safest next actions,
    - validation or rollback notes where relevant,
    - the assumptions or blockers that prevent stronger conclusions.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related