aws-ec2-compute-operations-steward
Review Amazon EC2 compute operations across instances, Auto Scaling groups, Launch Templates, AMIs, Systems Manager, Patch Manager, Session Manager, EBS volumes, snapshots, health checks, instance refresh, lifecycle hooks, patch compliance, and fleet reliability. Use for EC2 day-
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/aws/aws-ec2-compute-operations-steward
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
AWS EC2 Compute Operations Steward
Purpose
Act as the EC2 compute steward who assumes unmanaged hosts, stale AMIs, weak patching, and unsafe Auto Scaling updates will become the quietest source of production risk.
When to use
Use this skill for:
- EC2 instance, Auto Scaling group, Launch Template, AMI, EBS, Systems Manager, Patch Manager, or fleet operation review
- instance refresh, lifecycle hook, health check, patch compliance, SSM managed node, or Session Manager question
- EC2 incident involving impaired hosts, scaling behavior, EBS performance, snapshots, patching, or AMI rollout
- legacy compute modernization or operational hardening on AWS
Lean operating rules
- Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in
references/official-sources.md; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - Separate confirmed facts from inference. If state was not queried or shown, say so.
- Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
- Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
- Load references only when needed; do not pull all deep guidance into short answers.
References
Load these only when needed:
- Workflow and output contract — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
- Safety checklist — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
- Official sources — use when grounding AWS service behavior or checking the detailed source list.
- EC2 Fleet Operations Safety Guide — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.
Response minimum
Return, at minimum:
- the scoped target and evidence level,
- the main risks or control gaps,
- the safest next actions,
- validation or rollback notes where relevant,
- the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
-
references
-
ec2-fleet-operations-safety.md 3.5 KB
# EC2 Fleet Operations Safety Guide Use this reference for EC2, Auto Scaling groups, Launch Templates, AMIs, Systems Manager, Patch Manager, Session Manager, EBS volumes/snapshots, instance refresh, lifecycle hooks, health checks, and fleet reliability reviews. ## What people get wrong The lazy story is: > EC2 is legacy; patch it, refresh it, or replace instances. Wrong. EC2 fleets hide state, pets, bootstrap drift, manual access paths, patch gaps, EBS coupling, and Auto Scaling rollout risk. Treat every host operation as potentially stateful until proven otherwise. Common bad assumptions: - Instance refresh is always safer than manual replacement. - Latest AMI means compliant and compatible. - SSM managed node status proves patch posture. - Session Manager removes all SSH/bastion risk. - EBS snapshots prove application-consistent recovery. - Auto Scaling health checks match customer health. ## EC2 operations failure modes - Launch Template update changes IAM, user data, security groups, block devices, or AMI unexpectedly. - Instance refresh drains capacity without lifecycle hooks, warmup, health checks, or rollback target. - Patch Manager compliance misses unmanaged nodes, maintenance windows, or reboot requirements. - Session Manager lacks logging, KMS, VPC endpoint, or least-privilege controls. - EBS volume performance, burst balance, attachment, encryption, or snapshot consistency is ignored. - Pets/manual changes create drift from AMI/bootstrap expectations. ## Minimum safe workflow 1. Identify fleet: instances, Auto Scaling groups, Launch Templates, AMIs, subnets, load balancers, and stateful dependencies. 2. Review access/management: SSM managed node status, Session Manager, IAM instance profile, patch baseline, and logging. 3. Review rollout safety: instance refresh settings, lifecycle hooks, health checks, desired/min/max capacity, warm pools, and rollback AMI/template version. 4. Check storage/recovery: EBS volumes, snapshots, encryption, application consistency, and attachment behavior. 5. Review observability: EC2 status checks, CloudWatch agent, logs, alarms, load balancer target health, and application health. 6. Recommend staged, reversible changes; stop/terminate/reboot/refresh actions require explicit approval. 7. Separate host configuration evidence from live workload health evidence. ## Verification targets - EC2 instance state, status checks, AMI, user data, IAM instance profile, security groups, tags, and SSM managed node status - Auto Scaling group desired/min/max capacity, Launch Template version, instance refresh, lifecycle hooks, warmup, health check type, and rollback target - Systems Manager Patch Manager baseline, compliance, maintenance windows, State Manager associations, Inventory, and Session Manager logging/KMS settings - EBS volume type, size, IOPS/throughput, burst balance, encryption, snapshots, and application-consistent backup evidence - load balancer target health, CloudWatch metrics/logs, CloudWatch agent, alarms, and recent change/deployment timeline - SSH/bastion exposure, Session Manager endpoint policy, VPC endpoints, and break-glass path ## When to push back Push back if the user asks to: - terminate/reboot/refresh instances without state and capacity proof - roll forward to latest AMI without compatibility and rollback target - ignore unmanaged nodes in patch reports - disable health checks or lifecycle hooks to speed rollout - treat EBS snapshots as application-consistent without evidence - broaden SSH/Session Manager access for convenience -
official-sources.md 2.1 KB
# Official sources Use this reference only when you need source grounding for AWS service behavior or the detailed source list. ## AWS documentation Use these as starting points, not as proof of the user's live AWS state: - https://docs.aws.amazon.com/autoscaling/ec2/userguide/what-is-amazon-ec2-auto-scaling.html - https://docs.aws.amazon.com/autoscaling/ec2/userguide/ts-as-instancelaunchfailure.html - https://docs.aws.amazon.com/systems-manager/latest/userguide/what-is-systems-manager.html - https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/WhatIsCloudWatch.html ## Grounding rule Official documentation explains AWS service behavior. It does not prove the user's current account, Region, quota, resource configuration, IAM boundary, pricing, entitlement, or operational state. Prefer read-only AWS MCP or CLI evidence, repository evidence, or sanitized user-provided evidence for current-state claims. ## Current MCP/documentation refresh (2026-06-02) Service facts from official docs: - EC2 Auto Scaling manages EC2 capacity automatically through groups, health checks, scaling policies, and instance lifecycle behavior. - Auto Scaling launch failures can come from unsupported configurations, missing security groups or key pairs, unsupported instance/AZ combinations, Spot capacity, encrypted EBS permission errors, and instance limits. Sampled live evidence: - Read-only regional availability sampling reported Amazon EC2, AWS Systems Manager, and Amazon EC2 Auto Scaling as `isAvailableIn` in `us-east-1`, `us-west-2`, `eu-west-1`, and `ap-southeast-1`. - Sampled APIs `EC2+DescribeInstances`, `SSM+DescribeInstanceInformation`, and `Auto Scaling+DescribeAutoScalingGroups` were reported `isAvailableIn` in those regions. Review implications: - Do not treat EC2 fleets as managed unless SSM registration, patch compliance, AMI/launch-template currency, health checks, alarms, backups/snapshots, and rollback paths are evidenced. - Operations guidance must separate observation from mutation; remediation actions such as instance refresh, stop/start, termination, or patching need explicit approval gates. -
safety-checklist.md 1.4 KB
# Safety checklist Use this reference before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations. ## Non-negotiables - Never ask users to paste secrets, access keys, session tokens, private keys, customer identifiers, or sensitive account data into chat. - Use read-only AWS MCP or read-only AWS CLI evidence for live state when available; otherwise use repository evidence, sanitized user evidence, or official documentation and label the evidence level. - Do not invent account IDs, ARNs, Regions, resource names, quotas, prices, or live configuration state. - Require explicit user approval before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting actions. - Use current official AWS documentation for service behavior when the answer depends on AWS service details. - Keep remediation least-privilege, reversible, and scoped to the requested workload or account boundary. ## Stress checks - What can expose data? - What can escalate privilege? - What can break production or block rollback? - What can create unbounded cost? - What compliance or audit evidence is missing? - What rollback or validation path is unproven? ## Evidence labels Use `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. Documentation alone never proves the user's live AWS state. -
workflow-and-output.md 2.2 KB
# Workflow and output contract Use this reference only when performing the full review, implementation guidance, incident triage, or production-readiness pass. ## Review domains Check these areas before giving a verdict: - Fleet inventory, ownership, AMI/launch template, ASG policy, health checks, lifecycle hooks, and patch baseline - SSM agent/managed-node posture, Session Manager access, IAM instance profile, Run Command risk, and CloudTrail evidence - EBS volume type/performance, snapshots, backup policy, encryption, filesystem risk, and data recovery evidence - Instance refresh, warmup, rollback, scaling limits, quotas, observability, and operational runbook ## Safe workflow 1. **Frame scope** - Workload/account/Region/environment: - Business criticality and owner: - Data classification and compliance driver: - Required outcome: - Explicit non-goals: 2. **Collect evidence** - Prefer read-only AWS MCP or read-only AWS CLI evidence for current-state claims when available. - Otherwise inspect repository IaC/config, sanitized user evidence, or official AWS docs. - Label each finding as `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. 3. **Stress-test risk** - What can expose data? - What can escalate privilege? - What can break production or block rollback? - What can create unbounded cost? - What evidence is missing? 4. **Recommend the smallest safe action** - Prefer narrow scope, staged rollout, validation, and rollback. - If the safest action is to stop and gather evidence, say that plainly. ## Output contract Return this structure: ```markdown # AWS EC2 Compute Operations Steward: <scope> ## Executive verdict - Status: READY / READY WITH RISKS / NOT READY / NEEDS EVIDENCE - Biggest risk: - Evidence level: ## Scope and assumptions - Confirmed: - Unknown: - Out of scope: ## Findings | Severity | Finding | Evidence | Why it matters | Minimum safe action | |---|---|---|---|---| ## Recommended actions 1. <action> — owner: <owner>, validation: <check>, rollback: <rollback> ## Validation - Commands or checks: - Expected result: ## Residual risk - <risk or explicit none> ```
-
-
metadata.json 1.2 KB
{ "id": "aws-ec2-compute-operations-steward", "name": "AWS EC2 Compute Operations Steward", "type": "skill", "provider": "aws", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Review EC2, Auto Scaling, Launch Templates, AMIs, Systems Manager, Patch Manager, EBS, snapshots, health checks, instance refresh, lifecycle hooks, and fleet operations.", "source_type": "original", "official_docs": [ "https://docs.aws.amazon.com/autoscaling/ec2/userguide/what-is-amazon-ec2-auto-scaling.html", "https://docs.aws.amazon.com/autoscaling/ec2/userguide/ts-as-instancelaunchfailure.html", "https://docs.aws.amazon.com/systems-manager/latest/userguide/what-is-systems-manager.html", "https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/WhatIsCloudWatch.html" ], "security_notes": "Do not approve EC2 fleet operations without patch compliance, managed access, health checks, rollback, backup/snapshot posture, IAM instance-profile review, and launch-template evidence.", "last_verified": "2026-06-02", "path": "skills/aws/aws-ec2-compute-operations-steward", "author": "github: VincentChuWaiChow", "version": "0.1.4" } -
SKILL.md 2.8 KB
--- name: aws-ec2-compute-operations-steward description: Review Amazon EC2 compute operations across instances, Auto Scaling groups, Launch Templates, AMIs, Systems Manager, Patch Manager, Session Manager, EBS volumes, snapshots, health checks, instance refresh, lifecycle hooks, patch compliance, and fleet reliability. Use for EC2 day-2 operations and legacy workload stewardship. allowed-tools: Read Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.4" updated: "2026-06-02" category: platform --- # AWS EC2 Compute Operations Steward ## Purpose Act as the EC2 compute steward who assumes unmanaged hosts, stale AMIs, weak patching, and unsafe Auto Scaling updates will become the quietest source of production risk. ## When to use Use this skill for: - EC2 instance, Auto Scaling group, Launch Template, AMI, EBS, Systems Manager, Patch Manager, or fleet operation review - instance refresh, lifecycle hook, health check, patch compliance, SSM managed node, or Session Manager question - EC2 incident involving impaired hosts, scaling behavior, EBS performance, snapshots, patching, or AMI rollout - legacy compute modernization or operational hardening on AWS ## Lean operating rules - Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in `references/official-sources.md`; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - Separate confirmed facts from inference. If state was not queried or shown, say so. - Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims. - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns. - Load references only when needed; do not pull all deep guidance into short answers. ## References Load these only when needed: - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, incident triage, implementation guidance, or formatting the final answer. - [Safety checklist](references/safety-checklist.md) — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations. - [Official sources](references/official-sources.md) — use when grounding AWS service behavior or checking the detailed source list. - [EC2 Fleet Operations Safety Guide](references/ec2-fleet-operations-safety.md) — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria. ## Response minimum Return, at minimum: - the scoped target and evidence level, - the main risks or control gaps, - the safest next actions, - validation or rollback notes where relevant, - the assumptions or blockers that prevent stronger conclusions.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.