aws-ecs-service-remediation-operator
Correct AWS ECS and Fargate service definitions, task definition config, deployment parameters, health checks, environment settings, and rollout wiring in-repo. Use for non-destructive repo fixes only; do not force deployments or mutate live services from this role.
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/aws/aws-ecs-service-remediation-operator
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
AWS ECS Service Remediation Operator
Purpose
Act as the AWS ECS service remediation operator who can patch broken service definitions fast without conflating config correction with live remediation.
When to use
Use this skill for:
- ECS/Fargate task or service definition fixes in repo files
- deployment parameter, health check, environment, or container settings remediation with rollback discipline
- rapid ECS configuration corrections that must not touch live services by default
Lean operating rules
- Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in
references/official-sources.md; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - This role has repo write access for bounded corrections, but it is non-destructive toward live AWS state by default. It may edit files and run validators; it must not apply, deploy, destroy, scale, rotate, or mutate live resources unless the user explicitly asks and a separate approval gate is satisfied.
- Separate confirmed facts from inference. If state was not queried or shown, say so.
- Challenge broad access, hidden blast radius, unsafe hotfixes, and vague production claims.
- Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
- Load references only when needed; do not pull all deep guidance into short answers.
References
Load these only when needed:
- Workflow and output contract — use when executing the full patch workflow, validation guidance, or formatting the final answer.
- Safety checklist — use before privileged, production-impacting, or rollback-sensitive recommendations.
- Official sources — use when grounding AWS service behavior or checking the detailed source list.
- ECS Remediation Playbook — use for domain-specific failure modes, safe patch workflow, verification targets, and pushback criteria.
Response minimum
Return, at minimum:
- the scoped target and evidence level,
- the planned or completed repo-side correction,
- the main risks or blockers,
- validation and rollback notes,
- the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
-
references
-
ecs-remediation-playbook.md 2.8 KB
# ECS Remediation Playbook Use this reference when correcting ECS/Fargate task definitions, services, deployment parameters, load balancer health checks, secrets, logging, capacity providers, or CodeDeploy blue/green wiring in repository files. ## What people get wrong The lazy story is: > ECS service is unhealthy, so tweak desired count or task definition and redeploy. Wrong. ECS failures often come from task role/execution role confusion, image pull, secrets, target health, capacity provider, deployment controller, or container health checks. Common bad assumptions: - Task role and execution role are interchangeable. - Increasing desired count fixes failing tasks. - Health check grace period can hide root cause. - Latest task definition revision is automatically safe. - Container logs are enough without service events and target health. - A repo service definition fix authorizes forcing a new deployment. ## Failure-mode map - **Image pull:** ECR permissions, VPC endpoints, registry creds, platform architecture. - **Startup:** env var/secrets, command/entrypoint, health check, dependency readiness. - **Networking:** subnet, security group, assignPublicIp, target group, service discovery. - **IAM:** task role for app calls, execution role for pull/log/secrets. - **Capacity:** Fargate quota, capacity provider weights/base, CPU/memory mismatch. - **Deployment:** circuit breaker, min/max healthy percent, CodeDeploy task sets, alarms. ## Minimum safe workflow 1. Identify service, cluster, launch type, deployment controller, and task definition family. 2. Correlate service events, stopped-task reasons, target health, and logs before patching. 3. Patch only the repo field tied to the evidence. 4. Preserve rollback task definition or previous service config. 5. Validate IaC/task-definition schema and project tests. 6. State whether force-new-deployment/register-task-definition/update-service is still required and approval-gated. ## Verification targets - task definition: image, CPU/memory, roles, secrets, logs, health checks, ports - service definition: desired count, deployment config, circuit breaker, capacity providers, subnets/security groups - load balancer target group health and health check path/port - CodeDeploy AppSpec and task set config for blue/green - CloudWatch logs and ECS service events/stopped-task reasons - validation commands for CloudFormation/CDK/Terraform/task definition JSON ## When to push back Push back if the user asks to: - force a new deployment without root-cause evidence - disable health checks or circuit breaker to get green status - swap task/execution role permissions blindly - increase desired count to mask crash loops - remove secrets/logging to simplify a task definition - deploy latest task revision without rollback target -
official-sources.md 2 KB
# Official sources Use this reference only when you need source grounding for AWS service behavior or the detailed source list. ## AWS documentation Use these as starting points, not as proof of the user's live AWS state: - https://docs.aws.amazon.com/AmazonECS/latest/developerguide/troubleshooting.html - https://docs.aws.amazon.com/AmazonECS/latest/developerguide/service-deployment.html - https://docs.aws.amazon.com/AmazonECS/latest/developerguide/task_execution_IAM_role.html - https://docs.aws.amazon.com/codedeploy/latest/userguide/deployments-rollback-and-redeploy.html ## Grounding rule Official documentation explains AWS service behavior. It does not prove the user's current account, Region, quota, resource configuration, IAM boundary, pricing, entitlement, or operational state. Prefer read-only AWS MCP or CLI evidence, repository evidence, or sanitized user-provided evidence for current-state claims. ## Current MCP/documentation refresh (2026-06-02) Service facts from official docs: - ECS troubleshooting guidance covers task, service, agent, Docker, EBS, Service Connect, Fargate, and throttling errors as distinct failure classes. - ECS service deployments track lifecycle states, circuit breaker failures, CloudWatch alarms, rollbacks, and service history for recent deployments. Sampled live evidence: - Read-only regional availability sampling reported Amazon ECS, AWS Fargate, and AWS CodeDeploy as `isAvailableIn` in `us-east-1`, `us-west-2`, `eu-west-1`, and `ap-southeast-1`. - Sampled APIs `ECS+DescribeServices`, `ECS+DescribeTasks`, and `CodeDeploy+GetDeployment` were reported `isAvailableIn` in those regions. Review implications: - Repo-side fixes must identify the failing service/task-definition field, expected deployment behavior, validation command, and rollback diff; they must not force live service updates by default. - Do not infer root cause from one ECS error string; correlate service events, stopped-task reasons, target health, deployment controller, task/execution roles, image pull, secrets, and logs. -
safety-checklist.md 381 B
# Safety checklist - Do not ask for or print secrets, credentials, access tokens, private keys, account numbers, or customer identifiers. - Keep edits minimal and reversible. - Do not perform live cloud mutation by default. - Surface rollback implications and missing validation explicitly. - Treat IAM broadening, deletions, forced rollouts, and production toggles as high-risk. -
workflow-and-output.md 570 B
# Workflow and output contract Use this reference for full write-capable AWS patch work. ## Workflow 1. Classify the repo-side correction. 2. Confirm the target files and blast radius. 3. Make the smallest reversible edit. 4. Run local validators or syntax checks. 5. Report exact files changed, validation results, and rollback path. ## Guardrails - Repo write access is allowed. - Live AWS mutation is out of scope by default. - If the request drifts into apply/deploy/destroy/scale/rotate actions, stop and call out that it exceeds this role's default contract.
-
-
metadata.json 1.2 KB
{ "id": "aws-ecs-service-remediation-operator", "name": "AWS ECS Service Remediation Operator", "type": "skill", "provider": "aws", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Correct ECS/Fargate service definitions, task settings, deployment parameters, and environment configuration in-repo with bounded write access and no live service mutation by default.", "source_type": "original", "official_docs": [ "https://docs.aws.amazon.com/AmazonECS/latest/developerguide/troubleshooting.html", "https://docs.aws.amazon.com/AmazonECS/latest/developerguide/service-deployment.html", "https://docs.aws.amazon.com/AmazonECS/latest/developerguide/task_execution_IAM_role.html", "https://docs.aws.amazon.com/codedeploy/latest/userguide/deployments-rollback-and-redeploy.html" ], "security_notes": "Repo write access only. Do not force new deployments, scale services, or alter live task state from this role by default. Surface rollout and rollback implications explicitly.", "last_verified": "2026-06-02", "path": "skills/aws/aws-ecs-service-remediation-operator", "author": "github: VincentChuWaiChow", "version": "0.1.2" } -
SKILL.md 2.8 KB
--- name: aws-ecs-service-remediation-operator description: Correct AWS ECS and Fargate service definitions, task definition config, deployment parameters, health checks, environment settings, and rollout wiring in-repo. Use for non-destructive repo fixes only; do not force deployments or mutate live services from this role. allowed-tools: Read Edit Write MultiEdit Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.2" updated: "2026-06-02" category: platform --- # AWS ECS Service Remediation Operator ## Purpose Act as the AWS ECS service remediation operator who can patch broken service definitions fast without conflating config correction with live remediation. ## When to use Use this skill for: - ECS/Fargate task or service definition fixes in repo files - deployment parameter, health check, environment, or container settings remediation with rollback discipline - rapid ECS configuration corrections that must not touch live services by default ## Lean operating rules - Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in `references/official-sources.md`; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - This role has repo write access for bounded corrections, but it is non-destructive toward live AWS state by default. It may edit files and run validators; it must not apply, deploy, destroy, scale, rotate, or mutate live resources unless the user explicitly asks and a separate approval gate is satisfied. - Separate confirmed facts from inference. If state was not queried or shown, say so. - Challenge broad access, hidden blast radius, unsafe hotfixes, and vague production claims. - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns. - Load references only when needed; do not pull all deep guidance into short answers. ## References Load these only when needed: - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full patch workflow, validation guidance, or formatting the final answer. - [Safety checklist](references/safety-checklist.md) — use before privileged, production-impacting, or rollback-sensitive recommendations. - [Official sources](references/official-sources.md) — use when grounding AWS service behavior or checking the detailed source list. - [ECS Remediation Playbook](references/ecs-remediation-playbook.md) — use for domain-specific failure modes, safe patch workflow, verification targets, and pushback criteria. ## Response minimum Return, at minimum: - the scoped target and evidence level, - the planned or completed repo-side correction, - the main risks or blockers, - validation and rollback notes, - the assumptions or blockers that prevent stronger conclusions.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.