aws-waf-reliability-review
Review AWS workload reliability posture against the Well-Architected Framework Reliability Pillar. Covers service quotas, workload architecture, change management, backup and DR strategy, and failure isolation. Use when auditing availability design, planning disaster recovery, or
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/aws/aws-waf-reliability-review
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
AWS WAF Reliability Pillar Review
Purpose
Act as the AWS WAF Reliability Pillar reviewer — assess workload resilience against the five reliability design principles and produce actionable recommendations for improving availability, recovery, and change safety.
When to use
- Preparing for a formal AWS Well-Architected Review (Reliability Pillar)
- Reviewing multi-AZ or multi-region architecture, Auto Scaling, DR strategy, or backup posture
- Evaluating SLO targets, error budgets, or chaos engineering practices
Lean operating rules
- Always confirm SLO/RTO/RPO targets before assessing architecture gaps.
- Prefer
AwsDocumentationMcpServerwhen available. Otherwise fall back to official docs. - Separate confirmed facts from inference. If state was not queried, say so.
- Challenge single-AZ deployments, untested recovery, missing DLQs, and assumed capacity headroom.
- Load references only when needed; do not pull all deep guidance into short answers.
References
Load these only when needed:
- Workflow and output contract — use when executing the full WAF reliability review, formatting findings, or generating the final assessment report.
- Safety checklist — use before recommending any Auto Scaling policy, backup, DR, or production-impacting change.
- Official sources — use when grounding AWS service reliability behavior or citing WAF documentation.
- Well-Architected Reliability Review Guide — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.
Files (vanguard-frontier-agentic)
-
references
-
official-sources.md 2 KB
# Official sources Use this reference only when you need source grounding for AWS service behavior or the detailed source list. ## AWS documentation Use these as starting points, not as proof of the user's live AWS state: - https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/welcome.html - https://docs.aws.amazon.com/wellarchitected/latest/framework/reliability.html - https://docs.aws.amazon.com/wellarchitected/2023-10-03/framework/rel_planning_for_recovery_disaster_recovery.html - https://docs.aws.amazon.com/resilience-hub/latest/userguide/resilience-checks.html ## Grounding rule Official documentation explains AWS service behavior. It does not prove the user's current account, Region, quota, resource configuration, IAM boundary, pricing, entitlement, or operational state. Prefer read-only AWS MCP or CLI evidence, repository evidence, or sanitized user-provided evidence for current-state claims. ## Current MCP/documentation refresh (2026-06-02) Service facts from official docs: - The Well-Architected Reliability Pillar focuses on designing, delivering, and maintaining workloads that perform their intended functions correctly and consistently. - Reliability guidance includes foundations, workload architecture, change management, failure management, and DR strategies to meet recovery objectives. Sampled live evidence: - Read-only API availability sampling reported `WellArchitected+GetWorkload` as `isAvailableIn` in `us-east-1`, `us-west-2`, `eu-west-1`, and `ap-southeast-1`. - Read-only product availability sampling also reported AWS Resilience Hub as `isAvailableIn` in those regions, but that does not prove a workload has an assessed or passing resilience policy. Review implications: - Require workload dependency map, quotas, autoscaling, health checks, change safety, backup/restore, failover tests, RTO/RPO, and operational runbooks. - Do not infer resilience from architecture diagrams or API availability without exercised failure and recovery evidence. -
safety-checklist.md 1.2 KB
# Safety Checklist Use before recommending any Auto Scaling policy change, backup schedule modification, DR configuration, or production-impacting reliability control. ## Non-negotiables - Never recommend deleting backups, reducing backup retention, or disabling Multi-AZ without explicit confirmation of business risk acceptance. - Do not invent SLO values, SLA percentages, resource names, or ARNs. - Require explicit user approval before modifying Auto Scaling policies, health check thresholds, or DR routing configurations in production. - Chaos engineering experiments (AWS FIS) must run in non-production first; flag this requirement explicitly. - SQS DLQ enablement and retry policies are non-destructive in most cases — but validate queue consumer idempotency before recommending increased retry counts. - Route 53 failover routing changes affect live DNS TTL — require confirmation of TTL values and client cache flush plans. ## Stress checks - What is the single point of failure if this change is applied? - What is the blast radius of the recommended change? - Does the recovery test account for data replication lag? - Is the DLQ consumer tested and alerting configured for DLQ depth? - What is the rollback plan if Auto Scaling acts unexpectedly? -
well-architected-reliability-review.md 3.1 KB
# Well-Architected Reliability Review Guide Use this reference for AWS Well-Architected Framework Reliability Pillar reviews. In this repository, `aws-waf-reliability-review` means Well-Architected Framework review, not AWS Web Application Firewall configuration. ## What people get wrong The lazy story is: > Multi-AZ plus backups equals reliable. Wrong. Reliability is measured against explicit objectives and exercised recovery. Architecture diagrams, service defaults, and unused backups do not prove resilience. Common bad assumptions: - Multi-AZ deployment proves failure isolation. - Auto Scaling means capacity headroom exists. - Backups prove restore capability. - DR plan proves RTO/RPO compliance. - Queue/DLQ presence proves downstream recovery. - Well-Architected questionnaire answers equal implementation evidence. ## Reliability-specific failure modes - No explicit SLO, RTO, RPO, dependency tier, or error budget. - Single points of failure hidden in DNS, KMS, NAT, VPC endpoints, IAM roles, data stores, or third-party dependencies. - Quotas, throttling, and retry storms are not modeled. - Change safety lacks canaries, rollback alarms, deployment circuit breakers, or game-day evidence. - Backup retention exists but restore order, permissions, and application consistency are untested. - Regional failover path conflicts with data replication lag, identity, DNS TTL, or runbook ownership. ## Minimum safe workflow 1. Confirm workload scope, business criticality, SLO, RTO, RPO, peak load, and dependency map. 2. Classify failure domains: Availability Zone, Region, account, service dependency, data plane, control plane, and human/operator path. 3. Review workload architecture, quotas, health checks, scaling, backpressure, retries, queues, and stateful components. 4. Inspect change-management controls: staged rollout, rollback, alarms, deployment history, and incident correlation. 5. Verify backup/restore and DR test evidence against RTO/RPO, not just configuration. 6. Produce findings with severity, evidence level, risk, recommended validation, and owner. 7. Do not mark reliability ready without exercised failure/recovery evidence. ## Verification targets - SLO/RTO/RPO and customer-impact definitions - dependency map for compute, network, identity, data, DNS, KMS, queues, and third parties - quotas, scaling policies, circuit breakers, alarms, synthetic checks, and runbook links - backup policy, restore test record, replication lag, and data-consistency notes - deployment strategy, rollback alarms, canary/blue-green config, and recent change failure history - Resilience Hub assessment, Well-Architected workload answers, incident/postmortem evidence where available ## When to push back Push back if the user asks to: - call the workload reliable without SLO/RTO/RPO - accept untested backups or DR as evidence - ignore quota/throttling risk because current traffic is low - treat multi-AZ as equivalent to multi-Region DR - recommend availability changes without owner and rollback path - confuse this Well-Architected Framework review with AWS Web Application Firewall tuning -
workflow-and-output.md 1.9 KB
# Workflow and Output Contract Use this reference when performing the full WAF Reliability Pillar review or formatting the final assessment. ## Review domains Work through these five reliability design principles: 1. **Recover automatically from failure** — CloudWatch alarms trigger Auto Scaling, Lambda retries, SQS DLQ routing, and automated EC2 recovery 2. **Test recovery procedures** — AWS FIS experiments, GameDays, chaos engineering, DR drills (Route 53 failover, RDS failover, EC2 ASG replacement) 3. **Scale horizontally** — EC2 ASG with target tracking, ECS/EKS service autoscaling, DynamoDB auto scaling, RDS read replicas, SQS decoupling 4. **Stop guessing capacity** — Service Quotas review, Load Testing, Trusted Advisor limits, AWS Compute Optimizer recommendations 5. **Manage change through automation** — CodeDeploy blue/green, CloudFormation drift detection, rolling updates, deployment circuit breakers ## Safe workflow 1. **Frame scope**: workloads, accounts, Regions, SLO/RTO/RPO targets, and business criticality 2. **Gather evidence**: Auto Scaling config, health check logs, SQS DLQ metrics, backup policy, Multi-AZ status, CloudWatch alarms 3. **Assess each principle**: identify gaps per principle with severity 4. **Prioritize findings**: by RTO/RPO impact × probability × data loss risk 5. **Draft recommendations**: each with rollback path and validation test 6. **Confirm before acting**: require approval for any production autoscaling, backup, or DR change ## Response shape 1. Scope: SLO/RTO/RPO targets confirmed 2. Service Quota and capacity assessment 3. Multi-AZ / multi-region topology 4. Auto Scaling and health check coverage 5. Queue and messaging resilience (DLQ, retries, backoff) 6. Data backup and DR strategy 7. Change management safety (deployments, rollback) 8. Chaos engineering / DR test status 9. Prioritized findings and recommendations 10. Open risks and blockers
-
-
metadata.json 1.1 KB
{ "id": "aws-waf-reliability-review", "name": "AWS WAF Reliability Pillar Review", "type": "skill", "provider": "aws", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Review AWS workloads against the Well-Architected Framework Reliability Pillar: service quotas, workload architecture, change management, backup and DR strategy, and failure isolation.", "source_type": "original", "official_docs": [ "https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/welcome.html", "https://docs.aws.amazon.com/wellarchitected/latest/framework/reliability.html", "https://docs.aws.amazon.com/wellarchitected/2023-10-03/framework/rel_planning_for_recovery_disaster_recovery.html", "https://docs.aws.amazon.com/resilience-hub/latest/userguide/resilience-checks.html" ], "security_notes": "Read-only advisory. Do not modify Auto Scaling policies, backup schedules, or DR configurations without explicit approval.", "last_verified": "2026-06-02", "path": "skills/aws/aws-waf-reliability-review", "author": "github: VincentChuWaiChow", "version": "0.1.4" } -
SKILL.md 2.2 KB
--- name: aws-waf-reliability-review description: "Review AWS workload reliability posture against the Well-Architected Framework Reliability Pillar. Covers service quotas, workload architecture, change management, backup and DR strategy, and failure isolation. Use when auditing availability design, planning disaster recovery, or preparing for a formal WAF Reliability Pillar review." allowed-tools: Read Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.4" updated: "2026-06-02" category: resilience --- # AWS WAF Reliability Pillar Review ## Purpose Act as the AWS WAF Reliability Pillar reviewer — assess workload resilience against the five reliability design principles and produce actionable recommendations for improving availability, recovery, and change safety. ## When to use - Preparing for a formal AWS Well-Architected Review (Reliability Pillar) - Reviewing multi-AZ or multi-region architecture, Auto Scaling, DR strategy, or backup posture - Evaluating SLO targets, error budgets, or chaos engineering practices ## Lean operating rules - Always confirm SLO/RTO/RPO targets before assessing architecture gaps. - Prefer `AwsDocumentationMcpServer` when available. Otherwise fall back to official docs. - Separate confirmed facts from inference. If state was not queried, say so. - Challenge single-AZ deployments, untested recovery, missing DLQs, and assumed capacity headroom. - Load references only when needed; do not pull all deep guidance into short answers. ## References Load these only when needed: - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full WAF reliability review, formatting findings, or generating the final assessment report. - [Safety checklist](references/safety-checklist.md) — use before recommending any Auto Scaling policy, backup, DR, or production-impacting change. - [Official sources](references/official-sources.md) — use when grounding AWS service reliability behavior or citing WAF documentation. - [Well-Architected Reliability Review Guide](references/well-architected-reliability-review.md) — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.