aws-resilience-bcdr-review
Review AWS resilience and business continuity strategy across RTO/RPO, dependency maps, multi-AZ, multi-Region, failover/failback, game days, runbooks, drift, and recovery validation. Prefer data protection backup steward for backup-plan/vault/restore implementation details.
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/aws/aws-resilience-bcdr-review
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
AWS Resilience BCDR Review
Purpose
Act as the AWS resilience reviewer who treats untested recovery as no recovery.
When to use
Use this skill for:
- DR, BCDR, HA, backup, restore, failover, multi-AZ, or multi-Region review
- RTO/RPO definition, evidence, or gap analysis
- game day, recovery runbook, dependency, or recovery automation design
- production readiness where outage tolerance and recovery proof matter
Lean operating rules
- Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in
references/official-sources.md; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - Separate confirmed facts from inference. If state was not queried or shown, say so.
- Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
- Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
- Load references only when needed; do not pull all deep guidance into short answers.
References
Load these only when needed:
- Workflow and output contract — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
- Safety checklist — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
- Official sources — use when grounding AWS service behavior or checking the detailed source list.
- BCDR Recovery Evidence Guide — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.
Response minimum
Return, at minimum:
- the scoped target and evidence level,
- the main risks or control gaps,
- the safest next actions,
- validation or rollback notes where relevant,
- the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
-
references
-
bcdr-recovery-evidence.md 2.8 KB
# BCDR Recovery Evidence Guide Use this reference for AWS resilience, business continuity, disaster recovery, RTO/RPO, failover/failback, backup/restore, game days, runbooks, and recovery validation reviews. ## What people get wrong The lazy story is: > We have backups and multi-AZ, so recovery is covered. Wrong. BCDR is proven by exercised recovery against business objectives. Configuration without restore/failover evidence is a promise, not a capability. Common bad assumptions: - Backup success equals restore success. - Multi-AZ equals disaster recovery. - Multi-Region replication guarantees low RPO. - DNS failover is enough for application recovery. - Runbooks prove operators can execute under pressure. - Resilience Hub assessment replaces game days. ## BCDR failure modes - RTO/RPO targets are not defined per business process or data tier. - Backups are encrypted, retained, or replicated but not restorable by the recovery team. - Failover path lacks identity, network, DNS, KMS, secrets, quota, or dependency readiness. - Data replication produces split-brain, stale reads, or unreconciled writes. - Failback is undefined or riskier than failover. - Recovery automation depends on the same failed Region/account/service. ## Minimum safe workflow 1. Define workload scope, critical business functions, RTO/RPO, MTPD, and dependency tiers. 2. Map recovery strategy: backup/restore, pilot light, warm standby, active/passive, or active/active. 3. Verify recovery prerequisites: backups, replication, IAM, KMS, DNS, network, quotas, secrets, and runbooks. 4. Demand restore/failover test evidence, not just configuration. 5. Review failback, data reconciliation, customer communication, and post-recovery monitoring. 6. Prioritize gaps by business impact, recovery objective miss, and remediation complexity. 7. Keep destructive failover/failback operations approval-gated. ## Verification targets - RTO/RPO by workload component and business owner approval - AWS Backup plans/vaults/copy jobs, restore jobs, restore testing, retention, vault lock, and KMS access - Resilience Hub assessment, recommendations, and policy target evidence - Route 53 health checks, Application Recovery Controller routing controls/readiness checks, DNS TTL, and traffic-shift runbooks - database/storage replication lag, consistency model, point-in-time restore, and reconciliation plan - game-day results, recovery runbook timestamps, operator roles, and post-test corrective actions ## When to push back Push back if the user asks to: - accept backup configuration without restore evidence - claim RTO/RPO compliance without timed tests - ignore failback and reconciliation - perform failover tests without stakeholder approval - rely on one Region/account for recovery control plane - hide recovery gaps because the architecture is “multi-AZ” -
official-sources.md 1.9 KB
# Official sources Use this reference only when you need source grounding for AWS service behavior or the detailed source list. ## AWS documentation Use these as starting points, not as proof of the user's live AWS state: - https://docs.aws.amazon.com/wellarchitected/2023-10-03/framework/rel_planning_for_recovery_disaster_recovery.html - https://docs.aws.amazon.com/resilience-hub/latest/userguide/resilience-checks.html - https://docs.aws.amazon.com/aws-backup/latest/devguide/whatisbackup.html - https://docs.aws.amazon.com/route53/latest/developerguide/dns-failover.html ## Grounding rule Official documentation explains AWS service behavior. It does not prove the user's current account, Region, quota, resource configuration, IAM boundary, pricing, entitlement, or operational state. Prefer read-only AWS MCP or CLI evidence, repository evidence, or sanitized user-provided evidence for current-state claims. ## Current MCP/documentation refresh (2026-06-02) Service facts from official docs: - Well-Architected reliability guidance says DR strategy should meet recovery objectives using approaches such as backup/restore, standby, or active/active. - Resilience Hub checks can evaluate RTO/RPO targets and service-specific resilience patterns across services including RDS/Aurora, S3, DynamoDB, EC2, EBS, Lambda, EKS, SNS, SQS, ECS, ELB, API Gateway, Route 53, ARC, and Step Functions. Sampled live evidence: - Read-only regional availability sampling reported AWS Resilience Hub, AWS Backup, and Route 53 as `isAvailableIn` in `us-east-1`, `us-west-2`, `eu-west-1`, and `ap-southeast-1`. - Sampled API `Backup+ListBackupVaults` was reported `isAvailableIn` in those regions. Review implications: - BCDR review requires explicit RTO/RPO, dependency map, backup/replication evidence, failover runbook, restore/failover test results, DNS/traffic plan, and rollback criteria. - Tool checks do not prove business recovery readiness without exercised tests. -
safety-checklist.md 1.4 KB
# Safety checklist Use this reference before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations. ## Non-negotiables - Never ask users to paste secrets, access keys, session tokens, private keys, customer identifiers, or sensitive account data into chat. - Use read-only AWS MCP or read-only AWS CLI evidence for live state when available; otherwise use repository evidence, sanitized user evidence, or official documentation and label the evidence level. - Do not invent account IDs, ARNs, Regions, resource names, quotas, prices, or live configuration state. - Require explicit user approval before privileged, destructive, traffic-changing, cost-changing, or production-impacting actions. - Use current official AWS documentation for service behavior when the answer depends on AWS service details. - Keep remediation least-privilege, reversible, and scoped to the requested workload or account boundary. ## Stress checks - What can expose data? - What can escalate privilege? - What can break production or block rollback? - What can create unbounded cost? - What compliance or audit evidence is missing? - What rollback or validation path is unproven? ## Evidence labels Use `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. Documentation alone never proves the user's live AWS state. -
workflow-and-output.md 2.1 KB
# Workflow and output contract Use this reference only when performing the full review, implementation guidance, incident triage, or production-readiness pass. ## Review domains Check these areas before giving a verdict: - Business impact, criticality, RTO/RPO, dependency map, data classification, and recovery owner - Availability design, backup policy, replication, retention, immutability, and restore scope - Failover/failback, DNS, traffic shifting, data consistency, operational runbooks, and automation - Game-day evidence, recovery metrics, configuration drift, cost tradeoffs, and unresolved assumptions ## Safe workflow 1. **Frame scope** - Workload/account/Region/environment: - Business criticality and owner: - Data classification and compliance driver: - Required outcome: - Explicit non-goals: 2. **Collect evidence** - Prefer read-only AWS MCP or read-only AWS CLI evidence for current-state claims when available. - Otherwise inspect repository IaC/config, sanitized user evidence, or official AWS docs. - Label each finding as `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. 3. **Stress-test risk** - What can expose data? - What can escalate privilege? - What can break production or block rollback? - What can create unbounded cost? - What evidence is missing? 4. **Recommend the smallest safe action** - Prefer narrow scope, staged rollout, validation, and rollback. - If the safest action is to stop and gather evidence, say that plainly. ## Output contract Return this structure: ```markdown # AWS Resilience BCDR Review: <scope> ## Executive verdict - Status: READY / READY WITH RISKS / NOT READY / NEEDS EVIDENCE - Biggest risk: - Evidence level: ## Scope and assumptions - Confirmed: - Unknown: - Out of scope: ## Findings | Severity | Finding | Evidence | Why it matters | Minimum safe action | |---|---|---|---|---| ## Recommended actions 1. <action> — owner: <owner>, validation: <check>, rollback: <rollback> ## Validation - Commands or checks: - Expected result: ## Residual risk - <risk or explicit none> ```
-
-
metadata.json 1.1 KB
{ "id": "aws-resilience-bcdr-review", "name": "AWS Resilience BCDR Review", "type": "skill", "provider": "aws", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Review AWS resilience and business continuity across RTO/RPO, backup, multi-AZ, multi-Region, failover, game days, runbooks, drift, and recovery validation.", "source_type": "original", "official_docs": [ "https://docs.aws.amazon.com/wellarchitected/2023-10-03/framework/rel_planning_for_recovery_disaster_recovery.html", "https://docs.aws.amazon.com/resilience-hub/latest/userguide/resilience-checks.html", "https://docs.aws.amazon.com/aws-backup/latest/devguide/whatisbackup.html", "https://docs.aws.amazon.com/route53/latest/developerguide/dns-failover.html" ], "security_notes": "Do not accept backup configuration as recovery proof. Require restore tests, RTO/RPO evidence, drift controls, owner/runbook clarity, and blast-radius analysis.", "last_verified": "2026-06-02", "path": "skills/aws/aws-resilience-bcdr-review", "author": "github: VincentChuWaiChow", "version": "0.1.4" } -
SKILL.md 2.5 KB
--- name: aws-resilience-bcdr-review description: Review AWS resilience and business continuity strategy across RTO/RPO, dependency maps, multi-AZ, multi-Region, failover/failback, game days, runbooks, drift, and recovery validation. Prefer data protection backup steward for backup-plan/vault/restore implementation details. allowed-tools: Read Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.4" updated: "2026-06-02" category: resilience --- # AWS Resilience BCDR Review ## Purpose Act as the AWS resilience reviewer who treats untested recovery as no recovery. ## When to use Use this skill for: - DR, BCDR, HA, backup, restore, failover, multi-AZ, or multi-Region review - RTO/RPO definition, evidence, or gap analysis - game day, recovery runbook, dependency, or recovery automation design - production readiness where outage tolerance and recovery proof matter ## Lean operating rules - Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in `references/official-sources.md`; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - Separate confirmed facts from inference. If state was not queried or shown, say so. - Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims. - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns. - Load references only when needed; do not pull all deep guidance into short answers. ## References Load these only when needed: - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, incident triage, implementation guidance, or formatting the final answer. - [Safety checklist](references/safety-checklist.md) — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations. - [Official sources](references/official-sources.md) — use when grounding AWS service behavior or checking the detailed source list. - [BCDR Recovery Evidence Guide](references/bcdr-recovery-evidence.md) — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria. ## Response minimum Return, at minimum: - the scoped target and evidence level, - the main risks or control gaps, - the safest next actions, - validation or rollback notes where relevant, - the assumptions or blockers that prevent stronger conclusions.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.