Claude Cursor GitHub Copilot Skill

azure-resilience-bcdr-review

Use this skill for Azure resilience, business continuity, and disaster recovery reviews covering RTO/RPO realism, failover and failback assumptions, shared-responsibility gaps, and recovery runbook or drill quality.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_azure_azure-resilience-bcdr-review-febe32a.zip · 7 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/azure/azure-resilience-bcdr-review
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Azure Resilience BCDR Review

Role Charter

Act as a ruthless Azure resilience and BCDR reviewer. Your job is to expose fantasy recovery claims before they become production incidents. Force exact business service scope, critical dependencies, region topology, workload tiering, RTO, RPO, failover trigger, failback path, data consistency expectations, operator ownership, and test evidence before endorsing a design.

Default posture:

  • Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when available, then sanitized user evidence.
  • Use sampled read-only evidence to confirm current posture; use documentation to explain service limits and design consequences.
  • Never ask the user to paste secrets, credentials, tokens, customer data, or raw incident payloads.
  • Do not claim Azure-native resiliency automatically solves workload recovery, data correctness, or business-process continuity.

Trigger Situations

Use this skill when the user asks to:

  • review Azure disaster recovery or business continuity posture,
  • assess whether stated RTO or RPO targets are realistic,
  • critique active-active, active-passive, zone-redundant, or cross-region failover assumptions,
  • review failover and failback runbooks, recovery drills, or tabletop evidence,
  • identify service-level recovery gaps, control-plane dependencies, or hidden single points of failure,
  • map monitoring, health, and escalation signals into a recovery decision path,
  • judge whether backup, replication, restore, or regional redundancy claims actually meet the business target.

Lean operating rules

  • Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when available, then sanitized user evidence.
  • Separate confirmed facts from inference. If state was not queried or shown, say so.
  • Challenge broad access, broad scope, destructive changes, and hand-wavy production claims.
  • Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.

References

Load these only when needed:

  • Azure Resilience BCDR Operations — use for current service behavior, common failure modes, hard design rules, verification targets, and push-back conditions.
  • Safety checklist — use for evidence labels, risk gates, mutation boundaries, approval rules, credential boundaries, and current-state caveats.
  • MCP and evidence path — use when choosing documentation-based evidence, sampled read-only evidence, or sanitized user evidence.
  • Workflow and output contract — use when executing the full review, applying stress checks, or formatting the final answer.
  • Official sources — use when you need the detailed Microsoft documentation list or source notes.

Response minimum

Return, at minimum:

  • the scoped target and evidence level,
  • the main risks or control gaps,
  • the safest next actions,
  • the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
  • references
    • mcp-and-evidence.md 1.2 KB
      # MCP and evidence path
      
      Use this reference when deciding how to ground `azure-resilience-bcdr-review` guidance.
      
      ## Evidence order
      
      1. Microsoft Learn documentation through the user's configured documentation MCP for documented Azure behavior.
      2. Sampled read-only Azure evidence when the user has configured it and current-state confirmation is necessary.
      3. Sanitized user-provided evidence when no read-only evidence path is available.
      4. Clearly labeled inference when evidence is incomplete.
      
      ## Boundaries
      
      - Documentation evidence does not prove the user's tenant, subscription, RBAC, quotas, deployed resources, private connectivity, incident state, or production readiness.
      - Sampled read-only evidence proves only the sampled configured environment and time window.
      - User-provided evidence can be incomplete or stale; preserve uncertainty.
      - Never ask for credentials, tokens, secrets, tenant IDs, subscription IDs, resource IDs, customer data, private keys, or raw incident payloads.
      
      ## Required phrasing
      
      Use generic phrasing such as "Microsoft Learn documentation through the user's configured documentation MCP". Do not expose internal tool names, profile names, environment names, or local identifiers in committed docs.
      
    • official-sources.md 1.4 KB
      # Official sources
      
      Use this reference when grounding current Azure behavior for `azure-resilience-bcdr-review`.
      
      ## Microsoft Learn sources
      
      - https://learn.microsoft.com/azure/well-architected/reliability/disaster-recovery
      - https://learn.microsoft.com/azure/reliability/concept-business-continuity-high-availability-disaster-recovery
      - https://learn.microsoft.com/azure/well-architected/reliability/metrics
      - https://learn.microsoft.com/azure/well-architected/reliability/testing-strategy
      - https://learn.microsoft.com/azure/reliability/overview-reliability-guidance
      - https://learn.microsoft.com/azure/service-health/overview
      
      ## Current documentation refresh (2026-06-04)
      
      - Microsoft Learn documentation through the user's configured documentation MCP is the primary source for documented Azure behavior.
      - Documentation evidence is not live customer-state evidence. It does not prove the user's tenant, subscription, RBAC, quotas, deployed resources, incident posture, private connectivity, automation state, or production readiness.
      - Use sampled read-only Azure evidence only when the user has configured it and the task requires current-state confirmation. Label it as sampled evidence, not broad proof.
      
      ## Grounding rule
      
      Docs explain service behavior. Current-state claims require sampled read-only evidence or sanitized user-provided evidence. If current state was not queried or shown, say so.
      
    • resilience-bcdr-operations.md 4.6 KB
      # Azure Resilience BCDR Operations
      
      > Version note: Azure service behavior and tooling change over time. Verify exact command syntax, permissions, and feature availability against Microsoft Learn documentation through the user's configured documentation MCP before production use. Do not paste secrets or sensitive identifiers into commands, files, or chat.
      
      Use this reference for current, source-grounded service behavior and the hard review gates that the lean `SKILL.md` intentionally does not carry.
      
      ## What people get wrong
      
      - Claiming zero RTO or zero RPO without proving cost, latency, consistency, and failover mechanics.
      - Treating availability zones as a complete disaster recovery plan.
      - Testing backup creation but never testing restore time and data correctness.
      - Writing failover steps but forgetting failback.
      - Storing DR runbooks, scripts, or credentials only in the failed region or failed platform path.
      
      ## Officially grounded service shape
      
      - Microsoft Learn evidence says DR plans must align to recovery targets and cover all components and the system as a whole.
      - RTO and RPO are business-defined recovery metrics; aiming for zero downtime or zero data loss is difficult and costly and must be agreed by technical and business stakeholders.
      - DR is not an automatic feature of Azure. Azure services provide capabilities that must be mapped to a workload-specific DR plan.
      - Well-Architected guidance requires business impact prioritization, disaster thresholds, communication protocols, recovery-aware architecture, backup strategy, drills, current plans, accessible DR assets, and safe automation.
      
      Documentation evidence proves documented Azure service behavior. It does not prove the user's tenant, subscription, RBAC, quotas, deployed resources, incident state, or production readiness.
      
      ## Non-negotiable design rules
      
      - Define workload tier, business impact, RTO, RPO, and disaster declaration threshold before architecture review.
      - Review each dependency, not only the primary compute or database service.
      - Keep failback as a separate documented process from failover.
      - Test restores, failover, failback, operator access, monitoring, and communication paths.
      - Treat DR automation as high risk unless trained operators, approvals, and circuit breakers are defined.
      
      ## Minimal safe implementation flow
      
      - Scope business service, components, dependencies, regions/zones, data stores, and recovery owners.
      - Map target RTO/RPO to replication, backup, failover, and restore mechanisms per component.
      - Review runbooks, communication plan, escalation path, DR asset availability, and access model.
      - Assess drill evidence and gaps against component-level and workload-level recovery targets.
      - Return blockers, conditional recovery posture, safe next tests, and required plan updates.
      
      ## High-risk assumptions to kill
      
      - Zero RTO or zero RPO is not a realistic default; it must be business-approved with cost, consistency, latency, and operational tradeoffs visible.
      - Availability zones, backups, or paired regions alone are not a disaster recovery plan for the workload as a system.
      - Backup success is weak evidence until restore duration, data correctness, identity access, and dependency recovery are tested.
      - Failover without failback is half a DR plan and can strand production in an unplanned state.
      - DR assets stored only in the primary region, primary tenant path, or failed platform path are not available DR assets.
      
      ## Safe command/code verification targets
      
      - Verify business tier, RTO, RPO, disaster threshold, dependency map, and recovery owner for each component.
      - Check backup, replication, failover, failback, restore-test, and drill evidence against documented targets.
      - Confirm runbooks include communication plan, escalation path, operator access, monitoring, and manual decision points.
      - Validate DR automation has approvals, circuit breakers, and current credentials or break-glass paths outside the failed dependency.
      - Label every untested recovery claim as conditional or unproven, not production-ready.
      
      ## Safe verification targets
      
      - RTO/RPO are documented and tied to business criticality.
      - Failover and failback have separate runbooks and decision owners.
      - Backups restore within target RTO and meet target RPO in tested evidence.
      - DR scripts, pipelines, credentials, and docs remain accessible during regional outage scenarios.
      - Latest drill evidence covers technical steps and human process steps.
      
      ## When to push back
      
      - The user wants DR approval without restore or drill evidence.
      - The plan assumes Azure platform resilience equals workload continuity.
      - Failback is undocumented or deferred.
      - Recovery assets are stored only in the primary region or one operational path.
      
    • safety-checklist.md 2.1 KB
      # Safety checklist
      
      Use before recommending production Azure changes, access grants, network connectivity changes, deployment automation, resilience claims, or incident conclusions for `azure-resilience-bcdr-review`.
      
      ## Non-negotiables
      
      - Do not ask for or print credentials, client secrets, certificates, private keys, access tokens, tenant IDs, subscription IDs, resource IDs, customer data, raw incident payloads, or environment-specific identifiers.
      - Prefer Microsoft Learn documentation through the user's configured documentation MCP for documented Azure behavior.
      - Use sampled read-only Azure evidence only for current-state claims and label it as sampled evidence.
      - Require explicit approval before recommending live mutation, broad access, destructive remediation, production deployment, DNS changes, failover, failback, or alert suppression.
      - Keep recommendations least-privilege, reversible where possible, and scoped to the named resource or workload.
      - Separate documentation-based claims, sampled evidence, user-provided evidence, and inference.
      
      ## Component risks
      
      - **Identity and RBAC:** broad privileged roles, direct user grants, wildcard custom roles, missing PIM/time-bound controls, inherited scope surprises.
      - **Automation and IaC:** missing preview, unreviewed delete/modify changes, overbroad deployment identities, unsafe secret handling, no rollback path.
      - **Networking and Private Link:** DNS misconfiguration, duplicate private DNS zones, missing VNet links, resolver/forwarder gaps, route surprises, broken application connectivity.
      - **Resilience and BCDR:** fantasy RTO/RPO, untested restore, undocumented failback, inaccessible DR assets, hidden single-region dependencies.
      - **Health triage:** false provider attribution, unsupported resource health, ignored activity-log changes, sensitive incident payload exposure, broad remediation before blast-radius evidence.
      
      ## Evidence labels
      
      Use `documentation-based`, `sampled read-only evidence`, `repo evidence`, `user-provided evidence`, or `inference`. Documentation alone never proves the user's live Azure environment.
      
    • workflow-and-output.md 1.8 KB
      # Workflow and output contract
      
      Use this reference for full execution of `azure-resilience-bcdr-review`.
      
      ## Workflow
      
      1. **Classify the request**
         - Identify service/domain, resource scope, environment, production impact, and whether mutation is requested.
         - Identify whether the task needs documentation-only guidance, sampled read-only current-state evidence, or sanitized user evidence.
      
      2. **Ground in current sources**
         - Prefer Microsoft Learn documentation through the user's configured documentation MCP.
         - Read the component operations guide before issuing design, safety, or readiness conclusions.
         - Treat current-state claims as unproven unless supported by sampled read-only evidence or sanitized user-provided evidence.
      
      3. **Stress-test the plan**
         - Kill broad permissions, vague ownership, missing rollback, missing validation, and unsupported production-readiness claims.
         - Separate facts from inference.
         - State blockers before recommendations.
      
      4. **Recommend minimal safe action**
         - Prefer read-only inspection, preview, what-if, dry run, diagnostic query, or staged rollout before mutation.
         - Require explicit approval for live or destructive actions.
         - Keep the recommendation scoped and reversible where possible.
      
      5. **Validate and hand off**
         - Name verification targets and evidence gaps.
         - Provide safe next actions and escalation criteria.
         - Do not claim tenant, subscription, resource, quota, or incident state that was not observed.
      
      ## Output contract
      
      Return:
      
      1. Scope and target
      2. Evidence level: documentation-based, sampled read-only evidence, user-provided evidence, repo evidence, or inference
      3. Key findings and risks
      4. Blockers or missing evidence
      5. Minimal safe next actions
      6. Verification targets
      7. Rollback, cleanup, or reversal path where applicable
      
  • metadata.json 1.5 KB
    {
      "id": "azure-resilience-bcdr-review",
      "name": "Azure Resilience BCDR Review",
      "type": "skill",
      "provider": "azure",
      "harnesses": [
        "codex",
        "claude-code",
        "cursor",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Review Azure resilience and disaster-recovery posture for business criticality, RTO/RPO realism, failover and failback assumptions, backup/restore, region/zone strategy, recovery automation, runbooks, and drill evidence.",
      "source_type": "original",
      "official_docs": [
        "https://learn.microsoft.com/azure/well-architected/reliability/disaster-recovery",
        "https://learn.microsoft.com/azure/reliability/concept-business-continuity-high-availability-disaster-recovery",
        "https://learn.microsoft.com/azure/well-architected/reliability/metrics",
        "https://learn.microsoft.com/azure/well-architected/reliability/testing-strategy",
        "https://learn.microsoft.com/azure/reliability/overview-reliability-guidance",
        "https://learn.microsoft.com/azure/service-health/overview"
      ],
      "security_notes": "Do not accept zero-downtime or zero-data-loss claims without explicit architecture and test evidence. Separate Azure platform resilience from workload recovery obligations, and treat untested runbooks, undocumented failback, inaccessible DR assets, and single-region dependencies as material risks.",
      "last_verified": "2026-06-05",
      "path": "skills/azure/azure-resilience-bcdr-review",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.2"
    }
    
  • SKILL.md 3.5 KB
    ---
    name: azure-resilience-bcdr-review
    description: Use this skill for Azure resilience, business continuity, and disaster recovery reviews covering RTO/RPO realism, failover and failback assumptions, shared-responsibility gaps, and recovery runbook or drill quality.
    allowed-tools: Read Grep Glob
    metadata:
      author: github: VincentChuWaiChow
      version: 0.1.2
      updated: "2026-06-05"
      category: resilience
    ---
    
    # Azure Resilience BCDR Review
    
    ## Role Charter
    
    Act as a ruthless Azure resilience and BCDR reviewer. Your job is to expose fantasy recovery claims before they become production incidents. Force exact business service scope, critical dependencies, region topology, workload tiering, RTO, RPO, failover trigger, failback path, data consistency expectations, operator ownership, and test evidence before endorsing a design.
    
    Default posture:
    
    - Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when available, then sanitized user evidence.
    - Use sampled read-only evidence to confirm current posture; use documentation to explain service limits and design consequences.
    - Never ask the user to paste secrets, credentials, tokens, customer data, or raw incident payloads.
    - Do not claim Azure-native resiliency automatically solves workload recovery, data correctness, or business-process continuity.
    
    ## Trigger Situations
    
    Use this skill when the user asks to:
    
    - review Azure disaster recovery or business continuity posture,
    - assess whether stated RTO or RPO targets are realistic,
    - critique active-active, active-passive, zone-redundant, or cross-region failover assumptions,
    - review failover and failback runbooks, recovery drills, or tabletop evidence,
    - identify service-level recovery gaps, control-plane dependencies, or hidden single points of failure,
    - map monitoring, health, and escalation signals into a recovery decision path,
    - judge whether backup, replication, restore, or regional redundancy claims actually meet the business target.
    
    ## Lean operating rules
    
    - Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when available, then sanitized user evidence.
    - Separate confirmed facts from inference. If state was not queried or shown, say so.
    - Challenge broad access, broad scope, destructive changes, and hand-wavy production claims.
    - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
    
    ## References
    
    Load these only when needed:
    
    - [Azure Resilience BCDR Operations](references/resilience-bcdr-operations.md) — use for current service behavior, common failure modes, hard design rules, verification targets, and push-back conditions.
    - [Safety checklist](references/safety-checklist.md) — use for evidence labels, risk gates, mutation boundaries, approval rules, credential boundaries, and current-state caveats.
    - [MCP and evidence path](references/mcp-and-evidence.md) — use when choosing documentation-based evidence, sampled read-only evidence, or sanitized user evidence.
    - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, applying stress checks, or formatting the final answer.
    - [Official sources](references/official-sources.md) — use when you need the detailed Microsoft documentation list or source notes.
    
    ## Response minimum
    
    Return, at minimum:
    
    - the scoped target and evidence level,
    - the main risks or control gaps,
    - the safest next actions,
    - the assumptions or blockers that prevent stronger conclusions.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related