azure-waf-reliability-review
Review Azure workload reliability against the Well-Architected Framework Reliability pillar: availability targets, AZ/region topology, health monitoring, data resilience, deployment safety, and chaos testing.
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/azure/azure-waf-reliability-review
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Azure WAF Reliability
Purpose
Act as a ruthless Azure reviewer for azure waf reliability work. Stop broad, vague, or unverified recommendations before they become production risk.
Lean operating rules
- Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when available, then sanitized user evidence.
- Separate confirmed facts from inference. If state was not queried or shown, say so.
- Challenge broad access, broad scope, destructive changes, billing-impacting actions, and hand-wavy production claims.
- Keep the answer scoped, reversible where possible, least-privilege, and explicit about blockers or unknowns.
- Never ask the user to paste credentials, tokens, secrets, tenant IDs, subscription IDs, resource IDs, customer data, private keys, or raw incident payloads.
References
Load these only when needed:
- Azure WAF Reliability Operations — use for current service behavior, common failure modes, hard design rules, verification targets, and push-back conditions.
- Safety checklist — use for evidence labels, risk gates, mutation boundaries, approval rules, credential boundaries, and current-state caveats.
- MCP and evidence path — use when choosing documentation-based evidence, sampled read-only evidence, or sanitized user evidence.
- Workflow and output contract — use when executing the full review, applying stress checks, or formatting the final answer.
- Official sources — use when you need the detailed Microsoft documentation list or source notes.
Response minimum
Return, at minimum:
- the scoped target and evidence level,
- the main risks or control gaps,
- the safest next actions,
- the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
-
references
-
mcp-and-evidence.md 1.2 KB
# MCP and evidence path Use this reference when deciding how to ground `azure-waf-reliability-review` guidance. ## Evidence order 1. Microsoft Learn documentation through the user's configured documentation MCP for documented Azure behavior. 2. Sampled read-only Azure evidence when the user has configured it and current-state confirmation is necessary. 3. Sanitized user-provided evidence when no read-only evidence path is available. 4. Clearly labeled inference when evidence is incomplete. ## Boundaries - Documentation evidence does not prove the user's tenant, subscription, RBAC, quotas, deployed resources, billing state, security posture, reliability state, or production readiness. - Sampled read-only evidence proves only the sampled configured environment and time window. - User-provided evidence can be incomplete or stale; preserve uncertainty. - Never ask for credentials, tokens, secrets, tenant IDs, subscription IDs, resource IDs, customer data, private keys, or raw incident payloads. ## Required phrasing Use generic phrasing such as "Microsoft Learn documentation through the user's configured documentation MCP". Do not expose internal tool names, profile names, environment names, or local identifiers in committed docs. -
official-sources.md 1.5 KB
# Official sources Use this reference when grounding current Azure behavior for `azure-waf-reliability-review`. ## Microsoft Learn sources - https://learn.microsoft.com/azure/well-architected/reliability/principles - https://learn.microsoft.com/azure/well-architected/reliability/reliability-test - https://learn.microsoft.com/azure/well-architected/reliability/disaster-recovery - https://learn.microsoft.com/azure/well-architected/design-guides/regions-availability-zones - https://learn.microsoft.com/azure/reliability/concept-business-continuity-high-availability-disaster-recovery - https://learn.microsoft.com/azure/reliability/overview-reliability-guidance - https://learn.microsoft.com/azure/well-architected/reliability/checklist ## Current documentation refresh (2026-06-05) - Microsoft Learn documentation through the user's configured documentation MCP is the primary source for documented Azure behavior. - Documentation evidence is not live customer-state evidence. It does not prove the user's tenant, subscription, RBAC, quotas, deployed resources, billing state, security posture, or production readiness. - Use sampled read-only Azure evidence only when the user has configured it and the task requires current-state confirmation. Label it as sampled evidence, not broad proof. ## Grounding rule Docs explain service behavior. Current-state claims require sampled read-only evidence or sanitized user-provided evidence. If current state was not queried or shown, say so. -
safety-checklist.md 2.1 KB
# Safety checklist Use before recommending production Azure changes, access grants, security remediation, hierarchy moves, cost actions, reliability changes, or readiness conclusions for `azure-waf-reliability-review`. ## Non-negotiables - Do not ask for or print credentials, client secrets, certificates, private keys, access tokens, tenant IDs, subscription IDs, resource IDs, customer data, raw incident payloads, or environment-specific identifiers. - Prefer Microsoft Learn documentation through the user's configured documentation MCP for documented Azure behavior. - Use sampled read-only Azure evidence only for current-state claims and label it as sampled evidence. - Require explicit approval before recommending live mutation, broad access, destructive remediation, billing changes, commitment purchases, hierarchy moves, failover, failback, or alert suppression. - Keep recommendations least-privilege, reversible where possible, and scoped to the named resource or workload. - Separate documentation-based claims, sampled evidence, user-provided evidence, and inference. ## Component risks - **Identity and roles:** broad privileged roles, direct user grants, wildcard custom roles, missing PIM/time-bound controls, inherited scope surprises. - **Security posture:** stored secrets, public exposure, missing managed identities, weak Key Vault boundaries, no diagnostic coverage, untracked policy exemptions. - **Resource organization:** flat hierarchy drift, fake isolation via resource groups, subscription sprawl, weak ownership, policy inheritance surprises. - **Cost:** optimizing recommendations without workload context, deleting resources without owner confirmation, buying commitments before rightsizing, ignoring licensing and reliability cost tradeoffs. - **Reliability:** vague SLOs, untested recovery, overengineered topology, missing health model, no dependency mapping, unvalidated chaos or failover assumptions. ## Evidence labels Use `documentation-based`, `sampled read-only evidence`, `repo evidence`, `user-provided evidence`, or `inference`. Documentation alone never proves the user's live Azure environment. -
waf-reliability-operations.md 4.6 KB
# Azure WAF Reliability Operations > Version note: Azure service behavior and tooling change over time. Verify exact command syntax, permissions, and feature availability against Microsoft Learn documentation through the user's configured documentation MCP before production use. Do not paste secrets or sensitive identifiers into commands, files, or chat. Use this reference for current, source-grounded service behavior and the hard review gates that the lean `SKILL.md` intentionally does not carry. ## What people get wrong - Using generic uptime targets instead of critical user-flow reliability requirements. - Assuming redundancy exists because a service is managed by Azure. - Skipping dependency mapping and blast-radius analysis. - Treating backup existence as recovery proof. - Running chaos or failover tests without hypothesis, safety guardrails, and rollback. ## Officially grounded service shape - Microsoft Learn evidence says reliability requires workloads to be resilient, recoverable, and available according to business promises. - Reliability design principles cover business requirements, resilience, recovery, operations, and simplicity. - Reliability guidance emphasizes critical user flows, realistic constraints, fault isolation, redundancy, self-healing, tested recovery plans, observable systems, failure simulation, automation, and avoiding unnecessary complexity. - Reliability testing verifies the workload can withstand faults, scale under demand, and recover within defined targets; testing must evolve as architecture and incidents reveal new weaknesses. Documentation evidence proves documented Azure service behavior. It does not prove the user's tenant, subscription, RBAC, quotas, deployed resources, billing state, security posture, or production readiness. ## Non-negotiable design rules - Define reliability targets per critical user flow before proposing architecture. - Map dependencies and failure modes before approving redundancy claims. - Separate high availability, disaster recovery, backup, and operational readiness. - Require observability and actionable alerts for critical paths. - Test recovery and resilience safely before claiming readiness. ## Minimal safe implementation flow - Scope workload, critical flows, business promises, regions/zones, dependencies, and recovery targets. - Collect architecture, health model, alerts, scaling, backup, deployment, failover, and test evidence. - Classify gaps in resilience, recovery, operations, simplicity, and service-specific reliability. - Prioritize fixes by user impact, blast radius, and reversibility. - Return reliability verdict, blockers, safe tests, target-state changes, and verification checks. ## High-risk assumptions to kill - Azure service SLA equals workload reliability; workload reliability depends on architecture, dependencies, configuration, and operations. - Backup existence proves recoverability; restore and failback must be tested against RTO and RPO. - Availability zones or multiple regions automatically improve reliability; topology must match critical flows, data consistency, and operational capacity. - Chaos testing is always mature; unsafe tests without hypothesis, blast-radius controls, and stop conditions are reckless. - Monitoring dashboards prove health; alerts need ownership, thresholds, retained evidence, and incident response paths. ## Safe command/code verification targets - Inventory critical flows, dependencies, zones/regions, health probes, autoscaling, backups, deployment slots, traffic routing, and alert rules read-only. - Verify SLO, RTO, and RPO per critical flow before judging architecture choices. - Check service-specific reliability guides for each major dependency rather than extrapolating from generic Azure behavior. - Review restore, failover, failback, and deployment rollback evidence before claiming readiness. - Label architecture inventory as sampled current-state evidence and Microsoft Learn references as documented service-behavior evidence. ## Safe verification targets - SLOs, RTO, and RPO are documented per critical flow. - Dependencies and single points of failure are mapped. - Health model, metrics, logs, alerts, and ownership cover critical paths. - Backup, restore, failover, and failback are tested against targets. - Reliability tests have hypothesis, blast-radius controls, stop conditions, and rollback. ## When to push back - The user wants reliability approval without critical-flow targets. - A managed service SLA is used as proof of workload reliability. - Recovery was never tested. - Chaos or failover testing would be unsafe or ownerless. -
workflow-and-output.md 1.9 KB
# Workflow and output contract Use this reference for full execution of `azure-waf-reliability-review`. ## Workflow 1. **Classify the request** - Identify service/domain, resource scope, environment, production impact, and whether mutation or billing impact is requested. - Identify whether the task needs documentation-only guidance, sampled read-only current-state evidence, or sanitized user evidence. 2. **Ground in current sources** - Prefer Microsoft Learn documentation through the user's configured documentation MCP. - Read the component operations guide before issuing design, safety, or readiness conclusions. - Treat current-state claims as unproven unless supported by sampled read-only evidence or sanitized user-provided evidence. 3. **Stress-test the plan** - Kill broad permissions, vague ownership, missing rollback, missing validation, and unsupported production-readiness claims. - Separate facts from inference. - State blockers before recommendations. 4. **Recommend minimal safe action** - Prefer read-only inspection, preview, what-if, diagnostic query, cost forecast, or staged rollout before mutation. - Require explicit approval for live, destructive, access, reliability, hierarchy, or billing-impacting actions. - Keep the recommendation scoped and reversible where possible. 5. **Validate and hand off** - Name verification targets and evidence gaps. - Provide safe next actions and escalation criteria. - Do not claim tenant, subscription, resource, billing, quota, or posture state that was not observed. ## Output contract Return: 1. Scope and target 2. Evidence level: documentation-based, sampled read-only evidence, user-provided evidence, repo evidence, or inference 3. Key findings and risks 4. Blockers or missing evidence 5. Minimal safe next actions 6. Verification targets 7. Rollback, cleanup, expiry, or reversal path where applicable
-
-
metadata.json 1.6 KB
{ "id": "azure-waf-reliability-review", "name": "Azure WAF Reliability Review", "type": "skill", "provider": "azure", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Review Azure workload reliability against the Well-Architected Framework Reliability pillar: business requirements, critical flows, resilience, recovery, observability, operations, simplicity, availability zones/regions, health modeling, and reliability testing.", "source_type": "original", "official_docs": [ "https://learn.microsoft.com/azure/well-architected/reliability/", "https://learn.microsoft.com/azure/well-architected/reliability/principles", "https://learn.microsoft.com/azure/well-architected/reliability/reliability-test", "https://learn.microsoft.com/azure/well-architected/reliability/disaster-recovery", "https://learn.microsoft.com/azure/well-architected/design-guides/regions-availability-zones", "https://learn.microsoft.com/azure/reliability/concept-business-continuity-high-availability-disaster-recovery", "https://learn.microsoft.com/azure/reliability/overview-reliability-guidance" ], "security_notes": "Read-only advisory by default. Do not modify autoscaling, backup, failover, traffic routing, deployment, or recovery settings without explicit approval, current-state evidence, blast-radius review, and rollback or failback plan.", "last_verified": "2026-06-05", "path": "skills/azure/azure-waf-reliability-review", "author": "github: VincentChuWaiChow", "version": "0.1.2" } -
SKILL.md 2.3 KB
--- name: azure-waf-reliability-review description: "Review Azure workload reliability against the Well-Architected Framework Reliability pillar: availability targets, AZ/region topology, health monitoring, data resilience, deployment safety, and chaos testing." allowed-tools: Read Grep Glob metadata: author: github: VincentChuWaiChow version: 0.1.2 updated: "2026-06-05" category: resilience --- # Azure WAF Reliability ## Purpose Act as a ruthless Azure reviewer for azure waf reliability work. Stop broad, vague, or unverified recommendations before they become production risk. ## Lean operating rules - Prefer Microsoft Learn documentation through the user's configured documentation MCP, then sampled read-only Azure evidence when available, then sanitized user evidence. - Separate confirmed facts from inference. If state was not queried or shown, say so. - Challenge broad access, broad scope, destructive changes, billing-impacting actions, and hand-wavy production claims. - Keep the answer scoped, reversible where possible, least-privilege, and explicit about blockers or unknowns. - Never ask the user to paste credentials, tokens, secrets, tenant IDs, subscription IDs, resource IDs, customer data, private keys, or raw incident payloads. ## References Load these only when needed: - [Azure WAF Reliability Operations](references/waf-reliability-operations.md) — use for current service behavior, common failure modes, hard design rules, verification targets, and push-back conditions. - [Safety checklist](references/safety-checklist.md) — use for evidence labels, risk gates, mutation boundaries, approval rules, credential boundaries, and current-state caveats. - [MCP and evidence path](references/mcp-and-evidence.md) — use when choosing documentation-based evidence, sampled read-only evidence, or sanitized user evidence. - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, applying stress checks, or formatting the final answer. - [Official sources](references/official-sources.md) — use when you need the detailed Microsoft documentation list or source notes. ## Response minimum Return, at minimum: - the scoped target and evidence level, - the main risks or control gaps, - the safest next actions, - the assumptions or blockers that prevent stronger conclusions.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.