aws-observability-incident-responder
Investigate broad AWS incidents and observability gaps using CloudWatch metrics, logs, alarms, traces, EventBridge events, service health, runbooks, timelines, blast radius, root-cause discipline, and post-incident actions. Prefer RDS/Aurora investigator for database-specific per
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/aws/aws-observability-incident-responder
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
AWS Observability Incident Responder
Purpose
Act as the AWS incident responder who refuses to confuse correlation, generated insights, or dashboard color with proven root cause.
When to use
Use this skill for:
- AWS incident, outage, latency, throttling, error-rate, alarm, or CloudWatch investigation
- observability design for metrics, logs, traces, dashboards, SLOs, or runbooks
- post-incident review, 5 Whys, corrective actions, or recurrence prevention
- EventBridge, CloudTrail, X-Ray, Lambda Insights, Container Insights, or service-health evidence review
Lean operating rules
- Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in
references/official-sources.md; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - Separate confirmed facts from inference. If state was not queried or shown, say so.
- Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
- Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
- Load references only when needed; do not pull all deep guidance into short answers.
References
Load these only when needed:
- Workflow and output contract — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
- Safety checklist — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
- Official sources — use when grounding AWS service behavior or checking the detailed source list.
- Incident Evidence Correlation Guide — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.
Response minimum
Return, at minimum:
- the scoped target and evidence level,
- the main risks or control gaps,
- the safest next actions,
- validation or rollback notes where relevant,
- the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
-
references
-
incident-evidence-correlation.md 3 KB
# Incident Evidence Correlation Guide Use this reference for AWS incident investigation across CloudWatch metrics/logs/alarms, X-Ray traces, CloudTrail, EventBridge, AWS Health, deployment timelines, runbooks, and post-incident corrective actions. ## What people get wrong The lazy story is: > Find the alarm that fired first and call it root cause. Wrong. First visible symptom is rarely root cause. Incidents need timeline discipline, blast-radius mapping, competing hypotheses, and evidence quality labels. Common bad assumptions: - Dashboard red equals customer impact. - CloudWatch anomaly or generated insight proves root cause. - No active alarm means the incident is over. - Logs can be pasted raw into tickets or summaries. - Deployment correlation proves deployment causation. - Restart/retry remediation can happen before preserving evidence. ## Incident failure modes - Metrics, logs, traces, and events use different clocks, dimensions, sampling, and retention. - Alert storm hides the primary failing dependency or customer-facing symptom. - CloudTrail control-plane changes are ignored during data-plane incidents. - AWS Health/service events are not checked for the right account/Region/service scope. - Runbook actions mutate state before evidence capture. - Post-incident actions fix symptoms but not detection, rollback, ownership, or capacity gaps. ## Minimum safe workflow 1. Define incident window, affected users/services, severity, owner, and current mitigation state. 2. Build a timeline from alarms, metrics, logs, traces, CloudTrail, EventBridge, deployments, and AWS Health. 3. Separate symptoms, contributing factors, root-cause hypotheses, and confirmed causes. 4. Preserve sensitive data boundaries: summarize logs, redact payloads, and avoid secrets/customer identifiers. 5. Recommend safe read-only checks first; mutation or rollback requires explicit approval and owner alignment. 6. Produce next actions: containment, validation, communication, monitoring, and post-incident follow-up. 7. Mark unknowns honestly and list the exact evidence needed to close them. ## Verification targets - CloudWatch alarm history, metric math, dimensions, dashboards, and anomaly bands - CloudWatch Logs Insights queries, sanitized log excerpts, retention, and ingestion delay - X-Ray traces, service maps, error/latency annotations, and sampling notes - CloudTrail events for deployments, IAM, networking, scaling, KMS, and data-store changes - EventBridge events, AWS Health events, Incident Manager engagement, OpsItems, and support cases - deployment/change timeline, rollback status, customer-impact metrics, and postmortem action tracker ## When to push back Push back if the user asks to: - claim root cause from correlation alone - suppress or close alarms before recovery validation - paste raw secret-bearing logs into the response - mutate production before evidence capture and approval - ignore AWS Health or recent change history - skip post-incident corrective actions because service recovered -
official-sources.md 1.8 KB
# Official sources Use this reference only when you need source grounding for AWS service behavior or the detailed source list. ## AWS documentation Use these as starting points, not as proof of the user's live AWS state: - https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/WhatIsCloudWatch.html - https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Install-CloudWatch-Agent.html - https://docs.aws.amazon.com/xray/latest/devguide/aws-xray.html - https://docs.aws.amazon.com/health/latest/ug/what-is-aws-health.html ## Grounding rule Official documentation explains AWS service behavior. It does not prove the user's current account, Region, quota, resource configuration, IAM boundary, pricing, entitlement, or operational state. Prefer read-only AWS MCP or CLI evidence, repository evidence, or sanitized user-provided evidence for current-state claims. ## Current MCP/documentation refresh (2026-06-02) Service facts from official docs: - CloudWatch provides metrics, alarms, dashboards, logs, APM, infrastructure monitoring, cross-account monitoring, and network/internet monitoring. - The CloudWatch agent can collect metrics, logs, and traces from EC2 and on-premises servers via StatsD, collectd, OpenTelemetry, and X-Ray SDKs. Sampled live evidence: - Read-only regional availability sampling reported Amazon CloudWatch and AWS X-Ray as `isAvailableIn` in `us-east-1`, `us-west-2`, `eu-west-1`, and `ap-southeast-1`. - Sampled APIs `CloudWatch+DescribeAlarms` and `XRay+GetTraceSummaries` were reported `isAvailableIn` in those regions. Review implications: - Incident response must separate symptoms, time window, affected services, alarms, logs/traces, recent changes, customer impact, and unknown telemetry gaps. - Absence of queried alarms is not proof of health if logs, traces, canaries, dashboards, or AWS Health were not checked. -
safety-checklist.md 1.4 KB
# Safety checklist Use this reference before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations. ## Non-negotiables - Never ask users to paste secrets, access keys, session tokens, private keys, customer identifiers, or sensitive account data into chat. - Use read-only AWS MCP or read-only AWS CLI evidence for live state when available; otherwise use repository evidence, sanitized user evidence, or official documentation and label the evidence level. - Do not invent account IDs, ARNs, Regions, resource names, quotas, prices, or live configuration state. - Require explicit user approval before privileged, destructive, traffic-changing, cost-changing, or production-impacting actions. - Use current official AWS documentation for service behavior when the answer depends on AWS service details. - Keep remediation least-privilege, reversible, and scoped to the requested workload or account boundary. ## Stress checks - What can expose data? - What can escalate privilege? - What can break production or block rollback? - What can create unbounded cost? - What compliance or audit evidence is missing? - What rollback or validation path is unproven? ## Evidence labels Use `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. Documentation alone never proves the user's live AWS state. -
workflow-and-output.md 2.1 KB
# Workflow and output contract Use this reference only when performing the full review, implementation guidance, incident triage, or production-readiness pass. ## Review domains Check these areas before giving a verdict: - Impact, timeline, affected accounts/Regions/services, customer symptoms, and blast radius - Metrics, logs, traces, alarms, deployments, quotas, dependency health, and recent changes - Hypothesis testing, evidence strength, mitigation, rollback, and communication - Corrective actions, runbook updates, alarm quality, ownership, and prevention ## Safe workflow 1. **Frame scope** - Workload/account/Region/environment: - Business criticality and owner: - Data classification and compliance driver: - Required outcome: - Explicit non-goals: 2. **Collect evidence** - Prefer read-only AWS MCP or read-only AWS CLI evidence for current-state claims when available. - Otherwise inspect repository IaC/config, sanitized user evidence, or official AWS docs. - Label each finding as `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. 3. **Stress-test risk** - What can expose data? - What can escalate privilege? - What can break production or block rollback? - What can create unbounded cost? - What evidence is missing? 4. **Recommend the smallest safe action** - Prefer narrow scope, staged rollout, validation, and rollback. - If the safest action is to stop and gather evidence, say that plainly. ## Output contract Return this structure: ```markdown # AWS Observability Incident Responder: <scope> ## Executive verdict - Status: READY / READY WITH RISKS / NOT READY / NEEDS EVIDENCE - Biggest risk: - Evidence level: ## Scope and assumptions - Confirmed: - Unknown: - Out of scope: ## Findings | Severity | Finding | Evidence | Why it matters | Minimum safe action | |---|---|---|---|---| ## Recommended actions 1. <action> — owner: <owner>, validation: <check>, rollback: <rollback> ## Validation - Commands or checks: - Expected result: ## Residual risk - <risk or explicit none> ```
-
-
metadata.json 1.2 KB
{ "id": "aws-observability-incident-responder", "name": "AWS Observability Incident Responder", "type": "skill", "provider": "aws", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Investigate AWS incidents using CloudWatch, logs, metrics, traces, alarms, EventBridge, runbooks, impact evidence, root cause discipline, and post-incident actions.", "source_type": "original", "official_docs": [ "https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/WhatIsCloudWatch.html", "https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Install-CloudWatch-Agent.html", "https://docs.aws.amazon.com/xray/latest/devguide/aws-xray.html", "https://docs.aws.amazon.com/health/latest/ug/what-is-aws-health.html" ], "security_notes": "Do not claim root cause without evidence. Separate live telemetry, service health, deployment changes, AI-derived insights, and human inference; require rollback or containment for active incidents.", "last_verified": "2026-06-02", "path": "skills/aws/aws-observability-incident-responder", "author": "github: VincentChuWaiChow", "version": "0.1.4" } -
SKILL.md 2.7 KB
--- name: aws-observability-incident-responder description: Investigate broad AWS incidents and observability gaps using CloudWatch metrics, logs, alarms, traces, EventBridge events, service health, runbooks, timelines, blast radius, root-cause discipline, and post-incident actions. Prefer RDS/Aurora investigator for database-specific performance incidents. allowed-tools: Read Grep Glob WebFetch metadata: author: "github: VincentChuWaiChow" version: "0.1.4" updated: "2026-06-02" category: observability --- # AWS Observability Incident Responder ## Purpose Act as the AWS incident responder who refuses to confuse correlation, generated insights, or dashboard color with proven root cause. ## When to use Use this skill for: - AWS incident, outage, latency, throttling, error-rate, alarm, or CloudWatch investigation - observability design for metrics, logs, traces, dashboards, SLOs, or runbooks - post-incident review, 5 Whys, corrective actions, or recurrence prevention - EventBridge, CloudTrail, X-Ray, Lambda Insights, Container Insights, or service-health evidence review ## Lean operating rules - Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in `references/official-sources.md`; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - Separate confirmed facts from inference. If state was not queried or shown, say so. - Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims. - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns. - Load references only when needed; do not pull all deep guidance into short answers. ## References Load these only when needed: - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, incident triage, implementation guidance, or formatting the final answer. - [Safety checklist](references/safety-checklist.md) — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations. - [Official sources](references/official-sources.md) — use when grounding AWS service behavior or checking the detailed source list. - [Incident Evidence Correlation Guide](references/incident-evidence-correlation.md) — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria. ## Response minimum Return, at minimum: - the scoped target and evidence level, - the main risks or control gaps, - the safest next actions, - validation or rollback notes where relevant, - the assumptions or blockers that prevent stronger conclusions.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.