Claude Skill

aws-health-events

ALWAYS use this skill in the beginning of any incident investigation, root cause

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download aws-tools-for-devops-agent-skills_aws-health-events-1c971c7.zip · 33 KB
Part of aws/tools-for-devops-agent — 21 skills

Install

skills CLI npx skills add https://github.com/aws/tools-for-devops-agent/tree/main/skills/aws-health-events
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install aws-tools-for-devops-agent@llmmart
Git git clone https://github.com/aws/tools-for-devops-agent.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole aws/tools-for-devops-agent collection as a plugin from our marketplace. Git is the plain clone.

README

AWS Health Events Skill

This skill enables the AWS DevOps Agent to retrieve and analyze AWS Health events during incident investigation, root cause analysis, and operational troubleshooting.

Purpose

When an operational incident occurs, AWS Health events provide critical visibility into service disruptions, scheduled maintenance, and account-specific notifications that may explain or correlate with observed issues. This skill searches AWS Health events by service, time window, severity, region, and event type to surface AWS-side events as potential root causes or contributing factors.

Key Capabilities

  • Retrieve open and resolved AWS Health events via the Health API
  • Search and filter events by service, severity, time range, region, and event type
  • Identify affected resources impacted by Health events
  • Correlate Health events with current incidents to determine root cause or contributing factors
  • Generate health posture reports summarizing account health over configurable time periods

Prerequisites

  • IAM permissions for health:Describe* actions (health:DescribeEvents, health:DescribeEventDetails, health:DescribeAffectedEntities, health:DescribeEventTypes)
  • AWS Health API access (requires AWS Business Support+, Enterprise Support, or Unified Operations support plan)

Limitations

  • The AWS Health API is only available in the us-east-1 region for commercial accounts
  • Health event data retention is limited by the AWS Health service retention period
  • API rate limits apply — the skill implements exponential backoff for throttled requests

Agent Types

This skill is used by the following agent types:

  • Chat tasks — conversational health event lookup, reporting, and analysis
  • Incident RCA — root cause analysis during active incidents

Uploading to AWS DevOps Agent

You can add this skill to your Agent Space in three ways:

Option A: Import from GitHub (recommended)

If you have a GitHub connection configured in your Agent Space, you can import this skill directly from the repository. In the DevOps Agent web app, go to Settings → Add Skill → Import from repository, then point to the skills/aws-health-events directory. When prompted, select the agent types: Chat tasks and Incident RCA. See Importing a skill from a repository for full instructions.

Note: You cannot connect the aws GitHub organization directly because the GitHub connection setup requires admin rights on the organization. Instead, connect your personal GitHub account and select any repository from it during the connection setup. Once a GitHub connection is established, you can import skills from any public repository — including this one — even if it wasn't selected during the connection setup.

Option B: Upload as a zip file

  1. Zip the aws-health-events/ directory (only including allowed extensions):

    cd skills
    zip -r aws-health-events.zip aws-health-events/ -i '*.md' '*.txt' '*.json' '*.yaml' '*.yml' '*.xml' '*.csv' '*.tsv' '*.html' '*.htm' '*.png' '*.jpg' '*.jpeg' '*.gif' '*.svg' '*.webp' '*.pdf' -x '*/.claude/*' '*/scripts/*' '*/README.md' '*/.skilleval.yaml' '*/.skilleval.yml' '*/CHANGELOG.md' '*/evals/*'
    
  2. In the AWS DevOps Agent web app, navigate to the Skills page.

  3. Click Add skill → Upload skill.

  4. Drag and drop the aws-health-events.zip file (max 6 MB).

  5. Select the agent types: Chat tasks and Incident RCA.

  6. Click Upload.

Option C: Upload via the Asset API

Use the AWS DevOps Agent Asset API to programmatically manage skills — useful for CI/CD pipelines or automation workflows. See Managing a skill end-to-end for the full API workflow.

For more details, see Uploading a skill in the AWS DevOps Agent User Guide.

How to Use This Skill

This skill is most suitable for chat and investigation. Below are sample prompts for each use-case.

Chat

  • "Create a health events report for the last 30 days, including breakdowns by service and event type."
  • "Show me all open AWS Health events affecting my account."
  • "Are there any scheduled maintenance events for RDS in the next 7 days?"
  • "Summarize all EC2 health events from the past quarter."
  • "What health events have occurred in us-west-2 this month?"
  • "Give me a health posture summary for my account over the last 90 days."

Investigation

  • "Investigate the RDS connectivity alarm — are there any related AWS Health events?"
  • "We're seeing elevated error rates on our ECS service. Check if there's an AWS-side issue."
  • "Our Lambda functions are timing out. Is there an active Health event for Lambda in us-east-1?"
  • "There's a latency spike on our ALB starting around 2am UTC — correlate with any Health events."
  • "Multiple services in us-west-2 are degraded. Check for regional Health events."
  • "The same EBS performance issue keeps recurring. Are there related Health events or scheduled changes?"

Skill manifest

AWS Health Event Review

Use this skill when investigating an incident and you need to check for AWS-side service events that may be causing or contributing to the observed issue. Also use this skill when a user requests a summary report of AWS Health events over a configurable time period.

When to Use This Skill

Incident Investigation (automatic activation):

  • An active incident may be caused by an AWS service disruption or degradation.
  • You observe service degradation, elevated error rates, latency spikes, connection failures, throttling, or capacity issues.
  • You need to determine whether an AWS-side event is the root cause or a contributing factor to the current incident.
  • You want to correlate observed symptoms with known AWS Health events.

Chat Reporting (on-demand activation):

  • A user requests a health event summary or report for their account.
  • A user wants to review the health posture of their AWS environment over a specific time period.
  • A user asks about recent AWS service issues affecting their account or region.

Prerequisites

  • The account must have an AWS Business Support+, Enterprise Support, or Unified Operations support plan to access the AWS Health API.
  • The agent must have permissions to call the following IAM actions:
    • health:DescribeEvents
    • health:DescribeEventDetails
    • health:DescribeAffectedEntities
    • health:DescribeEventTypes
  • The AWS Health API is only available in the us-east-1 region. All API calls must target the us-east-1 endpoint regardless of where the affected resources are located.
  • Health event data is available for up to 90 days. Events older than 90 days cannot be retrieved via the API.

Step 1: Gather Incident Context

Before searching Health events, extract key details from the current incident:

  1. Affected AWS services — identify the service(s) experiencing issues (e.g., EC2, RDS, Lambda, ELB, ECS).
  2. Timeframe — determine when the incident started and its current duration. Use ISO 8601 timestamps.
  3. Affected resources — collect specific resource identifiers (instance IDs, ARNs, endpoint names, cluster names).
  4. Region and availability zone — identify the AWS region and, if known, the specific availability zone(s) affected.
  5. Symptoms — note the observed symptoms (latency spikes, 5xx errors, connection timeouts, throttling, capacity errors).

Use these details as filter criteria in subsequent steps.


Step 2: Search Health Events

Use the AWS Health API DescribeEvents operation to retrieve events matching the incident context. All calls must target the us-east-1 endpoint.

API call pattern

aws health describe-events \
  --region us-east-1 \
  --filter '{
    "services": ["<SERVICE_CODE>"],
    "startTimes": [{"from": "<ISO-8601-start>"}],
    "regions": ["<affected-region>"],
    "eventStatusCodes": ["open", "closed"],
    "eventTypeCategories": ["issue", "scheduledChange", "accountNotification"]
  }' \
  --max-results 100

Filtering strategies

Strategy How to Apply
By service Use the services filter with the AWS Health service code (e.g., EC2, RDS, ELASTICLOADBALANCING), because service-specific events are most likely to correlate with the incident.
By time range Use startTimes with a from value set to 7 days before the incident start, because events that started before the incident may still be active and causing impact.
By region Use the regions filter to scope events to the affected region, because regional events are more likely to impact the specific resources under investigation.
By availability zone Use the availabilityZones filter when the incident is isolated to a specific AZ, because AZ-scoped events have the highest correlation with AZ-specific failures.
By status Include both open and closed statuses, because recently closed events may have caused residual impact that is still being observed.
By event scope Include both ACCOUNT_SPECIFIC and PUBLIC events, because public service events affect all accounts in the region while account-specific events target your resources directly.

Pagination handling

  • Follow the nextToken from each response to retrieve subsequent pages.
  • Continue paginating until nextToken is null or a maximum of 500 events have been collected.
  • Set maxResults to 100 per page for efficient retrieval.

Step 3: Filter Relevant Events

Before retrieving full event details, filter the events returned in Step 2 to identify only those relevant to the current investigation. This avoids unnecessary DescribeEventDetails calls for events that are clearly unrelated.

Relevance filtering criteria

Evaluate each event from the DescribeEvents response using these fields (available without calling DescribeEventDetails):

Field Relevance Signal
service Must match one of the affected services from the incident context, or a related service from the Service Dependency Map
eventTypeCategory Prioritize issue events for active incidents; include scheduledChange if the incident coincides with a maintenance window
eventTypeCode Match against known operational event patterns (e.g., AWS_EC2_OPERATIONAL_ISSUE, AWS_RDS_MAINTENANCE)
statusCode Prioritize open events; include closed only if the event ended within 2 hours of the incident start
startTime / endTime The event's active period must overlap with the incident timeframe
region / availabilityZone Must match the incident's affected region or AZ

Filtering rules

  1. Keep events where the service matches an affected service or a related service from the Service Dependency Map.
  2. Keep events where the active period (startTime to endTime, or to present if open) overlaps with the incident timeframe.
  3. Keep events where the region or availabilityZone matches the incident's affected region/AZ.
  4. Discard accountNotification events unless the incident context specifically suggests an account-level issue (e.g., abuse notification, certificate expiry).
  5. Discard closed events that ended more than 2 hours before the incident started (unlikely to be contributing).

Result

After filtering, proceed to Step 4 only with the relevant subset of events. If all events are filtered out, report that no relevant Health events were found and suggest alternative investigation paths (see Step 7).


Step 4: Get Event Details

For each relevant event identified in Step 3, retrieve full descriptions and timelines using DescribeEventDetails.

API call pattern

aws health describe-event-details \
  --region us-east-1 \
  --event-arns '["<arn-1>", "<arn-2>", ..., "<arn-10>"]'

Batching rules

  • The API accepts a maximum of 10 event ARNs per request.
  • If more than 10 relevant events need details, issue multiple batched calls of up to 10 ARNs each until all relevant events are detailed.

Extract from each event detail

  • Event description — the latestDescription text explaining the event.
  • Timeline — start time, end time (null if ongoing), last updated time.
  • Status — current status (open, closed, upcoming).
  • Service and region — confirm the affected service and region.
  • Event type — the category (issue, scheduledChange, accountNotification) and specific type code.

Handling failedSet

  • The response contains a successfulSet and a failedSet.
  • If any event ARNs appear in failedSet, report the failed ARN and error message to the operator.
  • Continue processing all events from successfulSet without blocking on failures.

Step 5: Identify Affected Entities

For events with eventScopeCode of ACCOUNT_SPECIFIC, retrieve the list of affected resources using DescribeAffectedEntities.

Important: Only call DescribeAffectedEntities for ACCOUNT_SPECIFIC events. PUBLIC events do not return entity data.

API call pattern

aws health describe-affected-entities \
  --region us-east-1 \
  --filter '{"eventArns": ["<event-arn>"]}'
  --max-results 100

Entity matching

When the incident context includes specific resource identifiers:

  1. Retrieve all affected entities for the event (paginate up to 500 entities per event using nextToken).
  2. Perform exact string matching of each entity's entityValue against the incident context resource identifiers.
  3. Present matched entities in a separate section before non-matched entities.
  4. Include entity status (IMPAIRED, UNIMPAIRED, UNKNOWN, PENDING) and last updated time for each entity.

Error handling

  • If DescribeAffectedEntities returns an error for a specific event ARN, report the event ARN that failed and continue processing remaining events.

Step 6: Correlate with Incident

Score each Health event for relevance to the current incident using the following criteria:

Relevance scoring

Classification Criteria Label
High Matching service + overlapping timeframe + matching affected resource (or matching region/AZ if no resource IDs available) Likely contributing factor (if event is open)
Medium Matching service + overlapping timeframe (no resource match) Likely contributing factor (if event is open)
Low Matching service only (no timeframe overlap) Background context

Scoring rules

  • Service match: The event's service code matches one of the affected services from the incident context.
  • Timeframe overlap: The event's active period (start time through end time, or through present if still open) intersects with the incident's timeframe (start time through end time, or through present if ongoing).
  • Region/AZ match: The event's region or availability zone matches the incident's affected region or AZ.
  • Resource match: At least one affected entity's entityValue matches a resource identifier from the incident context.

Contributing factor labeling

  • Any open event classified as High or Medium relevance SHALL be labeled as a "likely contributing factor" in addition to its relevance classification.
  • Closed events with High relevance should be noted as potential recent causes if the incident started shortly after the event closed.

When resource identifiers are unavailable

If the incident context does not include specific resource identifiers, score relevance using only service, timeframe, and region/AZ factors:

  • High: Matching service + overlapping timeframe + matching region or AZ
  • Medium: Matching service + overlapping timeframe
  • Low: Matching service only

Step 7: Present Structured Output

Present findings in a clear, structured format organized for quick comprehension and action.

Output structure

  1. Summary — total events found, broken down by category and status.
  2. Correlated events — grouped by event type category in this order:
    • Issues (service disruptions) — present first
    • Scheduled changes (maintenance) — present second
    • Account notifications — present last
  3. Within each group — sort by:
    • Relevance classification (High → Medium → Low)
    • Then by start time descending (most recent first)
  4. Per event — include:
    • Event type category and service
    • Region and availability zone (if applicable)
    • Status (open/closed/upcoming)
    • Start time and end time (ISO 8601)
    • Description (summarized to 256 characters max)
    • Relevance classification and matching criteria
    • Contributing factor label (if applicable)
  5. Actionable next steps — for each correlated event, include at least one recommendation such as:
    • Check specific affected resources
    • Review related service limits or quotas
    • Verify recent configuration changes
    • Monitor the AWS Health Dashboard for updates
    • Contact AWS Support if the event is ongoing

When no events are found

If no relevant Health events are identified:

  • Explicitly state that no matching AWS Health events were found.
  • Confirm the search parameters used (service, time range, region).
  • Recommend checking other potential causes:
    • Recent deployments or configuration changes
    • Resource limits or quota exhaustion
    • Network connectivity issues
    • Application-level errors

Decision Tree: Event Search Strategy

Is this a chat-based health report request?
├── YES → Search the user-specified time period (default 30 days, max 90 days)
│         Organize results by category, service, and status
│         Present as a summary report
└── NO → Continue with incident investigation flow below

Is the affected AWS service known?
├── YES → Search events for that service within the past 7 days
│   ├── Events found → Proceed to Step 3 (Filter Relevant Events)
│   └── No events found → Broaden to related services (see Service Dependency Map)
│       ├── Events found → Proceed to Step 3
│       └── No events found → Expand time window to 14 days and retry
│           ├── Events found → Proceed to Step 3
│           └── No events found → Report no events found, suggest other investigation paths
└── NO → Search all services filtered by region and availability zone (past 7 days)
    ├── Events found → Proceed to Step 3
    └── No events found → Expand time window to 14 days
        ├── Events found → Proceed to Step 3
        └── No events found → Report no events found, suggest other investigation paths

Does the incident involve a specific availability zone?
├── YES → Include the AZ filter in all searches above
└── NO → Filter by region only

Service Dependency Map

When the initial service-specific search returns no results, broaden the search to related services that share infrastructure dependencies:

Primary Service Related Services to Check
ELB / ALB / NLB EC2, VPC, Route 53
RDS EC2, EBS
ECS / EKS EC2, VPC, ELB
Lambda VPC, CloudWatch
CloudFront S3, Route 53
API Gateway Lambda, VPC
ElastiCache EC2, VPC
DynamoDB VPC (if VPC endpoints used)
S3 CloudFront, VPC (if VPC endpoints used)
Kinesis EC2, VPC

Search up to 3 related services when broadening. Use the Health API service codes from the references document (e.g., ELASTICLOADBALANCING for ELB, ROUTE53 for Route 53).


Error Handling

Error Condition Agent Behavior
Missing health:Describe* permissions Report the missing permissions and specify the required IAM actions: health:DescribeEvents, health:DescribeEventDetails, health:DescribeAffectedEntities, health:DescribeEventTypes. Provide the IAM policy snippet needed.
Throttling (HTTP 429) Retry with exponential backoff: wait 1s → 2s → 4s (max 3 retries). If still throttled after 3 retries, report that the Health API is currently rate-limited and recommend trying again shortly.
Service error (HTTP 5xx) Report the error code and recommend the operator check the AWS Health Dashboard directly as a fallback.
Timeout (30 seconds) Cancel the request and report a timeout error. Suggest the operator check the Health Dashboard directly or retry with narrower filters.
Zero events found Report that no events matched the specified filters. Confirm the search parameters used. Suggest broadening the search or checking other investigation paths.
Invalid time range (start > end) Report the invalid time range error. Ask the operator to provide corrected timestamps.
DescribeEventDetails failedSet Report the failed event ARNs and error messages. Continue processing events from the successfulSet.
DescribeAffectedEntities error Report the event ARN for which entity retrieval failed. Continue processing remaining events.
Unknown service name (chat report) Inform the user the service was not recognized. List services that have events in the requested time period.

Tips for Effective Health Event Review

  • Always check us-east-1: The Health API endpoint is only in us-east-1, regardless of where your resources are located.
  • Start narrow, then broaden: Begin with the specific affected service and a 7-day window. Only expand if no results are found.
  • Check both open and closed events: A recently closed event may still be causing residual impact.
  • Correlate with support cases: If a Health event references a service disruption, check if related support cases exist using the support-cases skill.
  • Account-specific vs public events: Account-specific events directly affect your resources. Public events are service-wide but may still impact you.
  • Look at scheduled changes: Upcoming or recent maintenance windows can explain transient issues that resolve on their own.
Files (tools-for-devops-agent)
  • evals
    • benchmark.json 77.6 KB
      {
        "skill_name": "aws-health-events",
        "skill_path": "/Users/udid/AWS Work/Services/DevOps Agent/aws-devops-agent-skill-central/skills/aws-health-events",
        "eval_count": 8,
        "runs_per_eval": 1,
        "metadata": {
          "timestamp": "2026-05-27T13:31:37Z"
        },
        "runs": [
          {
            "eval_id": "incident-health-event-lookup",
            "run_index": 0,
            "with_skill": {
              "eval_id": "incident-health-event-lookup",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output calls DescribeEvents with the RDS service filter",
                  "passed": true,
                  "evidence": "The output states 'I'll check AWS Health events for RDS' and the search parameters table shows 'Services | RDS, EC2, VPC, EBS', indicating DescribeEvents was called with RDS as a service filter. However, the actual API call (DescribeEvents) is not explicitly shown in the output text.",
                  "method": "llm",
                  "confidence": 0.75,
                  "uncertain": false
                },
                {
                  "text": "The output filters events by the us-west-2 region",
                  "passed": true,
                  "evidence": "The search parameters table explicitly shows 'Region | us-west-2' confirming the events were filtered by us-west-2.",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output searches within the past 7 days from the incident start time",
                  "passed": false,
                  "evidence": "The search parameters show 'Time range | 2026-05-13 to present (14 days)'. The assertion expects a 7-day search window, but the output used a 14-day window instead.",
                  "method": "llm",
                  "confidence": 0.8,
                  "uncertain": false
                },
                {
                  "text": "The output includes both account-specific and public event scopes",
                  "passed": false,
                  "evidence": "The output does not mention event scopes (account-specific vs public). The search parameters table lists services, region, time range, statuses, and categories but does not reference event scope filtering.",
                  "method": "llm",
                  "confidence": 0.7,
                  "uncertain": false
                },
                {
                  "text": "The output presents matching events or clearly states no events were found",
                  "passed": true,
                  "evidence": "The output clearly states 'No matching AWS Health events were found for RDS, EC2, VPC, or EBS in us-west-2 over the past 14 days' under a heading 'Results: No AWS Health Events Found'.",
                  "method": "llm",
                  "confidence": 0.99,
                  "uncertain": false
                },
                {
                  "text": "The output does not produce an API error or permission denied message",
                  "passed": true,
                  "evidence": "The output contains no API errors, permission denied messages, or access denied responses. It successfully completed the search and reported results.",
                  "method": "llm",
                  "confidence": 0.99,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.6666666666666666,
              "summary": "With skill: 67% assertions passed",
              "execution_metrics": {
                "tool_calls": 4,
                "token_counts": {
                  "input_tokens": 5,
                  "output_tokens": 1269,
                  "cache_read_input_tokens": 56662,
                  "cache_creation_input_tokens": 28998
                }
              },
              "timing": {
                "elapsed_seconds": 61.42067291700005
              },
              "raw_output": "I'll check AWS Health events for RDS in us-west-2 to see if there's a service-side issue explaining your connection timeouts.Let me also check EBS (another RDS dependency) and broaden the search slightly.## Results: No AWS Health Events Found\n\nNo matching AWS Health events were found for RDS, EC2, VPC, or EBS in **us-west-2** over the past 14 days. This means AWS is not reporting any service-side disruptions or scheduled maintenance that would explain your connection timeouts.\n\n### Search Parameters Used\n| Filter | Value |\n|--------|-------|\n| Services | RDS, EC2, VPC, EBS |\n| Region | us-west-2 |\n| Time range | 2026-05-13 to present (14 days) |\n| Statuses | open, closed |\n| Categories | issue, scheduledChange, accountNotification |\n\n### Recommended Next Steps\n\nSince this doesn't appear to be an AWS-side issue, investigate these common causes of RDS connection timeouts:\n\n1. **Security group / network ACL changes** \u2014 verify inbound rules on the RDS instance's security group still allow traffic from your application subnets\n2. **Max connections exhausted** \u2014 check `DatabaseConnections` in CloudWatch; if at the limit, connections will be refused\n3. **Storage full** \u2014 check `FreeStorageSpace`; a full disk can cause RDS to reject connections\n4. **Recent deployments** \u2014 check if any application or infrastructure changes were deployed ~2 hours ago\n5. **DNS resolution** \u2014 if using a custom endpoint or Route 53, verify it still resolves to the correct RDS address\n6. **RDS instance status** \u2014 run `aws rds describe-db-instances --region us-west-2` to confirm the instance is in `available` state\n\nWould you like me to help investigate any of these?"
            },
            "without_skill": {
              "eval_id": "incident-health-event-lookup",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output calls DescribeEvents with the RDS service filter",
                  "passed": true,
                  "evidence": "The output shows the aws health describe-events command with services=RDS filter, but it was not actually executed - only provided as a recommendation for the user to run",
                  "method": "llm",
                  "confidence": 0.6,
                  "uncertain": false
                },
                {
                  "text": "The output filters events by the us-west-2 region",
                  "passed": true,
                  "evidence": "The describe-events command includes regions=us-west-2 in the filter parameter",
                  "method": "llm",
                  "confidence": 0.9,
                  "uncertain": false
                },
                {
                  "text": "The output searches within the past 7 days from the incident start time",
                  "passed": false,
                  "evidence": "No time-based filter (startTime/endTime for past 7 days) is included in the describe-events command; only eventStatusCodes are specified",
                  "method": "llm",
                  "confidence": 0.9,
                  "uncertain": false
                },
                {
                  "text": "The output includes both account-specific and public event scopes",
                  "passed": false,
                  "evidence": "The command does not include any eventScopeCode filter for account-specific or public events; neither scope is mentioned",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output presents matching events or clearly states no events were found",
                  "passed": false,
                  "evidence": "The output states it cannot query the API due to missing credentials and provides commands for the user to run - it does not present actual events or definitively state no events were found",
                  "method": "llm",
                  "confidence": 0.85,
                  "uncertain": false
                },
                {
                  "text": "The output does not produce an API error or permission denied message",
                  "passed": true,
                  "evidence": "No API error or permission denied message is produced; the output proactively explains credentials are unavailable before attempting any API call",
                  "method": "llm",
                  "confidence": 0.9,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.5,
              "summary": "Without skill: 50% assertions passed",
              "execution_metrics": {
                "tool_calls": 2,
                "token_counts": {
                  "input_tokens": 5,
                  "output_tokens": 719,
                  "cache_read_input_tokens": 57220,
                  "cache_creation_input_tokens": 13855
                }
              },
              "timing": {
                "elapsed_seconds": 36.07404095799939
              },
              "raw_output": "I'll check AWS Health for any events affecting RDS in us-west-2.There are no valid AWS credentials configured in this environment, so I can't query the AWS Health API directly. Here's how you can check yourself:\n\n**Check AWS Health events:**\n```bash\n# AWS Health API must be called from us-east-1 (it's a global endpoint)\naws health describe-events \\\n  --filter \"services=RDS,regions=us-west-2,eventStatusCodes=open,upcoming,closed\" \\\n  --region us-east-1 \\\n  --output table\n```\n\n**Get event details (if events are found):**\n```bash\naws health describe-event-details \\\n  --event-arns <arn-from-above> \\\n  --region us-east-1\n```\n\n**Other quick checks for your RDS connection timeouts:**\n\n1. **RDS instance status:**\n   ```bash\n   aws rds describe-db-instances --region us-west-2 \\\n     --query 'DBInstances[*].[DBInstanceIdentifier,DBInstanceStatus]' --output table\n   ```\n\n2. **Recent RDS events (maintenance, failover, etc.):**\n   ```bash\n   aws rds describe-events --region us-west-2 --duration 180 --output table\n   ```\n\n3. **Security group / network changes** \u2014 verify the instance's security group still allows inbound connections on the DB port from your application's CIDR/security group.\n\n4. **CloudWatch metrics** \u2014 check `DatabaseConnections`, `CPUUtilization`, and `FreeableMemory` for the past 3 hours to see if the instance is overwhelmed or unreachable.\n\nWould you like me to run any of these once you have credentials configured, or help investigate further in a different way?"
            },
            "delta_pass_rate": 0.16666666666666663
          },
          {
            "eval_id": "event-detail-extraction",
            "run_index": 0,
            "with_skill": {
              "eval_id": "event-detail-extraction",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output calls DescribeEventDetails to retrieve full event information",
                  "passed": true,
                  "evidence": "The output says 'Let me get the full details' and shows detailed event information (description, timeline, root cause) consistent with a DescribeEventDetails call, but doesn't explicitly name the API",
                  "method": "llm",
                  "confidence": 0.7,
                  "uncertain": false
                },
                {
                  "text": "The output includes the event description text for each event",
                  "passed": true,
                  "evidence": "The output includes extensive event description text including Root Cause, Impact, and Timeline Summary sections which correspond to the event description from DescribeEventDetails",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output includes start time and end time in ISO 8601 format",
                  "passed": false,
                  "evidence": "Times are shown as '2026-05-07 16:25 PDT' and '2026-05-08 20:04 PDT' which is not ISO 8601 format (would need T separator and proper timezone offset like 2026-05-07T16:25:00-07:00)",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output includes the current status (open, closed, or upcoming) for each event",
                  "passed": true,
                  "evidence": "The output explicitly shows 'Status | **Closed**' which is one of the expected status values (open, closed, or upcoming)",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output batches requests with no more than 10 event ARNs per API call",
                  "passed": false,
                  "evidence": "The output does not mention batching at all. Only 1 event was found so batching wasn't needed, but there's no evidence the output implements or describes the batching constraint",
                  "method": "llm",
                  "confidence": 0.6,
                  "uncertain": false
                },
                {
                  "text": "The output handles any failedSet entries by reporting them and continuing with successfulSet",
                  "passed": false,
                  "evidence": "The output does not mention failedSet handling at all - no reporting of failed entries or explicit mention of processing successfulSet separately",
                  "method": "llm",
                  "confidence": 0.75,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.5,
              "summary": "With skill: 50% assertions passed",
              "execution_metrics": {
                "tool_calls": 2,
                "token_counts": {
                  "input_tokens": 5,
                  "output_tokens": 1192,
                  "cache_read_input_tokens": 66606,
                  "cache_creation_input_tokens": 22127
                }
              },
              "timing": {
                "elapsed_seconds": 46.91727712500142
              },
              "raw_output": "I'll retrieve EC2 Health events in us-east-1 and then get their full details.Found one EC2 event. Let me get the full details.Here are the full details for the EC2 Health event found in us-east-1:\n\n---\n\n## EC2 Operational Issue \u2014 RESOLVED\n\n| Field | Value |\n|-------|-------|\n| **Event ARN** | `arn:aws:health:us-east-1::event/EC2/AWS_EC2_OPERATIONAL_ISSUE/AWS_EC2_OPERATIONAL_ISSUE_FFA94_A4DC515D30F` |\n| **Service** | EC2 |\n| **Type** | `AWS_EC2_OPERATIONAL_ISSUE` (issue) |\n| **Region** | us-east-1 |\n| **Availability Zone** | use1-az4 |\n| **Status** | **Closed** |\n| **Scope** | PUBLIC (affected all accounts in the zone) |\n| **Start Time** | 2026-05-07 16:25 PDT |\n| **End Time** | 2026-05-08 20:04 PDT |\n| **Duration** | ~27.5 hours |\n\n---\n\n### Root Cause\n\nA **thermal event** in a single data center facility within use1-az4 caused a **loss of power**. Servers automatically shut down when temperatures exceeded operating thresholds to protect hardware.\n\n### Impact\n\n- Impaired EC2 instances and degraded EBS volumes in use1-az4\n- Cascading impact to dependent services: ELB, EKS, ElastiCache, OpenSearch, Redshift, Kafka (MSK), NAT Gateway, IoT Core, SageMaker\n\n### Timeline Summary\n\n| Time (PDT) | Update |\n|------------|--------|\n| May 7, 4:20 PM | Thermal event begins, instance impairments start |\n| May 7, 5:06 PM | Traffic shifted away from impacted AZ for most services |\n| May 7, 9:12 PM | Power restored to a subset of infrastructure; early recovery signs |\n| May 8, 6:51 AM | Continued work on cooling capacity; Redshift errors also reported |\n| May 8, 9:30 AM | Redshift fully recovered |\n| May 8, 12:29 PM | Recovery slower than anticipated; measured cooling restoration underway |\n| May 8, 4:00 PM | Steady improvements observed; additional cooling capacity showing progress |\n| May 8, 8:04 PM | Cooling stabilized to pre-event levels; majority of EC2/EBS recovered. **Event resolved.** |\n\n### Previously Impacted Services (Now Normal)\n\nIOT, EKS, ElastiCache, OpenSearch, ELB, Ka"
            },
            "without_skill": {
              "eval_id": "event-detail-extraction",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output calls DescribeEventDetails to retrieve full event information",
                  "passed": false,
                  "evidence": "The output never progressed past authentication errors and profile selection. No DescribeEventDetails call was made.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                },
                {
                  "text": "The output includes the event description text for each event",
                  "passed": false,
                  "evidence": "No event data was retrieved at all. The output is entirely about credential/profile selection.",
                  "method": "llm",
                  "confidence": 0.98,
                  "uncertain": false
                },
                {
                  "text": "The output includes start time and end time in ISO 8601 format",
                  "passed": false,
                  "evidence": "No event data was retrieved, so no timestamps in any format are present.",
                  "method": "llm",
                  "confidence": 0.98,
                  "uncertain": false
                },
                {
                  "text": "The output includes the current status (open, closed, or upcoming) for each event",
                  "passed": false,
                  "evidence": "No event data was retrieved, so no status information is present.",
                  "method": "llm",
                  "confidence": 0.98,
                  "uncertain": false
                },
                {
                  "text": "The output batches requests with no more than 10 event ARNs per API call",
                  "passed": false,
                  "evidence": "No API requests were made at all, so no batching behavior can be observed.",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output handles any failedSet entries by reporting them and continuing with successfulSet",
                  "passed": false,
                  "evidence": "No API calls were made, so no failedSet handling occurred.",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.0,
              "summary": "Without skill: 0% assertions passed",
              "execution_metrics": {
                "tool_calls": 3,
                "token_counts": {
                  "input_tokens": 6,
                  "output_tokens": 1125,
                  "cache_read_input_tokens": 81618,
                  "cache_creation_input_tokens": 14790
                }
              },
              "timing": {
                "elapsed_seconds": 45.30411141699733
              },
              "raw_output": "I'll query AWS Health for EC2 events in us-east-1.The request failed with an authentication error \u2014 the current AWS credentials are invalid or not configured. Let me check what's available.I only see Admin profiles available. Since this is a read-only operation (describing Health events), I'll proceed with one of the available profiles. Which account would you like me to check? The available profiles are:\n\n- `udid-Admin` (main account)\n- `udid+linked1-Admin`\n- `udid+linked2-Admin`\n- `udid+linked3-Admin`\n- `udid+datacollection-Admin`\n- `udid+datacollection4-Admin`\n- `udid+datacollection5-Admin`\n- `udid+datacollection6-Admin`\n\nAlso note: the AWS Health API requires a Business or Enterprise AWS Support plan. Which profile/account should I use to query the Health events?Which AWS profile should I use to query the EC2 Health events? The default credentials aren't valid, and I see several profiles available (listed above). Let me know which account you'd like me to check."
            },
            "delta_pass_rate": 0.5
          },
          {
            "eval_id": "affected-entity-matching",
            "run_index": 0,
            "with_skill": {
              "eval_id": "affected-entity-matching",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output calls DescribeAffectedEntities for the account-specific event",
                  "passed": false,
                  "evidence": "The output did call DescribeAffectedEntities, but the event found was PUBLIC, not account-specific. The output explicitly states 'No account-specific AWS Health events were found' so the assertion's condition ('for the account-specific event') cannot be satisfied.",
                  "method": "llm",
                  "confidence": 0.85,
                  "uncertain": false
                },
                {
                  "text": "The output performs exact string matching of entity values against the incident resource identifiers",
                  "passed": false,
                  "evidence": "The output does not show any string matching of returned entity values against the instance IDs i-0abc123def456 and i-0def789ghi012. It only notes that the entities response showed 'UNKNOWN' and recommends checking instances via EC2 commands instead.",
                  "method": "llm",
                  "confidence": 0.92,
                  "uncertain": false
                },
                {
                  "text": "The output presents matched entities separately from non-matched entities",
                  "passed": false,
                  "evidence": "The output does not present any matched vs non-matched entities. No entity-level breakdown is provided since the public event didn't return per-account entity data.",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output includes entity status (IMPAIRED, UNIMPAIRED, UNKNOWN, or PENDING) for each entity",
                  "passed": false,
                  "evidence": "The output mentions 'UNKNOWN' once in the context of the API response but does not present entity status for each of the user's specific instances. It only notes generically that the response showed 'UNKNOWN'.",
                  "method": "llm",
                  "confidence": 0.8,
                  "uncertain": false
                },
                {
                  "text": "The output includes last updated time for each affected entity",
                  "passed": false,
                  "evidence": "The output contains no last updated time for any affected entities. Only the event's start/end period is shown, not per-entity timestamps.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.0,
              "summary": "With skill: 0% assertions passed",
              "execution_metrics": {
                "tool_calls": 7,
                "token_counts": {
                  "input_tokens": 8,
                  "output_tokens": 2081,
                  "cache_read_input_tokens": 159082,
                  "cache_creation_input_tokens": 24636
                }
              },
              "timing": {
                "elapsed_seconds": 76.7529924579976
              },
              "raw_output": "I'll investigate AWS Health events for EC2 in us-east-1 and check which of your specific instances are affected.The only EC2 event found in us-east-1 is a **PUBLIC** event (not account-specific). Let me broaden the search to include a wider time window and check for any account-specific events.Let me get the event details and check affected entities for the us-east-1 event, and also try to query affected entities directly for your instances.The event found is a **PUBLIC** event (`eventScopeCode: PUBLIC`), not an account-specific one. Public events don't return per-account affected entity data \u2014 the entities response only shows `\"UNKNOWN\"`. Let me search more broadly for any account-specific EC2 events that might not have appeared yet.## Findings\n\n**No account-specific AWS Health events were found for EC2 in us-east-1 for this account.**\n\nThe only EC2 event in us-east-1 is a **public** (service-wide) event:\n\n| Field | Value |\n|-------|-------|\n| **Event** | `AWS_EC2_OPERATIONAL_ISSUE` |\n| **Region** | us-east-1, AZ: use1-az4 |\n| **Status** | Closed |\n| **Period** | May 7 4:20 PM \u2013 May 9 3:04 AM PDT |\n| **Cause** | Thermal event (data center power loss) in use1-az4 |\n| **Scope** | PUBLIC (not account-specific) |\n\nBecause this is a public event, the Health API does not return per-account affected entity data \u2014 I cannot determine from the Health API alone whether your instances `i-0abc123def456` and `i-0def789ghi012` were impacted by this event.\n\n### Recommended next steps to identify if your resources were affected:\n\n1. **Check instance status directly:**\n   ```\n   aws ec2 describe-instance-status --instance-ids i-0abc123def456 i-0def789ghi012 --region us-east-1\n   ```\n\n2. **Check which AZ your instances are in** \u2014 the event only affected `use1-az4`:\n   ```\n   aws ec2 describe-instances --instance-ids i-0abc123def456 i-0def789ghi012 --region us-east-1 --query 'Reservations[].Instances[].[InstanceId,Placement.AvailabilityZone,State.Name]'\n   ```\n\n3. **Check the \"Your Ac"
            },
            "without_skill": {
              "eval_id": "affected-entity-matching",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output calls DescribeAffectedEntities for the account-specific event",
                  "passed": false,
                  "evidence": "The output mentions querying affected entities for the us-east-1 event, but since only a public (not account-specific) event was found, DescribeAffectedEntities was not called for an account-specific event specifically. The output explicitly states the event is 'public (not account-specific)'.",
                  "method": "llm",
                  "confidence": 0.7,
                  "uncertain": false
                },
                {
                  "text": "The output performs exact string matching of entity values against the incident resource identifiers",
                  "passed": false,
                  "evidence": "The output does not perform exact string matching of entity values against the incident resource identifiers. Instead, it validates the format of the provided IDs and checks AZ placement. No entity-to-resource-ID comparison is shown.",
                  "method": "llm",
                  "confidence": 0.85,
                  "uncertain": false
                },
                {
                  "text": "The output presents matched entities separately from non-matched entities",
                  "passed": false,
                  "evidence": "The output does not present matched entities separately from non-matched entities. It presents instances grouped by AZ and concludes no resources are affected, but doesn't use a matched/non-matched separation structure.",
                  "method": "llm",
                  "confidence": 0.9,
                  "uncertain": false
                },
                {
                  "text": "The output includes entity status (IMPAIRED, UNIMPAIRED, UNKNOWN, or PENDING) for each entity",
                  "passed": false,
                  "evidence": "The output shows EC2 instance states ('running', 'stopped') not Health API entity statuses (IMPAIRED, UNIMPAIRED, UNKNOWN, or PENDING). No Health entity status values are mentioned anywhere.",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output includes last updated time for each affected entity",
                  "passed": false,
                  "evidence": "No last updated time is provided for any affected entity. Only the event timeframe (May 7-9, 2026) is mentioned, not per-entity timestamps.",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.0,
              "summary": "Without skill: 0% assertions passed",
              "execution_metrics": {
                "tool_calls": 14,
                "token_counts": {
                  "input_tokens": 11,
                  "output_tokens": 3232,
                  "cache_read_input_tokens": 224300,
                  "cache_creation_input_tokens": 21966
                }
              },
              "timing": {
                "elapsed_seconds": 112.17324075000215
              },
              "raw_output": "I'll investigate the AWS Health event affecting your EC2 instances in us-east-1. Let me first check your AWS credentials and then query the Health API.I see only Admin profiles available. Since querying Health events is a read-only operation, I'll proceed with the default credentials. Let me check the caller identity and then query AWS Health for events affecting your instances.The default credentials are invalid. I'll need to use one of your configured profiles. Note that only Admin profiles are available \u2014 I'll use them strictly for read-only operations (Health API queries). Let me try with the primary account profile.Good, the `udid-Admin` profile works. I'm using Admin credentials (read-only operations only). Let me now query AWS Health for EC2 events and check which of your instances are affected.The Health API returned public-scope events. Let me now check for account-specific events that might affect your instances directly, and also query the affected entities for the us-east-1 event.The Health API returned a **public** event (not account-specific). Let me check if there are account-specific events, and also verify whether your instances are in the affected availability zone (use1-az4).The instance IDs you provided (`i-0abc123def456` and `i-0def789ghi012`) are not valid EC2 instance ID format \u2014 they appear to be placeholder/redacted IDs. EC2 instance IDs are 17 hex characters after the `i-` prefix (e.g., `i-0a1b2c3d4e5f67890`).\n\nLet me try a broader approach to find account-specific events and also check instances in your account.Here's what I found:\n\n## Summary\n\n**No resources in this account are affected by the Health event.**\n\n### Health Event Details\nThe only EC2 Health event in us-east-1 is a **closed, public** (not account-specific) operational issue:\n- **Event:** `AWS_EC2_OPERATIONAL_ISSUE` \u2014 a thermal event causing power loss in **use1-az4** on May 7-9, 2026\n- **Status:** Resolved\n- **Affected AZ:** `use1-az4`\n\n### Your Account's Instances\nNone of yo"
            },
            "delta_pass_rate": 0.0
          },
          {
            "eval_id": "relevance-correlation",
            "run_index": 0,
            "with_skill": {
              "eval_id": "relevance-correlation",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output classifies each event as High, Medium, or Low relevance",
                  "passed": true,
                  "evidence": "Correlated events in the Issues and Scheduled Changes tables are labeled with **Low** relevance, but the 'Discarded' events don't receive an explicit High/Medium/Low label \u2014 they're just listed as 'no correlation'. Not all events get a formal H/M/L classification.",
                  "method": "llm",
                  "confidence": 0.6,
                  "uncertain": false
                },
                {
                  "text": "The output explains the matching criteria used for each classification",
                  "passed": true,
                  "evidence": "Each classified event has a 'Rationale' column explaining why it was rated Low \u2014 citing factors like timeframe mismatch, wrong region, service relationship, and residual impact windows.",
                  "method": "llm",
                  "confidence": 0.9,
                  "uncertain": false
                },
                {
                  "text": "The output orders results by relevance (High first, then Medium, then Low)",
                  "passed": true,
                  "evidence": "All correlated events are rated Low, so there's no misordering possible. However, the output organizes by event type (Issues, then Scheduled Changes, then Discarded) rather than explicitly by relevance level, so the ordering principle isn't clearly demonstrated.",
                  "method": "llm",
                  "confidence": 0.5,
                  "uncertain": false
                },
                {
                  "text": "The output labels open events with High or Medium relevance as likely contributing factors",
                  "passed": true,
                  "evidence": "There are no open events with High or Medium relevance in the results. The assertion is vacuously satisfied \u2014 there's nothing to label, and the conclusion explicitly states no correlating events were found.",
                  "method": "llm",
                  "confidence": 0.6,
                  "uncertain": false
                },
                {
                  "text": "The output includes at least one actionable next step per correlated event",
                  "passed": false,
                  "evidence": "The correlated events (Route 53 issue, Lambda lifecycle events) do not have individual actionable next steps attached to them. The 'Recommended Next Steps' section provides general investigation guidance for the Lambda timeout, not per-event recommendations.",
                  "method": "llm",
                  "confidence": 0.85,
                  "uncertain": false
                },
                {
                  "text": "The output considers service match, timeframe overlap, and region when scoring",
                  "passed": true,
                  "evidence": "Rationales explicitly reference service match ('Route 53 could affect Lambda DNS', 'Same service', 'unrelated service'), timeframe overlap ('closed ~4 days before', 'no timeframe overlap', '5 months away'), and region ('global', 'wrong region us-east-1 vs us-west-2').",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.8333333333333334,
              "summary": "With skill: 83% assertions passed",
              "execution_metrics": {
                "tool_calls": 5,
                "token_counts": {
                  "input_tokens": 6,
                  "output_tokens": 2669,
                  "cache_read_input_tokens": 95702,
                  "cache_creation_input_tokens": 20713
                }
              },
              "timing": {
                "elapsed_seconds": 79.27190295899709
              },
              "raw_output": "I'll check AWS Health events for Lambda and related services in us-west-2 to correlate with your timeout issue.No events found for Lambda or its dependencies specifically. Let me broaden the search to all services in us-west-2 and also check account-wide events.## AWS Health Event Correlation Results\n\n### Summary\n\n| Total Events Found | Issues | Scheduled Changes | Account Notifications |\n|---|---|---|---|\n| 6 | 1 | 5 | 0 |\n\n---\n\n### Correlation Analysis\n\n**Incident context:**\n- **Service:** Lambda\n- **Region:** us-west-2\n- **Started:** ~2026-05-27 approx 3 hours ago\n- **Symptom:** Function timeouts\n- **Resource:** `arn:aws:lambda:us-west-2:123456789012:function:my-processor`\n\n---\n\n### Correlated Events\n\n#### Issues\n\n| Event | Relevance | Rationale |\n|---|---|---|\n| **AWS_ROUTE53_OPERATIONAL_ISSUE** (global, closed 2026-05-23 02:38\u201303:13 UTC) | **Low** | Route 53 is global and could affect Lambda DNS resolution, but this event closed ~4 days before the current incident \u2014 well outside the 2-hour residual impact window. |\n\n#### Scheduled Changes (Lambda-related)\n\n| Event | Relevance | Rationale |\n|---|---|---|\n| **AWS_LAMBDA_PLANNED_LIFECYCLE_EVENT** (us-east-1, starts 2026-10-31) | **Low** | Same service, but wrong region (us-east-1 vs us-west-2) and no timeframe overlap (5 months away). |\n| **AWS_LAMBDA_PLANNED_LIFECYCLE_EVENT** (us-east-1, starts 2026-08-31) | **Low** | Same service, but wrong region and no timeframe overlap (3 months away). |\n\n#### Discarded (no correlation)\n\n- **AWS_EKS_FARGATE_POD_EVICTIONS** \u2014 unrelated service, wrong region\n- **AWS_EKS_PLANNED_LIFECYCLE_EVENT** \u2014 unrelated service, wrong region, future event\n- **AWS_ELASTICMAPREDUCE_PLANNED_LIFECYCLE_EVENT** \u2014 unrelated service, wrong region\n\n---\n\n### Conclusion\n\n**No matching AWS Health events were found that correlate with your Lambda timeout incident in us-west-2.** There are no open or recently-closed events for Lambda (or its dependencies: VPC, CloudWatch) in us-west-2 that overlap with y"
            },
            "without_skill": {
              "eval_id": "relevance-correlation",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output classifies each event as High, Medium, or Low relevance",
                  "passed": false,
                  "evidence": "The output provides a generic classification framework/table with High/Medium/Low/None categories and examples, but does not classify any actual specific Health events. No real events were retrieved or classified.",
                  "method": "llm",
                  "confidence": 0.85,
                  "uncertain": false
                },
                {
                  "text": "The output explains the matching criteria used for each classification",
                  "passed": true,
                  "evidence": "The table includes a 'Criteria' column explaining what qualifies for each relevance level (e.g., High = directly affects Lambda/networking/downstream services), but these are generic criteria not applied to specific events.",
                  "method": "llm",
                  "confidence": 0.6,
                  "uncertain": false
                },
                {
                  "text": "The output orders results by relevance (High first, then Medium, then Low)",
                  "passed": true,
                  "evidence": "The classification table is ordered High \u2192 Medium \u2192 Low \u2192 None, presenting results by descending relevance.",
                  "method": "llm",
                  "confidence": 0.7,
                  "uncertain": false
                },
                {
                  "text": "The output labels open events with High or Medium relevance as likely contributing factors",
                  "passed": false,
                  "evidence": "No actual open events are identified or labeled as likely contributing factors. The output only provides a hypothetical framework since credentials were expired.",
                  "method": "llm",
                  "confidence": 0.9,
                  "uncertain": false
                },
                {
                  "text": "The output includes at least one actionable next step per correlated event",
                  "passed": false,
                  "evidence": "The output lists general 'Common causes' and asks user to refresh credentials, but does not provide actionable next steps per correlated event since no actual events were retrieved or correlated.",
                  "method": "llm",
                  "confidence": 0.85,
                  "uncertain": false
                },
                {
                  "text": "The output considers service match, timeframe overlap, and region when scoring",
                  "passed": true,
                  "evidence": "The framework considers service match (criteria column references Lambda, networking, downstream services), timeframe (query uses -3H filter), and region (filters on us-west-2, notes Health API is in us-east-1). However, these are in the proposed approach rather than applied to actual scoring.",
                  "method": "llm",
                  "confidence": 0.6,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.5,
              "summary": "Without skill: 50% assertions passed",
              "execution_metrics": {
                "tool_calls": 3,
                "token_counts": {
                  "input_tokens": 5,
                  "output_tokens": 1056,
                  "cache_read_input_tokens": 57584,
                  "cache_creation_input_tokens": 14196
                }
              },
              "timing": {
                "elapsed_seconds": 45.48013312499825
              },
              "raw_output": "I'll investigate the AWS Health events in us-west-2 to correlate them with your Lambda timeout issue. Let me start by checking available AWS credentials and querying the Health API.My AWS credentials are currently expired/invalid, so I can't directly query the AWS Health API. However, I can still help you investigate and classify Health events. Let me outline the approach:\n\n**To correlate AWS Health events with your Lambda timeout issue, here's what to check:**\n\n1. **Refresh your credentials** first:\n   ```bash\n   aws sso login --profile AdministratorAccess-123456789012-sso\n   ```\n\n2. **Then query Health events** (preferring ReadOnly credentials if available):\n   ```bash\n   # Get Health events in us-west-2 from the last 3 hours\n   aws health describe-events \\\n     --filter \"regions=us-west-2,startTimes=[{from=$(date -u -v-3H +%Y-%m-%dT%H:%M:%S)}],eventStatusCodes=open,upcoming,closed\" \\\n     --region us-east-1 \\\n     --profile <your-profile>\n   ```\n   (Note: AWS Health API is only available in `us-east-1` for global endpoint)\n\n3. **Classification framework** for correlating Health events with your Lambda timeouts:\n\n| Relevance | Criteria | Example |\n|---|---|---|\n| **High** | Directly affects Lambda, networking, or downstream services your function calls | Lambda operational issue, VPC/NAT degradation, DynamoDB/S3 latency in us-west-2 |\n| **Medium** | Affects infrastructure your function may depend on indirectly | EBS degradation (if using EFS), CloudWatch delays, SQS delivery delays |\n| **Low** | Same region but unrelated service | EC2 instance retirement, Redshift maintenance |\n| **None** | Different region or completely unrelated | Health events in other regions |\n\n**Common causes of Lambda timeouts that correlate with Health events:**\n- VPC/ENI provisioning delays (if your function is VPC-attached)\n- Downstream service latency (DynamoDB, RDS, S3, external APIs)\n- Cold start amplification during Lambda service degradation\n- NAT Gateway or networking issues\n\nWould"
            },
            "delta_pass_rate": 0.33333333333333337
          },
          {
            "eval_id": "chat-health-report",
            "run_index": 0,
            "with_skill": {
              "eval_id": "chat-health-report",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output organizes events by event type category (issues, scheduled changes, account notifications)",
                  "passed": true,
                  "evidence": "Output is explicitly organized with sections 'Issues (Service Disruptions)', 'Scheduled Changes', and 'Account Notifications' referenced in the overview table.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                },
                {
                  "text": "The output includes counts for each category",
                  "passed": true,
                  "evidence": "The overview table shows counts per category: Issues (3), Scheduled Changes (6), Account Notifications (0), with Total column.",
                  "method": "llm",
                  "confidence": 0.98,
                  "uncertain": false
                },
                {
                  "text": "The output includes a breakdown of events by AWS service",
                  "passed": false,
                  "evidence": "Events mention their AWS service (EC2, Route 53, EKS, Lambda, EMR, CloudShell) but there is no dedicated breakdown/summary table grouping or counting events by service. Events are organized by category type, not by service.",
                  "method": "llm",
                  "confidence": 0.72,
                  "uncertain": false
                },
                {
                  "text": "The output includes a breakdown by event status (open, closed, upcoming)",
                  "passed": true,
                  "evidence": "The overview table explicitly has columns for Open, Closed, and Upcoming counts per category, and each event is labeled with its status.",
                  "method": "llm",
                  "confidence": 0.98,
                  "uncertain": false
                },
                {
                  "text": "The output covers the past 30 days as the default time period",
                  "passed": true,
                  "evidence": "Report header states 'Report period: 2026-04-27 to 2026-05-27' and title says 'Past 30 Days'.",
                  "method": "llm",
                  "confidence": 0.99,
                  "uncertain": false
                },
                {
                  "text": "The output presents event timelines with start time and resolution time for closed events",
                  "passed": true,
                  "evidence": "Closed events show start and end times (e.g., '2026-05-08 00:25 UTC \u2192 2026-05-09 03:04 UTC (~27 hours)' for EC2 thermal event, and similar for other closed issues and scheduled changes).",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.8333333333333334,
              "summary": "With skill: 83% assertions passed",
              "execution_metrics": {
                "tool_calls": 3,
                "token_counts": {
                  "input_tokens": 6,
                  "output_tokens": 2861,
                  "cache_read_input_tokens": 100644,
                  "cache_creation_input_tokens": 32503
                }
              },
              "timing": {
                "elapsed_seconds": 73.23885583300216
              },
              "raw_output": "I'll retrieve AWS Health events for your account over the past 30 days using the Health API in us-east-1.Now let me get the event details to provide better descriptions.## AWS Health Events Summary Report \u2014 Past 30 Days\n\n**Account:** 123456789012 | **Report period:** 2026-04-27 to 2026-05-27 | **Total events:** 9\n\n---\n\n### Overview\n\n| Category | Open | Closed | Upcoming | Total |\n|----------|------|--------|----------|-------|\n| Issues (service disruptions) | 0 | 3 | 0 | **3** |\n| Scheduled Changes | 0 | 3 | 3 | **6** |\n| Account Notifications | 0 | 0 | 0 | **0** |\n\n| Scope | Count |\n|-------|-------|\n| Account-Specific | 6 |\n| Public | 3 |\n\n---\n\n### Issues (Service Disruptions)\n\n#### 1. EC2 \u2014 Thermal Event / Power Loss (us-east-1, use1-az4)\n- **Status:** Closed\n- **Duration:** 2026-05-08 00:25 UTC \u2192 2026-05-09 03:04 UTC (~27 hours)\n- **Scope:** PUBLIC\n- **Description:** A thermal event in a single data center caused loss of power, leading to impaired EC2 instances and degraded EBS volumes in use1-az4. Cascading impact to ELB, EKS, ElastiCache, Redshift, OpenSearch, Kafka, IoT Core, NAT Gateway, and SageMaker. Traffic was shifted away from the impacted AZ. Full recovery achieved after cooling systems were restored.\n- **Severity:** High \u2014 multi-service impact lasting over a day.\n\n#### 2. EC2 \u2014 Power Loss (eu-west-3, euw3-az2)\n- **Status:** Closed\n- **Duration:** 2026-04-27 11:27 UTC \u2192 2026-04-27 17:10 UTC (~6 hours)\n- **Scope:** PUBLIC\n- **Description:** Loss of power to infrastructure in euw3-az2 caused instance impairments and launch failures. Power restored by 4:20 AM PDT; full EC2/EBS recovery by 4:40 AM PDT. Numerous dependent services (50+) experienced temporary impact during the event.\n- **Severity:** Medium \u2014 rapid recovery, broad service impact window was short.\n\n#### 3. Route 53 \u2014 Increased API Error Rates (global)\n- **Status:** Closed\n- **Duration:** 2026-05-22 23:38 UTC \u2192 2026-05-23 00:13 UTC (~35 minutes)\n- **Scope:** PUBLIC\n- **Description:** Increased "
            },
            "without_skill": {
              "eval_id": "chat-health-report",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output organizes events by event type category (issues, scheduled changes, account notifications)",
                  "passed": false,
                  "evidence": "The output never retrieves or displays any AWS Health events. It gets stuck at the credential selection stage and asks the user which profile to use.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                },
                {
                  "text": "The output includes counts for each category",
                  "passed": false,
                  "evidence": "No event counts are presented anywhere in the output. The assistant never completes the query.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                },
                {
                  "text": "The output includes a breakdown of events by AWS service",
                  "passed": false,
                  "evidence": "No breakdown of events by AWS service is shown. The output only discusses credential/profile selection.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                },
                {
                  "text": "The output includes a breakdown by event status (open, closed, upcoming)",
                  "passed": false,
                  "evidence": "No event status breakdown (open, closed, upcoming) is presented. The query was never executed.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                },
                {
                  "text": "The output covers the past 30 days as the default time period",
                  "passed": false,
                  "evidence": "The output mentions intent to 'query AWS Health events for the past 30 days' but never actually retrieves or presents data for that time period.",
                  "method": "llm",
                  "confidence": 0.7,
                  "uncertain": false
                },
                {
                  "text": "The output presents event timelines with start time and resolution time for closed events",
                  "passed": false,
                  "evidence": "No event timelines, start times, or resolution times are shown anywhere in the output.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.0,
              "summary": "Without skill: 0% assertions passed",
              "execution_metrics": {
                "tool_calls": 4,
                "token_counts": {
                  "input_tokens": 6,
                  "output_tokens": 1143,
                  "cache_read_input_tokens": 82251,
                  "cache_creation_input_tokens": 14735
                }
              },
              "timing": {
                "elapsed_seconds": 44.08617000000231
              },
              "raw_output": "I'll check your AWS configuration and then query AWS Health events for the past 30 days.I see you have several profiles but only Admin roles are configured, and no active default credentials. I need to use one of your profiles to query AWS Health events. Since this is a read-only operation, I'd prefer a ReadOnly role, but I only see Admin roles in your config.I don't see any ReadOnly profiles configured \u2014 only Admin roles are available. Per best practices, I'd prefer using least-privilege credentials for this read-only query.I need to know which AWS profile to use. You only have Admin profiles configured \u2014 no ReadOnly roles are available. Which profile should I use to query AWS Health events? The operation itself is read-only (describe-events), but I want to confirm before using an Admin credential."
            },
            "delta_pass_rate": 0.8333333333333334
          },
          {
            "eval_id": "no-events-found",
            "run_index": 0,
            "with_skill": {
              "eval_id": "no-events-found",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output explicitly states that no matching AWS Health events were found",
                  "passed": true,
                  "evidence": "Output explicitly states 'No matching AWS Health events were found for ECS or related services'",
                  "method": "llm",
                  "confidence": 0.99,
                  "uncertain": false
                },
                {
                  "text": "The output confirms the search parameters used (service, time range, region)",
                  "passed": true,
                  "evidence": "Output lists search parameters including Services (ECS, EC2, VPC, ELB), Region (eu-west-1), Time range (2026-05-13 to present), Statuses, and Categories",
                  "method": "llm",
                  "confidence": 0.99,
                  "uncertain": false
                },
                {
                  "text": "The output suggests alternative investigation paths",
                  "passed": true,
                  "evidence": "Output provides 6 numbered 'Recommended next steps' covering deployments, resource limits, networking, application logs, target group health, and capacity provider",
                  "method": "llm",
                  "confidence": 0.99,
                  "uncertain": false
                },
                {
                  "text": "The output recommends checking recent deployments or configuration changes",
                  "passed": true,
                  "evidence": "First recommended next step explicitly says 'Recent deployments \u2014 check if a recent ECS task definition or service update correlates with the error spike'",
                  "method": "llm",
                  "confidence": 0.99,
                  "uncertain": false
                },
                {
                  "text": "The output does not fabricate or hallucinate health events that do not exist",
                  "passed": true,
                  "evidence": "Output clearly states no events were found and does not invent or fabricate any health events; it only suggests next investigation steps",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                }
              ],
              "pass_rate": 1.0,
              "summary": "With skill: 100% assertions passed",
              "execution_metrics": {
                "tool_calls": 3,
                "token_counts": {
                  "input_tokens": 5,
                  "output_tokens": 1166,
                  "cache_read_input_tokens": 66643,
                  "cache_creation_input_tokens": 18813
                }
              },
              "timing": {
                "elapsed_seconds": 45.409266541999386
              },
              "raw_output": "I'll check AWS Health events for ECS and related services in eu-west-1. Let me query the Health API.No events found in the past 7 days. Let me broaden the time window to 14 days.## Results: No AWS Health Events Found\n\nNo matching AWS Health events were found for ECS or related services (EC2, VPC, ELB) in **eu-west-1** over the past 14 days.\n\n**Search parameters used:**\n- Services: ECS, EC2, VPC, Elastic Load Balancing\n- Region: eu-west-1\n- Time range: 2026-05-13 to present\n- Statuses: open and closed\n- Categories: issues, scheduled changes, account notifications\n\n**Recommended next steps** \u2014 since this doesn't appear to be caused by an AWS-side service event:\n\n1. **Recent deployments** \u2014 check if a recent ECS task definition or service update correlates with the error spike\n2. **Resource limits** \u2014 verify ECS service quotas and task CPU/memory aren't being exhausted\n3. **Networking** \u2014 check security group rules, NACLs, and NAT gateway metrics for the VPC\n4. **Application logs** \u2014 inspect CloudWatch Logs for the ECS tasks to identify the error type (5xx, OOM kills, connection timeouts, etc.)\n5. **Target group health** \u2014 if behind an ALB, check target group health check failures\n6. **Capacity provider** \u2014 verify the capacity provider has sufficient capacity and isn't failing to place tasks\n\nWould you like me to help investigate any of these areas?"
            },
            "without_skill": {
              "eval_id": "no-events-found",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The output explicitly states that no matching AWS Health events were found",
                  "passed": false,
                  "evidence": "The output never states that no health events were found. Instead, it reports that API calls failed due to invalid/expired credentials and asks the user how to proceed.",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output confirms the search parameters used (service, time range, region)",
                  "passed": false,
                  "evidence": "The output mentions 'ECS in eu-west-1' in the opening line but does not confirm a time range. The search was never actually completed due to credential failures.",
                  "method": "llm",
                  "confidence": 0.7,
                  "uncertain": false
                },
                {
                  "text": "The output suggests alternative investigation paths",
                  "passed": false,
                  "evidence": "The output suggests ways to fix the credential problem (refresh credentials, use a different profile) but does not suggest alternative investigation paths for the original ECS health concern (e.g., checking CloudWatch, service logs, etc.).",
                  "method": "llm",
                  "confidence": 0.75,
                  "uncertain": false
                },
                {
                  "text": "The output recommends checking recent deployments or configuration changes",
                  "passed": false,
                  "evidence": "The output does not recommend checking recent deployments or configuration changes. It only discusses credential/authentication issues.",
                  "method": "llm",
                  "confidence": 0.95,
                  "uncertain": false
                },
                {
                  "text": "The output does not fabricate or hallucinate health events that do not exist",
                  "passed": true,
                  "evidence": "The output does not fabricate any health events. It honestly reports that it could not complete the query due to expired credentials.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                }
              ],
              "pass_rate": 0.2,
              "summary": "Without skill: 20% assertions passed",
              "execution_metrics": {
                "tool_calls": 3,
                "token_counts": {
                  "input_tokens": 5,
                  "output_tokens": 713,
                  "cache_read_input_tokens": 57372,
                  "cache_creation_input_tokens": 14388
                }
              },
              "timing": {
                "elapsed_seconds": 38.81947583299916
              },
              "raw_output": "I'll check AWS Health for any active events affecting ECS in eu-west-1.The AWS Health API calls are failing due to invalid security credentials. This could mean:\n\n1. **No active AWS credentials** \u2014 your session may have expired or no profile is configured.\n2. **Wrong profile** \u2014 you may need to specify a `--profile` flag.\n\nLet me check what's available:Your default credentials are expired. You have several profiles available. To check AWS Health events, I'd recommend using a **ReadOnly** profile if one exists, but I only see Admin profiles here.\n\nWhich profile should I use to query AWS Health? For example:\n- `udid-Admin` (appears to be your primary account)\n- One of the linked account profiles\n\nAlso, you may need to refresh your credentials first (e.g., via `ada credentials update --profile <profile-name>` or re-authenticating via SSO). Would you like me to try with a specific profile, or would you prefer to refresh your credentials first?"
            },
            "delta_pass_rate": 0.8
          },
          {
            "eval_id": "api-region-awareness",
            "run_index": 0,
            "with_skill": {
              "eval_id": "api-region-awareness",
              "run_index": 0,
              "assertion_results": [
                {
                  "text": "The Health API call is made to the us-east-1 endpoint",
                  "passed": true,
                  "evidence": "The output shows successful Health API queries returning results (even if empty), which implies correct endpoint usage (us-east-1 is required for AWS Health API). However, the output does not explicitly state which endpoint was used.",
                  "method": "llm",
                  "confidence": 0.4,
                  "uncertain": true
                },
                {
                  "text": "The output does not attempt to call the Health API in ap-southeast-1",
                  "passed": true,
                  "evidence": "The output uses ap-southeast-1 as a filter parameter for events, not as the API endpoint. The successful execution without errors suggests the API was not called in ap-southeast-1 directly. However, the output doesn't explicitly show the endpoint used.",
                  "method": "llm",
                  "confidence": 0.5,
                  "uncertain": false
                },
                {
                  "text": "The output filters events by the ap-southeast-1 region where resources are located",
                  "passed": true,
                  "evidence": "The summary table explicitly shows 'Region | ap-southeast-1' as a filter, and the output states it checked for 'open issues across all services in ap-southeast-1' and 'DynamoDB events globally'.",
                  "method": "llm",
                  "confidence": 0.97,
                  "uncertain": false
                },
                {
                  "text": "The output returns relevant health event information regardless of the resource region",
                  "passed": true,
                  "evidence": "The output successfully returned health event information (no events found) and also checked 'DynamoDB events globally' beyond just ap-southeast-1, providing a complete answer with next steps regardless of resource region.",
                  "method": "llm",
                  "confidence": 0.85,
                  "uncertain": false
                },
                {
                  "text": "The output does not fail due to region misconfiguration",
                  "passed": true,
                  "evidence": "The output completed successfully with a full summary, searc
    • evals.json 8.2 KB
      [
        {
          "id": "incident-health-event-lookup",
          "prompt": "I'm investigating an incident with our RDS database in us-west-2. The database started returning connection timeout errors about 2 hours ago. Check if there are any AWS Health events that could explain this.",
          "expected_output": "The skill searches AWS Health events using DescribeEvents filtered by the RDS service and us-west-2 region within the past 7 days. It retrieves both account-specific and public events, then presents any matching events with their status, timeline, and description, or states that no matching events were found.",
          "assertions": [
            "The output calls DescribeEvents with the RDS service filter",
            "The output filters events by the us-west-2 region",
            "The output searches within the past 7 days from the incident start time",
            "The output includes both account-specific and public event scopes",
            "The output presents matching events or clearly states no events were found",
            "The output does not produce an API error or permission denied message"
          ]
        },
        {
          "id": "event-detail-extraction",
          "prompt": "I found AWS Health events related to EC2 in us-east-1. Get me the full details for these events including descriptions, timelines, and current status.",
          "expected_output": "The skill calls DescribeEventDetails with the identified event ARNs, batching up to 10 per request. For each event, it presents the event description, start time, end time, last updated time, status, service, region, and event type category.",
          "assertions": [
            "The output calls DescribeEventDetails to retrieve full event information",
            "The output includes the event description text for each event",
            "The output includes start time and end time in ISO 8601 format",
            "The output includes the current status (open, closed, or upcoming) for each event",
            "The output batches requests with no more than 10 event ARNs per API call",
            "The output handles any failedSet entries by reporting them and continuing with successfulSet"
          ]
        },
        {
          "id": "affected-entity-matching",
          "prompt": "We have an account-specific Health event affecting EC2. Our incident involves instances i-0abc123def456 and i-0def789ghi012 in us-east-1. Identify which of our resources are affected by this event.",
          "expected_output": "The skill calls DescribeAffectedEntities for the account-specific event and matches the returned entities against the provided instance IDs. Matched resources are presented separately with their status (IMPAIRED, UNIMPAIRED, UNKNOWN, or PENDING) and last updated time.",
          "assertions": [
            "The output calls DescribeAffectedEntities for the account-specific event",
            "The output performs exact string matching of entity values against the incident resource identifiers",
            "The output presents matched entities separately from non-matched entities",
            "The output includes entity status (IMPAIRED, UNIMPAIRED, UNKNOWN, or PENDING) for each entity",
            "The output includes last updated time for each affected entity"
          ]
        },
        {
          "id": "relevance-correlation",
          "prompt": "I'm investigating a Lambda function timeout issue in us-west-2 that started 3 hours ago. The function ARN is arn:aws:lambda:us-west-2:123456789012:function:my-processor. Several Health events were found. Correlate them with my incident and classify their relevance.",
          "expected_output": "The skill scores each Health event for relevance by evaluating service match, timeframe overlap, region/AZ match, and affected resource match. Events are classified as High, Medium, or Low relevance. Open events with High or Medium relevance are labeled as likely contributing factors. Results are ordered by relevance classification then by start time descending.",
          "assertions": [
            "The output classifies each event as High, Medium, or Low relevance",
            "The output explains the matching criteria used for each classification",
            "The output orders results by relevance (High first, then Medium, then Low)",
            "The output labels open events with High or Medium relevance as likely contributing factors",
            "The output includes at least one actionable next step per correlated event",
            "The output considers service match, timeframe overlap, and region when scoring"
          ]
        },
        {
          "id": "chat-health-report",
          "prompt": "Generate a summary report of all AWS Health events for my account over the past 30 days.",
          "expected_output": "The skill retrieves all Health events within the past 30 days and presents a structured report organized by event type category (issues, scheduled changes, account notifications) with counts for each. The report includes a breakdown by AWS service and by event status (open, closed, upcoming), with event timelines showing start, resolution, and duration.",
          "assertions": [
            "The output organizes events by event type category (issues, scheduled changes, account notifications)",
            "The output includes counts for each category",
            "The output includes a breakdown of events by AWS service",
            "The output includes a breakdown by event status (open, closed, upcoming)",
            "The output covers the past 30 days as the default time period",
            "The output presents event timelines with start time and resolution time for closed events"
          ]
        },
        {
          "id": "no-events-found",
          "prompt": "I'm investigating elevated error rates on our ECS cluster in eu-west-1. Check if there are any AWS Health events that could be causing this.",
          "expected_output": "The skill searches for Health events filtered by ECS service in eu-west-1 and finds no matching events. It explicitly states no events were found, confirms the search parameters used, and suggests alternative investigation paths such as checking recent deployments, configuration changes, resource limits, or network connectivity.",
          "assertions": [
            "The output explicitly states that no matching AWS Health events were found",
            "The output confirms the search parameters used (service, time range, region)",
            "The output suggests alternative investigation paths",
            "The output recommends checking recent deployments or configuration changes",
            "The output does not fabricate or hallucinate health events that do not exist"
          ]
        },
        {
          "id": "api-region-awareness",
          "prompt": "Check AWS Health events for our DynamoDB tables experiencing throttling in ap-southeast-1.",
          "expected_output": "The skill calls the AWS Health API using the us-east-1 endpoint regardless of the fact that the affected DynamoDB resources are in ap-southeast-1. The events filter includes ap-southeast-1 as the region filter to scope results to that region, but the API call itself targets us-east-1.",
          "assertions": [
            "The Health API call is made to the us-east-1 endpoint",
            "The output does not attempt to call the Health API in ap-southeast-1",
            "The output filters events by the ap-southeast-1 region where resources are located",
            "The output returns relevant health event information regardless of the resource region",
            "The output does not fail due to region misconfiguration"
          ]
        },
        {
          "id": "broadened-search",
          "prompt": "I'm investigating connection failures to our Application Load Balancer in us-east-1. The ALB started returning 503 errors 30 minutes ago. Check for any related AWS Health events.",
          "expected_output": "The skill first searches for Health events filtered by the ELB service in us-east-1 within the past 7 days. Finding no results, it broadens the search to related services (EC2, VPC, Route 53) that share infrastructure dependencies with ELB. If still no results, it expands the time window to 14 days before reporting no events found.",
          "assertions": [
            "The output first searches for events specific to the ELB/ELASTICLOADBALANCING service",
            "The output broadens to related services (EC2, VPC, Route 53) when no ELB events are found",
            "The output expands the time window to 14 days if related service search also returns no results",
            "The output follows the decision tree: specific service → related services → expanded time window",
            "The output reports the search progression to the operator"
          ]
        }
      ]
      
    • eval_queries.json 831 B
      [
        {
          "query": "How do I set up a multi-AZ deployment for my RDS instance?",
          "should_trigger": false
        },
        {
          "query": "Write a Python script to list all S3 buckets in my account",
          "should_trigger": false
        },
        {
          "query": "What is the difference between an IAM role and an IAM user?",
          "should_trigger": false
        },
        {
          "query": "How can I reduce my monthly AWS bill for EC2 instances?",
          "should_trigger": false
        },
        {
          "query": "What are the best practices for designing a VPC with public and private subnets?",
          "should_trigger": false
        },
        {
          "query": "How do I configure CloudWatch alarms for Lambda function errors?",
          "should_trigger": false
        },
        {
          "query": "Explain how to attach an IAM policy that allows read-only access to DynamoDB",
          "should_trigger": false
        }
      ]
      
    • report.json 1.6 KB
      {
        "skill_name": "aws-health-events",
        "skill_path": "/Users/udid/AWS Work/Services/DevOps Agent/aws-devops-agent-skill-central/skills/aws-health-events",
        "timestamp": "2026-05-27T13:34:33Z",
        "overall_score": 0.9147,
        "overall_grade": "A",
        "passed": true,
        "sections": {
          "audit": {
            "score": 100,
            "grade": "A",
            "passed": true,
            "normalized": 1.0,
            "critical": 0,
            "warning": 0,
            "info": 0
          },
          "functional": {
            "overall": 0.8067,
            "grade": "B",
            "passed": true,
            "scores": {
              "outcome": 0.6792,
              "process": 0.8684,
              "style": 0.6792,
              "efficiency": 1.0,
              "overall": 0.8067
            },
            "cost_efficiency": {
              "quality_delta": 0.3792,
              "cost_delta_pct": 33.8,
              "classification": "TRADEOFF",
              "emoji": "\ud83d\udfe1",
              "description": "Skill improves quality but increases cost"
            },
            "estimated_cost": {
              "with_skill_per_run": {
                "input_cost": 1.7e-05,
                "output_cost": 0.026423,
                "total_cost": 0.02644,
                "model": "sonnet",
                "currency": "USD"
              },
              "without_skill_per_run": {
                "input_cost": 1.8e-05,
                "output_cost": 0.019725,
                "total_cost": 0.019743,
                "model": "sonnet",
                "currency": "USD"
              },
              "per_eval_pair": 0.046183,
              "total_runs": 8,
              "total_cost": 0.3695,
              "model": "sonnet",
              "currency": "USD"
            }
          },
          "trigger": {
            "pass_rate": 1.0,
            "grade": "A",
            "passed": true,
            "total_queries": 7
          }
        }
      }
    • trigger_report.json 2.9 KB
      {
        "skill_name": "aws-health-events",
        "skill_path": "/Users/udid/AWS Work/Services/DevOps Agent/aws-devops-agent-skill-central/skills/aws-health-events",
        "query_results": [
          {
            "query": "How do I set up a multi-AZ deployment for my RDS instance?",
            "should_trigger": false,
            "trigger_count": 0,
            "run_count": 1,
            "trigger_rate": 0.0,
            "passed": true,
            "mean_input_tokens": 3.0,
            "mean_output_tokens": 438.0,
            "mean_total_tokens": 441.0
          },
          {
            "query": "Write a Python script to list all S3 buckets in my account",
            "should_trigger": false,
            "trigger_count": 0,
            "run_count": 1,
            "trigger_rate": 0.0,
            "passed": true,
            "mean_input_tokens": 3.0,
            "mean_output_tokens": 181.0,
            "mean_total_tokens": 184.0
          },
          {
            "query": "What is the difference between an IAM role and an IAM user?",
            "should_trigger": false,
            "trigger_count": 0,
            "run_count": 1,
            "trigger_rate": 0.0,
            "passed": true,
            "mean_input_tokens": 3.0,
            "mean_output_tokens": 299.0,
            "mean_total_tokens": 302.0
          },
          {
            "query": "How can I reduce my monthly AWS bill for EC2 instances?",
            "should_trigger": false,
            "trigger_count": 0,
            "run_count": 1,
            "trigger_rate": 0.0,
            "passed": true,
            "mean_input_tokens": 3.0,
            "mean_output_tokens": 518.0,
            "mean_total_tokens": 521.0
          },
          {
            "query": "What are the best practices for designing a VPC with public and private subnets?",
            "should_trigger": false,
            "trigger_count": 0,
            "run_count": 1,
            "trigger_rate": 0.0,
            "passed": true,
            "mean_input_tokens": 3.0,
            "mean_output_tokens": 601.0,
            "mean_total_tokens": 604.0
          },
          {
            "query": "How do I configure CloudWatch alarms for Lambda function errors?",
            "should_trigger": false,
            "trigger_count": 0,
            "run_count": 1,
            "trigger_rate": 0.0,
            "passed": true,
            "mean_input_tokens": 3.0,
            "mean_output_tokens": 792.0,
            "mean_total_tokens": 795.0
          },
          {
            "query": "Explain how to attach an IAM policy that allows read-only access to DynamoDB",
            "should_trigger": false,
            "trigger_count": 0,
            "run_count": 1,
            "trigger_rate": 0.0,
            "passed": true,
            "mean_input_tokens": 3.0,
            "mean_output_tokens": 497.0,
            "mean_total_tokens": 500.0
          }
        ],
        "summary": {
          "total_queries": 7,
          "passed": 7,
          "failed": 0,
          "trigger_precision": 1.0,
          "no_trigger_precision": 1.0,
          "mean_total_tokens_per_run": 478.1,
          "estimated_cost": {
            "per_run": {
              "input_cost": 9e-06,
              "output_cost": 0.007127,
              "total_cost": 0.007136,
              "model": "sonnet",
              "currency": "USD"
            },
            "total_runs": 7,
            "total_cost": 0.05,
            "model": "sonnet",
            "currency": "USD"
          }
        },
        "passed": true
      }
  • references
    • health-api-reference.md 7.3 KB
      # AWS Health API Reference
      
      Quick reference for the AWS Health API operations used by this skill.
      
      ## DescribeEvents
      
      Returns a list of Health events matching specified filter criteria.
      
      ### Key Parameters
      
      | Parameter | Type | Description | Constraints |
      |-----------|------|-------------|-------------|
      | `filter.services` | List of strings | AWS service codes to filter by (e.g., `EC2`, `RDS`) | Max 10 values |
      | `filter.regions` | List of strings | AWS regions to filter by | Max 10 values |
      | `filter.availabilityZones` | List of strings | Availability zones to filter by | Max 10 values |
      | `filter.startTimes` | List of DateTimeRange | Filter by event start time (from/to) | ISO 8601 timestamps |
      | `filter.endTimes` | List of DateTimeRange | Filter by event end time (from/to) | ISO 8601 timestamps |
      | `filter.lastUpdatedTimes` | List of DateTimeRange | Filter by last updated time | ISO 8601 timestamps |
      | `filter.eventTypeCategories` | List of strings | Event type categories to filter by | `issue`, `scheduledChange`, `accountNotification` |
      | `filter.eventStatusCodes` | List of strings | Event statuses to filter by | `open`, `closed`, `upcoming` |
      | `filter.eventTypeCodes` | List of strings | Specific event type codes | e.g., `AWS_EC2_OPERATIONAL_ISSUE` |
      | `filter.entityValues` | List of strings | Resource ARNs or IDs to filter by | Max 100 values |
      | `filter.eventArns` | List of strings | Specific event ARNs to retrieve | Max 10 values |
      | `maxResults` | Integer | Max events per page | 10–100 (default 10) |
      | `nextToken` | String | Pagination token | From previous response |
      
      ### Response Fields
      
      | Field | Description |
      |-------|-------------|
      | `events[].arn` | Unique event identifier (ARN) |
      | `events[].service` | AWS service namespace (e.g., `EC2`, `RDS`) |
      | `events[].eventTypeCode` | Specific event type (e.g., `AWS_EC2_OPERATIONAL_ISSUE`) |
      | `events[].eventTypeCategory` | Category: `issue`, `scheduledChange`, or `accountNotification` |
      | `events[].region` | AWS region where the event occurred |
      | `events[].availabilityZone` | Specific AZ if applicable |
      | `events[].startTime` | When the event started (ISO 8601) |
      | `events[].endTime` | When the event ended (ISO 8601); null if ongoing |
      | `events[].lastUpdatedTime` | Last update timestamp (ISO 8601) |
      | `events[].statusCode` | Current status: `open`, `closed`, or `upcoming` |
      | `events[].eventScopeCode` | Scope: `ACCOUNT_SPECIFIC` or `PUBLIC` |
      | `nextToken` | Pagination token for next page (null if no more results) |
      
      ---
      
      ## DescribeEventDetails
      
      Returns detailed information for one or more Health events, including full descriptions.
      
      ### Key Parameters
      
      | Parameter | Type | Description | Constraints |
      |-----------|------|-------------|-------------|
      | `eventArns` | List of strings (required) | Event ARNs to retrieve details for | **Max 10 per request** |
      | `locale` | String | Language for event descriptions | e.g., `en` (default) |
      
      ### Response Fields
      
      | Field | Description |
      |-------|-------------|
      | `successfulSet[].event` | The event object (same fields as DescribeEvents response) |
      | `successfulSet[].eventDescription.latestDescription` | Full text description of the event |
      | `successfulSet[].eventMetadata` | Additional metadata key-value pairs |
      | `failedSet[].eventArn` | ARN of the event that failed to retrieve |
      | `failedSet[].errorName` | Error code for the failure |
      | `failedSet[].errorMessage` | Human-readable error message |
      
      ---
      
      ## DescribeAffectedEntities
      
      Returns a list of resources (entities) affected by a specific Health event.
      
      ### Key Parameters
      
      | Parameter | Type | Description | Constraints |
      |-----------|------|-------------|-------------|
      | `filter.eventArns` | List of strings (required) | Event ARNs to get affected entities for | Max 10 values |
      | `filter.entityValues` | List of strings | Filter by specific resource ARNs or IDs | Max 100 values |
      | `filter.entityArns` | List of strings | Filter by entity ARNs | Max 100 values |
      | `filter.lastUpdatedTimes` | List of DateTimeRange | Filter by entity last updated time | ISO 8601 timestamps |
      | `filter.statusCodes` | List of strings | Filter by entity status | `IMPAIRED`, `UNIMPAIRED`, `UNKNOWN`, `PENDING` |
      | `maxResults` | Integer | Max entities per page | 10–100 (default 10) |
      | `nextToken` | String | Pagination token | From previous response |
      
      ### Response Fields
      
      | Field | Description |
      |-------|-------------|
      | `entities[].entityValue` | Resource ARN or ID |
      | `entities[].eventArn` | Associated event ARN |
      | `entities[].awsAccountId` | Account owning the resource |
      | `entities[].lastUpdatedTime` | Last status update (ISO 8601) |
      | `entities[].statusCode` | Entity status: `IMPAIRED`, `UNIMPAIRED`, `UNKNOWN`, or `PENDING` |
      | `entities[].tags` | Resource tags (key-value map) |
      | `nextToken` | Pagination token for next page (null if no more results) |
      
      ---
      
      ## Common Service Codes
      
      | Service | Health API Code |
      |---------|----------------|
      | EC2 | `EC2` |
      | RDS | `RDS` |
      | Lambda | `LAMBDA` |
      | ELB | `ELASTICLOADBALANCING` |
      | S3 | `S3` |
      | CloudFront | `CLOUDFRONT` |
      | DynamoDB | `DYNAMODB` |
      | ECS | `ECS` |
      | EKS | `EKS` |
      | VPC | `VPC` |
      | Route 53 | `ROUTE53` |
      | IAM | `IAM` |
      | CloudWatch | `CLOUDWATCH` |
      | SNS | `SNS` |
      | SQS | `SQS` |
      | EBS | `EBS` |
      | ElastiCache | `ELASTICACHE` |
      | Kinesis | `KINESIS` |
      | API Gateway | `APIGATEWAY` |
      | SageMaker | `SAGEMAKER` |
      
      ---
      
      ## Event Type Categories
      
      | Category | Description |
      |----------|-------------|
      | `issue` | An AWS service issue or outage affecting resources |
      | `scheduledChange` | Planned maintenance or infrastructure change |
      | `accountNotification` | Account-specific notification (e.g., certificate expiry, abuse report) |
      
      ---
      
      ## Event Statuses
      
      | Status | Description |
      |--------|-------------|
      | `open` | Event is currently active and ongoing |
      | `closed` | Event has been resolved |
      | `upcoming` | Scheduled event that has not yet started |
      
      ---
      
      ## Entity Status Codes
      
      | Status | Description |
      |--------|-------------|
      | `IMPAIRED` | Resource is confirmed impaired by the event |
      | `UNIMPAIRED` | Resource is confirmed not impaired |
      | `UNKNOWN` | Impact status cannot be determined |
      | `PENDING` | Impact assessment is in progress |
      
      ---
      
      ## Important Constraints
      
      - **API endpoint**: The AWS Health API is available only in the `us-east-1`
        region. All API calls must target this region regardless of where the affected
        resources are located.
      - **Batch size**: `DescribeEventDetails` accepts a maximum of **10 event ARNs**
        per request. For more than 10 events, issue multiple batched calls.
      - **Pagination**: All list operations are paginated. Use `nextToken` from the
        response to retrieve subsequent pages until `nextToken` is null.
      - **Rate limiting**: Standard AWS API rate limits apply. Use exponential backoff
        for retries on throttling errors (HTTP 429).
      - **Event scope**: Use `eventScopeCode` to distinguish between account-specific
        events (`ACCOUNT_SPECIFIC`) and public service events (`PUBLIC`).
      - **Data availability**: Health events are available for up to 90 days. Events
        older than 90 days cannot be retrieved via the API.
      - **IAM permissions required**: `health:DescribeEvents`,
        `health:DescribeEventDetails`, `health:DescribeAffectedEntities`,
        `health:DescribeEventTypes`.
      - **Affected entities**: Only events with `eventScopeCode=ACCOUNT_SPECIFIC` have
        affected entities. Public events do not return entity data.
      
  • .skilleval.yaml 77 B
    audit:
      ignore:
        - STR-016    # README alongside SKILL.md is intentional
    
  • CHANGELOG.md 41 B
    # Changelog
    
    ## 1.0.0
    
    - Initial version
    
  • README.md 5.3 KB
    # AWS Health Events Skill
    
    This skill enables the AWS DevOps Agent to retrieve and analyze AWS Health events during incident investigation, root cause analysis, and operational troubleshooting.
    
    ## Purpose
    
    When an operational incident occurs, AWS Health events provide critical visibility into service disruptions, scheduled maintenance, and account-specific notifications that may explain or correlate with observed issues. This skill searches AWS Health events by service, time window, severity, region, and event type to surface AWS-side events as potential root causes or contributing factors.
    
    ## Key Capabilities
    
    - Retrieve open and resolved AWS Health events via the Health API
    - Search and filter events by service, severity, time range, region, and event type
    - Identify affected resources impacted by Health events
    - Correlate Health events with current incidents to determine root cause or contributing factors
    - Generate health posture reports summarizing account health over configurable time periods
    
    ## Prerequisites
    
    - IAM permissions for `health:Describe*` actions (`health:DescribeEvents`, `health:DescribeEventDetails`, `health:DescribeAffectedEntities`, `health:DescribeEventTypes`)
    - AWS Health API access (requires AWS Business Support+, Enterprise Support, or Unified Operations support plan)
    
    ## Limitations
    
    - The AWS Health API is only available in the `us-east-1` region for commercial accounts
    - Health event data retention is limited by the AWS Health service retention period
    - API rate limits apply — the skill implements exponential backoff for throttled requests
    
    ## Agent Types
    
    This skill is used by the following agent types:
    
    - **Chat tasks** — conversational health event lookup, reporting, and analysis
    - **Incident RCA** — root cause analysis during active incidents
    
    ## Uploading to AWS DevOps Agent
    
    You can add this skill to your Agent Space in three ways:
    
    **Option A: Import from GitHub (recommended)**
    
    If you have a [GitHub connection configured](https://docs.aws.amazon.com/devopsagent/latest/userguide/connecting-to-cicd-pipelines-connecting-github.html) in your Agent Space, you can import this skill directly from the repository. In the DevOps Agent web app, go to Settings → Add Skill → Import from repository, then point to the `skills/aws-health-events` directory. When prompted, select the agent types: **Chat tasks** and **Incident RCA**. See [Importing a skill from a repository](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-devops-agent-skills.html#creating-skills) for full instructions.
    
    > **Note:** You cannot connect the `aws` GitHub organization directly because the GitHub connection setup requires admin rights on the organization. Instead, connect your personal GitHub account and select any repository from it during the connection setup. Once a GitHub connection is established, you can import skills from any public repository — including this one — even if it wasn't selected during the connection setup.
    
    **Option B: Upload as a zip file**
    
    1. Zip the `aws-health-events/` directory (only including allowed extensions):
    
       ```bash
       cd skills
       zip -r aws-health-events.zip aws-health-events/ -i '*.md' '*.txt' '*.json' '*.yaml' '*.yml' '*.xml' '*.csv' '*.tsv' '*.html' '*.htm' '*.png' '*.jpg' '*.jpeg' '*.gif' '*.svg' '*.webp' '*.pdf' -x '*/.claude/*' '*/scripts/*' '*/README.md' '*/.skilleval.yaml' '*/.skilleval.yml' '*/CHANGELOG.md' '*/evals/*'
       ```
    
    2. In the AWS DevOps Agent web app, navigate to the **Skills** page.
    3. Click **Add skill** → **Upload skill**.
    4. Drag and drop the `aws-health-events.zip` file (max 6 MB).
    5. Select the agent types: **Chat tasks** and **Incident RCA**.
    6. Click **Upload**.
    
    **Option C: Upload via the Asset API**
    
    Use the AWS DevOps Agent Asset API to programmatically manage skills — useful for CI/CD pipelines or automation workflows. See [Managing a skill end-to-end](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-managing-assets.html#managing-a-skill-end-to-end) for the full API workflow.
    
    For more details, see [Uploading a skill](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-devops-agent-skills.html#creating-skills) in the AWS DevOps Agent User Guide.
    
    ## How to Use This Skill
    
    This skill is most suitable for chat and investigation. Below are sample prompts for each use-case.
    
    ### Chat
    
    - "Create a health events report for the last 30 days, including breakdowns by service and event type."
    - "Show me all open AWS Health events affecting my account."
    - "Are there any scheduled maintenance events for RDS in the next 7 days?"
    - "Summarize all EC2 health events from the past quarter."
    - "What health events have occurred in us-west-2 this month?"
    - "Give me a health posture summary for my account over the last 90 days."
    
    ### Investigation
    
    - "Investigate the RDS connectivity alarm — are there any related AWS Health events?"
    - "We're seeing elevated error rates on our ECS service. Check if there's an AWS-side issue."
    - "Our Lambda functions are timing out. Is there an active Health event for Lambda in us-east-1?"
    - "There's a latency spike on our ALB starting around 2am UTC — correlate with any Health events."
    - "Multiple services in us-west-2 are degraded. Check for regional Health events."
    - "The same EBS performance issue keeps recurring. Are there related Health events or scheduled changes?"
    
  • SKILL.md 18.1 KB
    ---
    name: aws-health-events
    description: ALWAYS use this skill in the beginning of any incident investigation, root cause
      analysis, or operational troubleshooting. This skill retrieves and analyzes
      AWS Health events (service issues, scheduled changes, and account notifications)
      to identify AWS-side events that may explain or correlate with observed
      operational issues. Activate this skill when investigating an issue and you
      observe service degradation, elevated error rates, latency spikes, connection
      failures, throttling, capacity issues, deployment-related failures, alarms, or
      any operational event or issue. This skill searches AWS Health events by
      service, time window, region, and status to surface active or recent service
      disruptions, scheduled maintenance, and account-specific notifications that
      inform the current investigation. Also activate when a user requests a health
      event summary or report for their account over a specified time period.
    metadata:
      author: udid-aws
      version: "1.0.0"
      aws-devops-agent-skills.agent-types: "Chat tasks, Incident RCA"
      aws-devops-agent-skills.aws-services: "AWS Health"
      aws-devops-agent-skills.technical-domains: "Operations"
    ---
    
    # AWS Health Event Review
    
    Use this skill when investigating an incident and you need to check for AWS-side
    service events that may be causing or contributing to the observed issue. Also
    use this skill when a user requests a summary report of AWS Health events over a
    configurable time period.
    
    ## When to Use This Skill
    
    **Incident Investigation (automatic activation):**
    
    - An active incident may be caused by an AWS service disruption or degradation.
    - You observe service degradation, elevated error rates, latency spikes,
      connection failures, throttling, or capacity issues.
    - You need to determine whether an AWS-side event is the root cause or a
      contributing factor to the current incident.
    - You want to correlate observed symptoms with known AWS Health events.
    
    **Chat Reporting (on-demand activation):**
    
    - A user requests a health event summary or report for their account.
    - A user wants to review the health posture of their AWS environment over a
      specific time period.
    - A user asks about recent AWS service issues affecting their account or region.
    
    ## Prerequisites
    
    - The account must have an **AWS Business Support+**, **Enterprise Support**, or
      **Unified Operations** support plan to access the AWS Health API.
    - The agent must have permissions to call the following IAM actions:
      - `health:DescribeEvents`
      - `health:DescribeEventDetails`
      - `health:DescribeAffectedEntities`
      - `health:DescribeEventTypes`
    - The AWS Health API is only available in the **us-east-1** region. All API
      calls must target the `us-east-1` endpoint regardless of where the affected
      resources are located.
    - Health event data is available for up to 90 days. Events older than 90 days
      cannot be retrieved via the API.
    
    ---
    
    ## Step 1: Gather Incident Context
    
    Before searching Health events, extract key details from the current incident:
    
    1. **Affected AWS services** — identify the service(s) experiencing issues
       (e.g., EC2, RDS, Lambda, ELB, ECS).
    2. **Timeframe** — determine when the incident started and its current duration.
       Use ISO 8601 timestamps.
    3. **Affected resources** — collect specific resource identifiers (instance IDs,
       ARNs, endpoint names, cluster names).
    4. **Region and availability zone** — identify the AWS region and, if known, the
       specific availability zone(s) affected.
    5. **Symptoms** — note the observed symptoms (latency spikes, 5xx errors,
       connection timeouts, throttling, capacity errors).
    
    Use these details as filter criteria in subsequent steps.
    
    ---
    
    ## Step 2: Search Health Events
    
    Use the AWS Health API `DescribeEvents` operation to retrieve events matching
    the incident context. All calls must target the **us-east-1** endpoint.
    
    ### API call pattern
    
    ```
    aws health describe-events \
      --region us-east-1 \
      --filter '{
        "services": ["<SERVICE_CODE>"],
        "startTimes": [{"from": "<ISO-8601-start>"}],
        "regions": ["<affected-region>"],
        "eventStatusCodes": ["open", "closed"],
        "eventTypeCategories": ["issue", "scheduledChange", "accountNotification"]
      }' \
      --max-results 100
    ```
    
    ### Filtering strategies
    
    | Strategy | How to Apply |
    |----------|-------------|
    | By service | Use the `services` filter with the AWS Health service code (e.g., `EC2`, `RDS`, `ELASTICLOADBALANCING`), because service-specific events are most likely to correlate with the incident. |
    | By time range | Use `startTimes` with a `from` value set to 7 days before the incident start, because events that started before the incident may still be active and causing impact. |
    | By region | Use the `regions` filter to scope events to the affected region, because regional events are more likely to impact the specific resources under investigation. |
    | By availability zone | Use the `availabilityZones` filter when the incident is isolated to a specific AZ, because AZ-scoped events have the highest correlation with AZ-specific failures. |
    | By status | Include both `open` and `closed` statuses, because recently closed events may have caused residual impact that is still being observed. |
    | By event scope | Include both `ACCOUNT_SPECIFIC` and `PUBLIC` events, because public service events affect all accounts in the region while account-specific events target your resources directly. |
    
    ### Pagination handling
    
    - Follow the `nextToken` from each response to retrieve subsequent pages.
    - Continue paginating until `nextToken` is null or a maximum of **500 events**
      have been collected.
    - Set `maxResults` to 100 per page for efficient retrieval.
    
    ---
    
    ## Step 3: Filter Relevant Events
    
    Before retrieving full event details, filter the events returned in Step 2 to
    identify only those relevant to the current investigation. This avoids
    unnecessary `DescribeEventDetails` calls for events that are clearly unrelated.
    
    ### Relevance filtering criteria
    
    Evaluate each event from the `DescribeEvents` response using these fields
    (available without calling `DescribeEventDetails`):
    
    | Field | Relevance Signal |
    |-------|-----------------|
    | `service` | Must match one of the affected services from the incident context, or a related service from the Service Dependency Map |
    | `eventTypeCategory` | Prioritize `issue` events for active incidents; include `scheduledChange` if the incident coincides with a maintenance window |
    | `eventTypeCode` | Match against known operational event patterns (e.g., `AWS_EC2_OPERATIONAL_ISSUE`, `AWS_RDS_MAINTENANCE`) |
    | `statusCode` | Prioritize `open` events; include `closed` only if the event ended within 2 hours of the incident start |
    | `startTime` / `endTime` | The event's active period must overlap with the incident timeframe |
    | `region` / `availabilityZone` | Must match the incident's affected region or AZ |
    
    ### Filtering rules
    
    1. **Keep** events where the `service` matches an affected service or a related
       service from the Service Dependency Map.
    2. **Keep** events where the active period (startTime to endTime, or to present
       if open) overlaps with the incident timeframe.
    3. **Keep** events where the `region` or `availabilityZone` matches the
       incident's affected region/AZ.
    4. **Discard** `accountNotification` events unless the incident context
       specifically suggests an account-level issue (e.g., abuse notification,
       certificate expiry).
    5. **Discard** `closed` events that ended more than 2 hours before the incident
       started (unlikely to be contributing).
    
    ### Result
    
    After filtering, proceed to Step 4 only with the relevant subset of events.
    If all events are filtered out, report that no relevant Health events were found
    and suggest alternative investigation paths (see Step 7).
    
    ---
    
    ## Step 4: Get Event Details
    
    For each **relevant** event identified in Step 3, retrieve full descriptions
    and timelines using `DescribeEventDetails`.
    
    ### API call pattern
    
    ```
    aws health describe-event-details \
      --region us-east-1 \
      --event-arns '["<arn-1>", "<arn-2>", ..., "<arn-10>"]'
    ```
    
    ### Batching rules
    
    - The API accepts a **maximum of 10 event ARNs per request**.
    - If more than 10 relevant events need details, issue multiple batched calls of
      up to 10 ARNs each until all relevant events are detailed.
    
    ### Extract from each event detail
    
    - **Event description** — the `latestDescription` text explaining the event.
    - **Timeline** — start time, end time (null if ongoing), last updated time.
    - **Status** — current status (open, closed, upcoming).
    - **Service and region** — confirm the affected service and region.
    - **Event type** — the category (issue, scheduledChange, accountNotification)
      and specific type code.
    
    ### Handling failedSet
    
    - The response contains a `successfulSet` and a `failedSet`.
    - If any event ARNs appear in `failedSet`, report the failed ARN and error
      message to the operator.
    - Continue processing all events from `successfulSet` without blocking on
      failures.
    
    ---
    
    ## Step 5: Identify Affected Entities
    
    For events with `eventScopeCode` of **ACCOUNT_SPECIFIC**, retrieve the list of
    affected resources using `DescribeAffectedEntities`.
    
    > **Important**: Only call DescribeAffectedEntities for ACCOUNT_SPECIFIC events.
    > PUBLIC events do not return entity data.
    
    ### API call pattern
    
    ```
    aws health describe-affected-entities \
      --region us-east-1 \
      --filter '{"eventArns": ["<event-arn>"]}'
      --max-results 100
    ```
    
    ### Entity matching
    
    When the incident context includes specific resource identifiers:
    
    1. Retrieve all affected entities for the event (paginate up to **500 entities**
       per event using `nextToken`).
    2. Perform exact string matching of each entity's `entityValue` against the
       incident context resource identifiers.
    3. Present matched entities in a separate section before non-matched entities.
    4. Include entity status (`IMPAIRED`, `UNIMPAIRED`, `UNKNOWN`, `PENDING`) and
       last updated time for each entity.
    
    ### Error handling
    
    - If `DescribeAffectedEntities` returns an error for a specific event ARN,
      report the event ARN that failed and continue processing remaining events.
    
    ---
    
    ## Step 6: Correlate with Incident
    
    Score each Health event for relevance to the current incident using the
    following criteria:
    
    ### Relevance scoring
    
    | Classification | Criteria | Label |
    |---------------|----------|-------|
    | **High** | Matching service + overlapping timeframe + matching affected resource (or matching region/AZ if no resource IDs available) | Likely contributing factor (if event is open) |
    | **Medium** | Matching service + overlapping timeframe (no resource match) | Likely contributing factor (if event is open) |
    | **Low** | Matching service only (no timeframe overlap) | Background context |
    
    ### Scoring rules
    
    - **Service match**: The event's service code matches one of the affected
      services from the incident context.
    - **Timeframe overlap**: The event's active period (start time through end time,
      or through present if still open) intersects with the incident's timeframe
      (start time through end time, or through present if ongoing).
    - **Region/AZ match**: The event's region or availability zone matches the
      incident's affected region or AZ.
    - **Resource match**: At least one affected entity's `entityValue` matches a
      resource identifier from the incident context.
    
    ### Contributing factor labeling
    
    - Any **open** event classified as High or Medium relevance SHALL be labeled as
      a "likely contributing factor" in addition to its relevance classification.
    - Closed events with High relevance should be noted as potential recent causes
      if the incident started shortly after the event closed.
    
    ### When resource identifiers are unavailable
    
    If the incident context does not include specific resource identifiers, score
    relevance using only service, timeframe, and region/AZ factors:
    
    - **High**: Matching service + overlapping timeframe + matching region or AZ
    - **Medium**: Matching service + overlapping timeframe
    - **Low**: Matching service only
    
    ---
    
    ## Step 7: Present Structured Output
    
    Present findings in a clear, structured format organized for quick
    comprehension and action.
    
    ### Output structure
    
    1. **Summary** — total events found, broken down by category and status.
    2. **Correlated events** — grouped by event type category in this order:
       - Issues (service disruptions) — present first
       - Scheduled changes (maintenance) — present second
       - Account notifications — present last
    3. **Within each group** — sort by:
       - Relevance classification (High → Medium → Low)
       - Then by start time descending (most recent first)
    4. **Per event** — include:
       - Event type category and service
       - Region and availability zone (if applicable)
       - Status (open/closed/upcoming)
       - Start time and end time (ISO 8601)
       - Description (summarized to 256 characters max)
       - Relevance classification and matching criteria
       - Contributing factor label (if applicable)
    5. **Actionable next steps** — for each correlated event, include at least one
       recommendation such as:
       - Check specific affected resources
       - Review related service limits or quotas
       - Verify recent configuration changes
       - Monitor the AWS Health Dashboard for updates
       - Contact AWS Support if the event is ongoing
    
    ### When no events are found
    
    If no relevant Health events are identified:
    
    - Explicitly state that no matching AWS Health events were found.
    - Confirm the search parameters used (service, time range, region).
    - Recommend checking other potential causes:
      - Recent deployments or configuration changes
      - Resource limits or quota exhaustion
      - Network connectivity issues
      - Application-level errors
    
    ---
    
    ## Decision Tree: Event Search Strategy
    
    ```
    Is this a chat-based health report request?
    ├── YES → Search the user-specified time period (default 30 days, max 90 days)
    │         Organize results by category, service, and status
    │         Present as a summary report
    └── NO → Continue with incident investigation flow below
    
    Is the affected AWS service known?
    ├── YES → Search events for that service within the past 7 days
    │   ├── Events found → Proceed to Step 3 (Filter Relevant Events)
    │   └── No events found → Broaden to related services (see Service Dependency Map)
    │       ├── Events found → Proceed to Step 3
    │       └── No events found → Expand time window to 14 days and retry
    │           ├── Events found → Proceed to Step 3
    │           └── No events found → Report no events found, suggest other investigation paths
    └── NO → Search all services filtered by region and availability zone (past 7 days)
        ├── Events found → Proceed to Step 3
        └── No events found → Expand time window to 14 days
            ├── Events found → Proceed to Step 3
            └── No events found → Report no events found, suggest other investigation paths
    
    Does the incident involve a specific availability zone?
    ├── YES → Include the AZ filter in all searches above
    └── NO → Filter by region only
    ```
    
    ---
    
    ## Service Dependency Map
    
    When the initial service-specific search returns no results, broaden the search
    to related services that share infrastructure dependencies:
    
    | Primary Service | Related Services to Check |
    |----------------|--------------------------|
    | ELB / ALB / NLB | EC2, VPC, Route 53 |
    | RDS | EC2, EBS |
    | ECS / EKS | EC2, VPC, ELB |
    | Lambda | VPC, CloudWatch |
    | CloudFront | S3, Route 53 |
    | API Gateway | Lambda, VPC |
    | ElastiCache | EC2, VPC |
    | DynamoDB | VPC (if VPC endpoints used) |
    | S3 | CloudFront, VPC (if VPC endpoints used) |
    | Kinesis | EC2, VPC |
    
    Search up to 3 related services when broadening. Use the Health API service
    codes from the references document (e.g., `ELASTICLOADBALANCING` for ELB,
    `ROUTE53` for Route 53).
    
    ---
    
    ## Error Handling
    
    | Error Condition | Agent Behavior |
    |----------------|----------------|
    | Missing `health:Describe*` permissions | Report the missing permissions and specify the required IAM actions: `health:DescribeEvents`, `health:DescribeEventDetails`, `health:DescribeAffectedEntities`, `health:DescribeEventTypes`. Provide the IAM policy snippet needed. |
    | Throttling (HTTP 429) | Retry with exponential backoff: wait 1s → 2s → 4s (max 3 retries). If still throttled after 3 retries, report that the Health API is currently rate-limited and recommend trying again shortly. |
    | Service error (HTTP 5xx) | Report the error code and recommend the operator check the AWS Health Dashboard directly as a fallback. |
    | Timeout (30 seconds) | Cancel the request and report a timeout error. Suggest the operator check the Health Dashboard directly or retry with narrower filters. |
    | Zero events found | Report that no events matched the specified filters. Confirm the search parameters used. Suggest broadening the search or checking other investigation paths. |
    | Invalid time range (start > end) | Report the invalid time range error. Ask the operator to provide corrected timestamps. |
    | DescribeEventDetails failedSet | Report the failed event ARNs and error messages. Continue processing events from the successfulSet. |
    | DescribeAffectedEntities error | Report the event ARN for which entity retrieval failed. Continue processing remaining events. |
    | Unknown service name (chat report) | Inform the user the service was not recognized. List services that have events in the requested time period. |
    
    ---
    
    ## Tips for Effective Health Event Review
    
    - **Always check us-east-1**: The Health API endpoint is only in us-east-1,
      regardless of where your resources are located.
    - **Start narrow, then broaden**: Begin with the specific affected service and
      a 7-day window. Only expand if no results are found.
    - **Check both open and closed events**: A recently closed event may still be
      causing residual impact.
    - **Correlate with support cases**: If a Health event references a service
      disruption, check if related support cases exist using the support-cases skill.
    - **Account-specific vs public events**: Account-specific events directly affect
      your resources. Public events are service-wide but may still impact you.
    - **Look at scheduled changes**: Upcoming or recent maintenance windows can
      explain transient issues that resolve on their own.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related