Claude Skill

cloudwatch

AWS CloudWatch monitoring for logs, metrics, alarms, and dashboards. Use when setting up monitoring, creating alarms, querying logs with Insights, configuring metric filters, building dashboards, or troubleshooting application issues.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download itsmostafa-aws-agent-skills-skills_cloudwatch-e786d25.zip · 9 KB
Part of itsmostafa/aws-agent-skills — 17 skills

Install

skills CLI npx skills add https://github.com/itsmostafa/aws-agent-skills/tree/main/skills/cloudwatch
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install itsmostafa-aws-agent-skills@llmmart
Git git clone https://github.com/itsmostafa/aws-agent-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole itsmostafa/aws-agent-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

AWS CloudWatch

Amazon CloudWatch provides monitoring and observability for AWS resources and applications. It collects metrics, logs, and events, enabling you to monitor, troubleshoot, and optimize your AWS environment.

Table of Contents

Core Concepts

Metrics

Time-ordered data points published to CloudWatch. Key components:

  • Namespace: Container for metrics (e.g., AWS/Lambda)
  • Metric name: Name of the measurement (e.g., Invocations)
  • Dimensions: Name-value pairs for filtering (e.g., FunctionName=MyFunc)
  • Statistics: Aggregations (Sum, Average, Min, Max, SampleCount, pN)

Logs

Log data from AWS services and applications:

  • Log groups: Collections of log streams
  • Log streams: Sequences of log events from same source
  • Log events: Individual log entries with timestamp and message

Alarms

Automated actions based on metric thresholds:

  • States: OK, ALARM, INSUFFICIENT_DATA
  • Actions: SNS notifications, Lambda functions, Auto Scaling, EC2 actions, Systems Manager OpsItems
  • Types: metric (incl. metric math, anomaly detection, Metrics Insights), composite, log (scheduled Logs Insights query evaluated M-out-of-N on query results)
  • Mute rules: scheduled windows that mute actions while alarms keep evaluating

Common Patterns

Create a Metric Alarm

AWS CLI:

# CPU utilization alarm for EC2
aws cloudwatch put-metric-alarm \
  --alarm-name "HighCPU-i-1234567890abcdef0" \
  --metric-name CPUUtilization \
  --namespace AWS/EC2 \
  --statistic Average \
  --period 300 \
  --threshold 80 \
  --comparison-operator GreaterThanThreshold \
  --evaluation-periods 2 \
  --dimensions Name=InstanceId,Value=i-1234567890abcdef0 \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts \
  --ok-actions arn:aws:sns:us-east-1:123456789012:alerts

boto3:

import boto3

cloudwatch = boto3.client('cloudwatch')

cloudwatch.put_metric_alarm(
    AlarmName='HighCPU-i-1234567890abcdef0',
    MetricName='CPUUtilization',
    Namespace='AWS/EC2',
    Statistic='Average',
    Period=300,
    Threshold=80.0,
    ComparisonOperator='GreaterThanThreshold',
    EvaluationPeriods=2,
    Dimensions=[
        {'Name': 'InstanceId', 'Value': 'i-1234567890abcdef0'}
    ],
    AlarmActions=['arn:aws:sns:us-east-1:123456789012:alerts'],
    OKActions=['arn:aws:sns:us-east-1:123456789012:alerts']
)

Lambda Error Rate Alarm

aws cloudwatch put-metric-alarm \
  --alarm-name "LambdaErrorRate-MyFunction" \
  --metrics '[
    {
      "Id": "errors",
      "MetricStat": {
        "Metric": {
          "Namespace": "AWS/Lambda",
          "MetricName": "Errors",
          "Dimensions": [{"Name": "FunctionName", "Value": "MyFunction"}]
        },
        "Period": 60,
        "Stat": "Sum"
      },
      "ReturnData": false
    },
    {
      "Id": "invocations",
      "MetricStat": {
        "Metric": {
          "Namespace": "AWS/Lambda",
          "MetricName": "Invocations",
          "Dimensions": [{"Name": "FunctionName", "Value": "MyFunction"}]
        },
        "Period": 60,
        "Stat": "Sum"
      },
      "ReturnData": false
    },
    {
      "Id": "errorRate",
      "Expression": "errors/invocations*100",
      "Label": "Error Rate",
      "ReturnData": true
    }
  ]' \
  --threshold 5 \
  --comparison-operator GreaterThanThreshold \
  --evaluation-periods 3 \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts

Query Logs with Insights

# Find errors in Lambda logs
aws logs start-query \
  --log-group-name /aws/lambda/MyFunction \
  --start-time $(date -d '1 hour ago' +%s) \
  --end-time $(date +%s) \
  --query-string '
    fields @timestamp, @message
    | filter @message like /ERROR/
    | sort @timestamp desc
    | limit 50
  '

# Get query results
aws logs get-query-results --query-id <query-id>

boto3:

import boto3
import time

logs = boto3.client('logs')

# Start query
response = logs.start_query(
    logGroupName='/aws/lambda/MyFunction',
    startTime=int(time.time()) - 3600,
    endTime=int(time.time()),
    queryString='''
        fields @timestamp, @message
        | filter @message like /ERROR/
        | sort @timestamp desc
        | limit 50
    '''
)

query_id = response['queryId']

# Wait for results
while True:
    result = logs.get_query_results(queryId=query_id)
    if result['status'] == 'Complete':
        break
    time.sleep(1)

for row in result['results']:
    print(row)

Create Metric Filter

Extract metrics from log patterns:

# Create metric filter for error count
aws logs put-metric-filter \
  --log-group-name /aws/lambda/MyFunction \
  --filter-name ErrorCount \
  --filter-pattern "ERROR" \
  --metric-transformations \
    metricName=ErrorCount,metricNamespace=MyApp,metricValue=1,defaultValue=0

Create a Log Alarm

Alarm directly on a Logs Insights query (no metric filter needed). CloudWatch creates and manages the underlying scheduled query. The ScheduledQueryRoleARN role must trust logs.amazonaws.com and allow logs:StartQuery and logs:GetQueryResults on the log group ARNs, plus logs:StopQuery and logs:DescribeLogGroups on "Resource": "*" (these two support no resource-level permissions, so a log-group-scoped statement never matches them).

# ALARM when >100 errors in 3 of the last 5 query runs
aws cloudwatch put-log-alarm \
  --alarm-name "HighErrorCount" \
  --comparison-operator GreaterThanThreshold \
  --threshold 100 \
  --query-results-to-evaluate 5 \
  --query-results-to-alarm 3 \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts \
  --scheduled-query-configuration '{
    "QueryString": "fields @timestamp, @message | filter @message like /ERROR/",
    "LogGroupIdentifiers": ["/aws/lambda/MyFunction"],
    "ScheduledQueryRoleARN": "arn:aws:iam::123456789012:role/ScheduledQueryRole",
    "AggregationExpression": "count(*)",
    "ScheduleConfiguration": {
      "ScheduleExpression": "rate(10 minutes)",
      "StartTimeOffset": 600
    }
  }'

# Log alarms are omitted from describe-alarms unless requested
aws cloudwatch describe-alarms --alarm-types LogAlarm

Publish Custom Metrics

import boto3

cloudwatch = boto3.client('cloudwatch')

cloudwatch.put_metric_data(
    Namespace='MyApp',
    MetricData=[
        {
            'MetricName': 'OrdersProcessed',
            'Value': 1,
            'Unit': 'Count',
            'Dimensions': [
                {'Name': 'Environment', 'Value': 'Production'},
                {'Name': 'OrderType', 'Value': 'Standard'}
            ]
        }
    ]
)

Create Dashboard

cat > dashboard.json << 'EOF'
{
  "widgets": [
    {
      "type": "metric",
      "x": 0, "y": 0, "width": 12, "height": 6,
      "properties": {
        "title": "Lambda Invocations",
        "metrics": [
          ["AWS/Lambda", "Invocations", "FunctionName", "MyFunction"]
        ],
        "period": 60,
        "stat": "Sum",
        "region": "us-east-1"
      }
    },
    {
      "type": "log",
      "x": 12, "y": 0, "width": 12, "height": 6,
      "properties": {
        "title": "Recent Errors",
        "query": "SOURCE '/aws/lambda/MyFunction' | filter @message like /ERROR/ | limit 20",
        "region": "us-east-1"
      }
    }
  ]
}
EOF

aws cloudwatch put-dashboard \
  --dashboard-name MyAppDashboard \
  --dashboard-body file://dashboard.json

CLI Reference

Metrics Commands

Command Description
aws cloudwatch put-metric-data Publish custom metrics
aws cloudwatch get-metric-data Retrieve metric values
aws cloudwatch get-metric-statistics Get aggregated statistics
aws cloudwatch list-metrics List available metrics

Alarms Commands

Command Description
aws cloudwatch put-metric-alarm Create or update alarm
aws cloudwatch put-log-alarm Create or update log query alarm
aws cloudwatch describe-alarms List alarms (--alarm-types LogAlarm for log alarms)
aws cloudwatch describe-alarm-contributors Show breaching contributors of a multi-contributor alarm
aws cloudwatch put-alarm-mute-rule Create or update scheduled mute window
aws cloudwatch list-alarm-mute-rules List mute rules (--statuses SCHEDULED ACTIVE EXPIRED)
aws cloudwatch delete-alarm-mute-rule Delete mute rule (unmutes immediately)
aws cloudwatch set-alarm-state Manually set alarm state
aws cloudwatch delete-alarms Delete alarms

Logs Commands

Command Description
aws logs create-log-group Create log group
aws logs put-log-events Write log events
aws logs filter-log-events Search log events
aws logs start-query Start Insights query
aws logs put-metric-filter Create metric filter
aws logs put-retention-policy Set log retention

Best Practices

Metrics

  • Use dimensions wisely — too many creates metric explosion
  • Aggregate before publishing — batch custom metrics
  • Use high-resolution metrics (1-second) only when needed
  • Set meaningful units for custom metrics

Alarms

  • Use composite alarms for complex conditions
  • Set appropriate evaluation periods to avoid flapping
  • Include OK actions to track recovery
  • Use anomaly detection for dynamic thresholds
  • Use log alarms instead of metric filter + metric alarm for query-based conditions; set --treat-missing-data notBreaching for sparse errors, breaching to detect logs that stop arriving
  • Stagger log alarm schedules — concurrent scheduled query executions per account are capped at 100
  • Set --warm-up-configuration on alarms created alongside new resources (1-2880 min) to avoid noise before metrics publish
  • Use --evaluation-window WallClockWindow={Timezone=...} for daily/weekly batch or backup alarms; keep the default sliding window for Auto Scaling
  • Use mute rules for maintenance, not disable-alarm-actions; enable-alarm-actions does not unmute an active mute rule

Logs

  • Set retention policies — don't keep logs forever
  • Use structured logging (JSON) for better querying
  • Create metric filters for key events
  • Use Contributor Insights for top-N analysis

Cost Optimization

  • Delete unused dashboards
  • Reduce log retention for non-critical logs
  • Avoid high-resolution metrics unless necessary
  • Use log subscription filters instead of polling

Troubleshooting

Missing Metrics

Causes:

  • Service not publishing yet (wait 1-5 minutes)
  • Wrong namespace/dimensions
  • Detailed monitoring not enabled (EC2)

Debug:

# List metrics for a namespace
aws cloudwatch list-metrics \
  --namespace AWS/Lambda \
  --dimensions Name=FunctionName,Value=MyFunction

Alarm Stuck in INSUFFICIENT_DATA

Causes:

  • Metric not being published
  • Dimensions mismatch
  • Evaluation period too short
  • Warm-up period still active (ends early once data fills the window unless OnlyStartEvaluatingAfterWarmUpPeriodEnds=true)
  • Log alarm: just created/updated (query, schedule, or log group changes reset state), scheduled query role lacks permissions (EvaluationState = EVALUATION_ERROR, see StateReason), or ingestion lag (shift window back with EndTimeOffset)

Debug:

# Check if metric has data
aws cloudwatch get-metric-statistics \
  --namespace AWS/Lambda \
  --metric-name Invocations \
  --dimensions Name=FunctionName,Value=MyFunction \
  --start-time $(date -d '1 hour ago' -u +%Y-%m-%dT%H:%M:%SZ) \
  --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
  --period 60 \
  --statistics Sum

# Log alarm: check evaluation state and reason
aws cloudwatch describe-alarms --alarm-types LogAlarm --alarm-names HighErrorCount \
  --query 'LogAlarms[].[StateValue,EvaluationState,StateReason]'

Log Events Not Appearing

Causes:

  • IAM permissions missing
  • CloudWatch Logs agent not running
  • Log group doesn't exist

Debug:

# Check log streams
aws logs describe-log-streams \
  --log-group-name /aws/lambda/MyFunction \
  --order-by LastEventTime \
  --descending \
  --limit 5

High CloudWatch Costs

Check usage:

# Get PutLogEvents usage
aws cloudwatch get-metric-statistics \
  --namespace AWS/Logs \
  --metric-name IncomingBytes \
  --dimensions Name=LogGroupName,Value=/aws/lambda/MyFunction \
  --start-time $(date -d '7 days ago' -u +%Y-%m-%dT%H:%M:%SZ) \
  --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
  --period 86400 \
  --statistics Sum

References

Files (aws-agent-skills)
  • alarms-metrics.md 14.8 KB
    # CloudWatch Alarms and Metrics
    
    Detailed patterns for CloudWatch alarms and custom metrics.
    
    ## Alarm Types
    
    ### Standard Metric Alarm
    
    Triggers based on a single metric threshold:
    
    ```bash
    aws cloudwatch put-metric-alarm \
      --alarm-name "HighCPUUtilization" \
      --metric-name CPUUtilization \
      --namespace AWS/EC2 \
      --statistic Average \
      --period 300 \
      --threshold 80 \
      --comparison-operator GreaterThanOrEqualToThreshold \
      --evaluation-periods 2 \
      --dimensions Name=InstanceId,Value=i-1234567890abcdef0 \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    ### Metric Math Alarm
    
    Combines multiple metrics with expressions:
    
    ```bash
    aws cloudwatch put-metric-alarm \
      --alarm-name "HighErrorRate" \
      --metrics '[
        {
          "Id": "e1",
          "MetricStat": {
            "Metric": {
              "Namespace": "AWS/ApplicationELB",
              "MetricName": "HTTPCode_ELB_5XX_Count",
              "Dimensions": [{"Name": "LoadBalancer", "Value": "app/my-lb/1234567890123456"}]
            },
            "Period": 60,
            "Stat": "Sum"
          },
          "ReturnData": false
        },
        {
          "Id": "e2",
          "MetricStat": {
            "Metric": {
              "Namespace": "AWS/ApplicationELB",
              "MetricName": "RequestCount",
              "Dimensions": [{"Name": "LoadBalancer", "Value": "app/my-lb/1234567890123456"}]
            },
            "Period": 60,
            "Stat": "Sum"
          },
          "ReturnData": false
        },
        {
          "Id": "e3",
          "Expression": "IF(e2>0, e1/e2*100, 0)",
          "Label": "Error Rate %",
          "ReturnData": true
        }
      ]' \
      --threshold 5 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 3 \
      --datapoints-to-alarm 2 \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    ### Anomaly Detection Alarm
    
    Uses machine learning to detect anomalies:
    
    ```bash
    aws cloudwatch put-metric-alarm \
      --alarm-name "AnomalousLatency" \
      --metrics '[
        {
          "Id": "m1",
          "MetricStat": {
            "Metric": {
              "Namespace": "AWS/Lambda",
              "MetricName": "Duration",
              "Dimensions": [{"Name": "FunctionName", "Value": "MyFunction"}]
            },
            "Period": 60,
            "Stat": "Average"
          },
          "ReturnData": true
        },
        {
          "Id": "ad1",
          "Expression": "ANOMALY_DETECTION_BAND(m1, 2)",
          "Label": "AnomalyDetectionBand",
          "ReturnData": true
        }
      ]' \
      --threshold-metric-id ad1 \
      --comparison-operator LessThanLowerOrGreaterThanUpperThreshold \
      --evaluation-periods 3 \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    ### Composite Alarm
    
    Combines multiple alarms with boolean logic:
    
    ```bash
    # Create component alarms first
    aws cloudwatch put-metric-alarm \
      --alarm-name "HighCPU" \
      --metric-name CPUUtilization \
      --namespace AWS/EC2 \
      --statistic Average \
      --period 300 \
      --threshold 80 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 2 \
      --dimensions Name=InstanceId,Value=i-1234567890abcdef0
    
    aws cloudwatch put-metric-alarm \
      --alarm-name "HighMemory" \
      --metric-name MemoryUtilization \
      --namespace CWAgent \
      --statistic Average \
      --period 300 \
      --threshold 80 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 2 \
      --dimensions Name=InstanceId,Value=i-1234567890abcdef0
    
    # Create composite alarm
    aws cloudwatch put-composite-alarm \
      --alarm-name "InstanceUnhealthy" \
      --alarm-rule "ALARM(HighCPU) AND ALARM(HighMemory)" \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    ### Log Alarm (Multi-Contributor)
    
    Evaluates a scheduled Logs Insights query. A `by` clause in `AggregationExpression` tracks each group as a contributor; the alarm is in ALARM if any contributor breaches. SNS/Lambda actions fire per contributor.
    
    ```bash
    aws cloudwatch put-log-alarm \
      --alarm-name "EndpointLatency" \
      --comparison-operator GreaterThanThreshold \
      --threshold 2000 \
      --query-results-to-evaluate 3 \
      --query-results-to-alarm 2 \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts \
      --scheduled-query-configuration '{
        "QueryString": "fields endpoint, latency | filter ispresent(latency)",
        "LogGroupIdentifiers": ["/app/api"],
        "ScheduledQueryRoleARN": "arn:aws:iam::123456789012:role/ScheduledQueryRole",
        "AggregationExpression": "avg(latency) by endpoint | sort desc",
        "ScheduleConfiguration": {
          "ScheduleExpression": "rate(5 minutes)",
          "StartTimeOffset": 420,
          "EndTimeOffset": 120
        }
      }'
    
    # Which contributors are in ALARM
    aws cloudwatch describe-alarm-contributors --alarm-name EndpointLatency
    ```
    
    - Aggregation functions: `count(*)`, `avg`, `sum`, `min`, `max`; `bin()` not allowed in the `by` clause
    - Limits: 5 `by` fields, 500 contributors returned per run (sort by value so the worst are kept; `EvaluationState` = `PARTIAL_DATA` when exceeded), 100 contributors in ALARM
    - `--action-log-line-count` (0-50) adds raw log lines to SNS email notifications; requires `--action-log-line-role-arn` (role trusting `cloudwatch.amazonaws.com` with `logs:GetQueryResults`)
    - `EndTimeOffset` shifts the window back to allow for ingestion delay
    
    ### Wall Clock Window and Warm-Up
    
    ```bash
    # Daily backup check aligned to local midnight, quiet for 60 min after create/update
    aws cloudwatch put-metric-alarm \
      --alarm-name "DailyBackupMissing" \
      --namespace MyApp \
      --metric-name BackupsCompleted \
      --statistic Sum \
      --period 86400 \
      --evaluation-periods 1 \
      --threshold 1 \
      --comparison-operator LessThanThreshold \
      --treat-missing-data breaching \
      --evaluation-window 'WallClockWindow={Timezone=America/New_York}' \
      --warm-up-configuration WarmUpPeriodDurationInMinutes=60 \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    - Wall clock windows support periods 60, 300, 3600, 86400, 604800 only; not PromQL alarms; data is reflected only after the clock period ends
    - Metrics Insights alarms with a wall clock window: Period x (EvaluationPeriods + 1) must be <= 3 hours
    - Warm-up (1-2880 min) applies to metric and log alarms; alarm stays INSUFFICIENT_DATA with no actions. Ends early once data fills the window unless `OnlyStartEvaluatingAfterWarmUpPeriodEnds=true`
    
    ## Common Alarm Patterns
    
    ### Lambda Function Health
    
    ```bash
    # Errors alarm
    aws cloudwatch put-metric-alarm \
      --alarm-name "Lambda-MyFunction-Errors" \
      --metric-name Errors \
      --namespace AWS/Lambda \
      --statistic Sum \
      --period 60 \
      --threshold 5 \
      --comparison-operator GreaterThanOrEqualToThreshold \
      --evaluation-periods 1 \
      --dimensions Name=FunctionName,Value=MyFunction \
      --treat-missing-data notBreaching \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    
    # Throttles alarm
    aws cloudwatch put-metric-alarm \
      --alarm-name "Lambda-MyFunction-Throttles" \
      --metric-name Throttles \
      --namespace AWS/Lambda \
      --statistic Sum \
      --period 60 \
      --threshold 1 \
      --comparison-operator GreaterThanOrEqualToThreshold \
      --evaluation-periods 1 \
      --dimensions Name=FunctionName,Value=MyFunction \
      --treat-missing-data notBreaching \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    
    # Duration alarm (approaching timeout)
    aws cloudwatch put-metric-alarm \
      --alarm-name "Lambda-MyFunction-Duration" \
      --metric-name Duration \
      --namespace AWS/Lambda \
      --extended-statistic p99 \
      --period 300 \
      --threshold 25000 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 2 \
      --dimensions Name=FunctionName,Value=MyFunction \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    ### DynamoDB Table Health
    
    ```bash
    # Consumed capacity alarm
    aws cloudwatch put-metric-alarm \
      --alarm-name "DynamoDB-MyTable-HighReadCapacity" \
      --metric-name ConsumedReadCapacityUnits \
      --namespace AWS/DynamoDB \
      --statistic Average \
      --period 60 \
      --threshold 80 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 5 \
      --dimensions Name=TableName,Value=MyTable \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    
    # Throttled requests
    aws cloudwatch put-metric-alarm \
      --alarm-name "DynamoDB-MyTable-Throttled" \
      --metric-name ThrottledRequests \
      --namespace AWS/DynamoDB \
      --statistic Sum \
      --period 60 \
      --threshold 1 \
      --comparison-operator GreaterThanOrEqualToThreshold \
      --evaluation-periods 1 \
      --dimensions Name=TableName,Value=MyTable \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    ### API Gateway Health
    
    ```bash
    # 5XX errors
    aws cloudwatch put-metric-alarm \
      --alarm-name "APIGW-MyAPI-5XXErrors" \
      --metric-name 5XXError \
      --namespace AWS/ApiGateway \
      --statistic Sum \
      --period 60 \
      --threshold 10 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 2 \
      --dimensions Name=ApiName,Value=MyAPI \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    
    # Latency P99
    aws cloudwatch put-metric-alarm \
      --alarm-name "APIGW-MyAPI-HighLatency" \
      --metric-name Latency \
      --namespace AWS/ApiGateway \
      --extended-statistic p99 \
      --period 300 \
      --threshold 5000 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 3 \
      --dimensions Name=ApiName,Value=MyAPI \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    ### SQS Queue Health
    
    ```bash
    # Messages in queue (backlog)
    aws cloudwatch put-metric-alarm \
      --alarm-name "SQS-MyQueue-Backlog" \
      --metric-name ApproximateNumberOfMessagesVisible \
      --namespace AWS/SQS \
      --statistic Average \
      --period 300 \
      --threshold 1000 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 3 \
      --dimensions Name=QueueName,Value=MyQueue \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    
    # Age of oldest message
    aws cloudwatch put-metric-alarm \
      --alarm-name "SQS-MyQueue-OldMessages" \
      --metric-name ApproximateAgeOfOldestMessage \
      --namespace AWS/SQS \
      --statistic Maximum \
      --period 300 \
      --threshold 3600 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 2 \
      --dimensions Name=QueueName,Value=MyQueue \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    ## Custom Metrics
    
    ### High-Resolution Metrics
    
    Publish metrics with 1-second resolution:
    
    ```python
    import boto3
    
    cloudwatch = boto3.client('cloudwatch')
    
    cloudwatch.put_metric_data(
        Namespace='MyApp',
        MetricData=[
            {
                'MetricName': 'RequestLatency',
                'Value': 125.5,
                'Unit': 'Milliseconds',
                'StorageResolution': 1,  # 1 second resolution
                'Dimensions': [
                    {'Name': 'Endpoint', 'Value': '/api/orders'}
                ]
            }
        ]
    )
    ```
    
    ### Batched Metric Publishing
    
    ```python
    import boto3
    from datetime import datetime
    
    cloudwatch = boto3.client('cloudwatch')
    
    # Collect metrics
    metric_data = [
        {
            'MetricName': 'OrdersProcessed',
            'Value': 42,
            'Unit': 'Count',
            'Timestamp': datetime.utcnow(),
            'Dimensions': [
                {'Name': 'Environment', 'Value': 'Production'}
            ]
        },
        {
            'MetricName': 'ProcessingTime',
            'Value': 250.5,
            'Unit': 'Milliseconds',
            'Timestamp': datetime.utcnow(),
            'Dimensions': [
                {'Name': 'Environment', 'Value': 'Production'}
            ]
        }
    ]
    
    # Publish in batch (max 1000 per call)
    cloudwatch.put_metric_data(
        Namespace='MyApp',
        MetricData=metric_data
    )
    ```
    
    ### Embedded Metric Format (EMF)
    
    Publish metrics directly from Lambda logs:
    
    ```python
    import json
    
    def handler(event, context):
        # Process request
        processing_time = 125.5
    
        # Emit EMF log
        print(json.dumps({
            "_aws": {
                "Timestamp": int(time.time() * 1000),
                "CloudWatchMetrics": [{
                    "Namespace": "MyApp",
                    "Dimensions": [["Environment", "Endpoint"]],
                    "Metrics": [
                        {"Name": "RequestLatency", "Unit": "Milliseconds"},
                        {"Name": "RequestCount", "Unit": "Count"}
                    ]
                }]
            },
            "Environment": "Production",
            "Endpoint": "/api/orders",
            "RequestLatency": processing_time,
            "RequestCount": 1
        }))
    
        return {"statusCode": 200}
    ```
    
    ## Metric Math Functions
    
    | Function | Description | Example |
    |----------|-------------|---------|
    | `SUM` | Sum of metrics | `SUM([m1, m2, m3])` |
    | `AVG` | Average | `AVG([m1, m2])` |
    | `MIN`, `MAX` | Minimum/Maximum | `MAX([m1, m2])` |
    | `RATE` | Per-second rate of change | `RATE(m1)` |
    | `PERIOD` | Current period in seconds | `m1 / PERIOD(m1)` |
    | `FILL` | Replace missing data | `FILL(m1, 0)` |
    | `IF` | Conditional | `IF(m1 > 100, m1, 0)` |
    | `ANOMALY_DETECTION_BAND` | ML-based band | `ANOMALY_DETECTION_BAND(m1, 2)` |
    | `SEARCH` | Dynamic metrics | `SEARCH('{AWS/EC2,InstanceId} MetricName="CPUUtilization"', 'Average', 300)` |
    
    ### Example: Percentage Calculation
    
    ```
    e1 = errors metric
    e2 = requests metric
    errorRate = IF(e2 > 0, (e1 / e2) * 100, 0)
    ```
    
    ### Example: Aggregate Across Dimensions
    
    ```
    SEARCH('{AWS/Lambda,FunctionName} MetricName="Errors"', 'Sum', 60)
    ```
    
    ## Alarm Actions
    
    ### SNS Notification
    
    ```bash
    --alarm-actions arn:aws:sns:us-east-1:123456789012:my-topic
    ```
    
    ### Auto Scaling
    
    ```bash
    --alarm-actions arn:aws:autoscaling:us-east-1:123456789012:scalingPolicy:12345678-1234-1234-1234-123456789012:autoScalingGroupName/my-asg:policyName/scale-out
    ```
    
    ### EC2 Actions
    
    ```bash
    # Stop instance
    --alarm-actions arn:aws:automate:us-east-1:ec2:stop
    
    # Terminate instance
    --alarm-actions arn:aws:automate:us-east-1:ec2:terminate
    
    # Recover instance
    --alarm-actions arn:aws:automate:us-east-1:ec2:recover
    ```
    
    ### Lambda Function
    
    Alarms invoke Lambda directly (function, version, or alias ARN). Grant the alarm permission first:
    
    ```bash
    aws lambda add-permission \
      --function-name HandleAlarm \
      --statement-id AlarmAction \
      --action lambda:InvokeFunction \
      --principal lambda.alarms.cloudwatch.amazonaws.com \
      --source-account 123456789012 \
      --source-arn arn:aws:cloudwatch:us-east-1:123456789012:alarm:HighCPU
    
    # Then on the alarm:
    --alarm-actions arn:aws:lambda:us-east-1:123456789012:function:HandleAlarm
    ```
    
    ## Alarm Mute Rules
    
    Mute actions (all states) during scheduled windows; alarms keep evaluating and EventBridge events still emit. Up to 100 alarms per rule.
    
    ```bash
    # Mute every Sunday 02:00 for 4 hours
    aws cloudwatch put-alarm-mute-rule \
      --name weekly-maintenance \
      --rule 'Schedule={Expression=cron(0 2 * * SUN),Duration=PT4H,Timezone=America/New_York}' \
      --mute-targets AlarmNames=HighCPU,HighMemory
    
    # One-time window: Expression=at(2026-12-23T00:00),Duration=P7D
    aws cloudwatch list-alarm-mute-rules --statuses ACTIVE
    aws cloudwatch delete-alarm-mute-rule --alarm-mute-rule-name weekly-maintenance
    ```
    
    - When the window ends (or the rule is deleted/updated or the alarm removed from targets), muted actions run if the alarm is still in the state it was muted in
    - `enable-alarm-actions` does not unmute; `disable-alarm-actions` is permanent until re-enabled
    - IAM: `cloudwatch:PutAlarmMuteRule` is needed on the mute rule resource and on each targeted alarm
    
  • SKILL.md 13.6 KB
    ---
    name: cloudwatch
    description: AWS CloudWatch monitoring for logs, metrics, alarms, and dashboards. Use when setting up monitoring, creating alarms, querying logs with Insights, configuring metric filters, building dashboards, or troubleshooting application issues.
    last_updated: "2026-09-14"
    doc_source: https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/
    ---
    
    # AWS CloudWatch
    
    Amazon CloudWatch provides monitoring and observability for AWS resources and applications. It collects metrics, logs, and events, enabling you to monitor, troubleshoot, and optimize your AWS environment.
    
    ## Table of Contents
    
    - [Core Concepts](#core-concepts)
    - [Common Patterns](#common-patterns)
    - [CLI Reference](#cli-reference)
    - [Best Practices](#best-practices)
    - [Troubleshooting](#troubleshooting)
    - [References](#references)
    
    ## Core Concepts
    
    ### Metrics
    
    Time-ordered data points published to CloudWatch. Key components:
    - **Namespace**: Container for metrics (e.g., `AWS/Lambda`)
    - **Metric name**: Name of the measurement (e.g., `Invocations`)
    - **Dimensions**: Name-value pairs for filtering (e.g., `FunctionName=MyFunc`)
    - **Statistics**: Aggregations (Sum, Average, Min, Max, SampleCount, pN)
    
    ### Logs
    
    Log data from AWS services and applications:
    - **Log groups**: Collections of log streams
    - **Log streams**: Sequences of log events from same source
    - **Log events**: Individual log entries with timestamp and message
    
    ### Alarms
    
    Automated actions based on metric thresholds:
    - **States**: OK, ALARM, INSUFFICIENT_DATA
    - **Actions**: SNS notifications, Lambda functions, Auto Scaling, EC2 actions, Systems Manager OpsItems
    - **Types**: metric (incl. metric math, anomaly detection, Metrics Insights), composite, log (scheduled Logs Insights query evaluated M-out-of-N on query results)
    - **Mute rules**: scheduled windows that mute actions while alarms keep evaluating
    
    ## Common Patterns
    
    ### Create a Metric Alarm
    
    **AWS CLI:**
    
    ```bash
    # CPU utilization alarm for EC2
    aws cloudwatch put-metric-alarm \
      --alarm-name "HighCPU-i-1234567890abcdef0" \
      --metric-name CPUUtilization \
      --namespace AWS/EC2 \
      --statistic Average \
      --period 300 \
      --threshold 80 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 2 \
      --dimensions Name=InstanceId,Value=i-1234567890abcdef0 \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts \
      --ok-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    **boto3:**
    
    ```python
    import boto3
    
    cloudwatch = boto3.client('cloudwatch')
    
    cloudwatch.put_metric_alarm(
        AlarmName='HighCPU-i-1234567890abcdef0',
        MetricName='CPUUtilization',
        Namespace='AWS/EC2',
        Statistic='Average',
        Period=300,
        Threshold=80.0,
        ComparisonOperator='GreaterThanThreshold',
        EvaluationPeriods=2,
        Dimensions=[
            {'Name': 'InstanceId', 'Value': 'i-1234567890abcdef0'}
        ],
        AlarmActions=['arn:aws:sns:us-east-1:123456789012:alerts'],
        OKActions=['arn:aws:sns:us-east-1:123456789012:alerts']
    )
    ```
    
    ### Lambda Error Rate Alarm
    
    ```bash
    aws cloudwatch put-metric-alarm \
      --alarm-name "LambdaErrorRate-MyFunction" \
      --metrics '[
        {
          "Id": "errors",
          "MetricStat": {
            "Metric": {
              "Namespace": "AWS/Lambda",
              "MetricName": "Errors",
              "Dimensions": [{"Name": "FunctionName", "Value": "MyFunction"}]
            },
            "Period": 60,
            "Stat": "Sum"
          },
          "ReturnData": false
        },
        {
          "Id": "invocations",
          "MetricStat": {
            "Metric": {
              "Namespace": "AWS/Lambda",
              "MetricName": "Invocations",
              "Dimensions": [{"Name": "FunctionName", "Value": "MyFunction"}]
            },
            "Period": 60,
            "Stat": "Sum"
          },
          "ReturnData": false
        },
        {
          "Id": "errorRate",
          "Expression": "errors/invocations*100",
          "Label": "Error Rate",
          "ReturnData": true
        }
      ]' \
      --threshold 5 \
      --comparison-operator GreaterThanThreshold \
      --evaluation-periods 3 \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts
    ```
    
    ### Query Logs with Insights
    
    ```bash
    # Find errors in Lambda logs
    aws logs start-query \
      --log-group-name /aws/lambda/MyFunction \
      --start-time $(date -d '1 hour ago' +%s) \
      --end-time $(date +%s) \
      --query-string '
        fields @timestamp, @message
        | filter @message like /ERROR/
        | sort @timestamp desc
        | limit 50
      '
    
    # Get query results
    aws logs get-query-results --query-id <query-id>
    ```
    
    **boto3:**
    
    ```python
    import boto3
    import time
    
    logs = boto3.client('logs')
    
    # Start query
    response = logs.start_query(
        logGroupName='/aws/lambda/MyFunction',
        startTime=int(time.time()) - 3600,
        endTime=int(time.time()),
        queryString='''
            fields @timestamp, @message
            | filter @message like /ERROR/
            | sort @timestamp desc
            | limit 50
        '''
    )
    
    query_id = response['queryId']
    
    # Wait for results
    while True:
        result = logs.get_query_results(queryId=query_id)
        if result['status'] == 'Complete':
            break
        time.sleep(1)
    
    for row in result['results']:
        print(row)
    ```
    
    ### Create Metric Filter
    
    Extract metrics from log patterns:
    
    ```bash
    # Create metric filter for error count
    aws logs put-metric-filter \
      --log-group-name /aws/lambda/MyFunction \
      --filter-name ErrorCount \
      --filter-pattern "ERROR" \
      --metric-transformations \
        metricName=ErrorCount,metricNamespace=MyApp,metricValue=1,defaultValue=0
    ```
    
    ### Create a Log Alarm
    
    Alarm directly on a Logs Insights query (no metric filter needed). CloudWatch creates and manages the underlying scheduled query. The `ScheduledQueryRoleARN` role must trust `logs.amazonaws.com` and allow `logs:StartQuery` and `logs:GetQueryResults` on the log group ARNs, plus `logs:StopQuery` and `logs:DescribeLogGroups` on `"Resource": "*"` (these two support no resource-level permissions, so a log-group-scoped statement never matches them).
    
    ```bash
    # ALARM when >100 errors in 3 of the last 5 query runs
    aws cloudwatch put-log-alarm \
      --alarm-name "HighErrorCount" \
      --comparison-operator GreaterThanThreshold \
      --threshold 100 \
      --query-results-to-evaluate 5 \
      --query-results-to-alarm 3 \
      --treat-missing-data notBreaching \
      --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts \
      --scheduled-query-configuration '{
        "QueryString": "fields @timestamp, @message | filter @message like /ERROR/",
        "LogGroupIdentifiers": ["/aws/lambda/MyFunction"],
        "ScheduledQueryRoleARN": "arn:aws:iam::123456789012:role/ScheduledQueryRole",
        "AggregationExpression": "count(*)",
        "ScheduleConfiguration": {
          "ScheduleExpression": "rate(10 minutes)",
          "StartTimeOffset": 600
        }
      }'
    
    # Log alarms are omitted from describe-alarms unless requested
    aws cloudwatch describe-alarms --alarm-types LogAlarm
    ```
    
    ### Publish Custom Metrics
    
    ```python
    import boto3
    
    cloudwatch = boto3.client('cloudwatch')
    
    cloudwatch.put_metric_data(
        Namespace='MyApp',
        MetricData=[
            {
                'MetricName': 'OrdersProcessed',
                'Value': 1,
                'Unit': 'Count',
                'Dimensions': [
                    {'Name': 'Environment', 'Value': 'Production'},
                    {'Name': 'OrderType', 'Value': 'Standard'}
                ]
            }
        ]
    )
    ```
    
    ### Create Dashboard
    
    ```bash
    cat > dashboard.json << 'EOF'
    {
      "widgets": [
        {
          "type": "metric",
          "x": 0, "y": 0, "width": 12, "height": 6,
          "properties": {
            "title": "Lambda Invocations",
            "metrics": [
              ["AWS/Lambda", "Invocations", "FunctionName", "MyFunction"]
            ],
            "period": 60,
            "stat": "Sum",
            "region": "us-east-1"
          }
        },
        {
          "type": "log",
          "x": 12, "y": 0, "width": 12, "height": 6,
          "properties": {
            "title": "Recent Errors",
            "query": "SOURCE '/aws/lambda/MyFunction' | filter @message like /ERROR/ | limit 20",
            "region": "us-east-1"
          }
        }
      ]
    }
    EOF
    
    aws cloudwatch put-dashboard \
      --dashboard-name MyAppDashboard \
      --dashboard-body file://dashboard.json
    ```
    
    ## CLI Reference
    
    ### Metrics Commands
    
    | Command | Description |
    |---------|-------------|
    | `aws cloudwatch put-metric-data` | Publish custom metrics |
    | `aws cloudwatch get-metric-data` | Retrieve metric values |
    | `aws cloudwatch get-metric-statistics` | Get aggregated statistics |
    | `aws cloudwatch list-metrics` | List available metrics |
    
    ### Alarms Commands
    
    | Command | Description |
    |---------|-------------|
    | `aws cloudwatch put-metric-alarm` | Create or update alarm |
    | `aws cloudwatch put-log-alarm` | Create or update log query alarm |
    | `aws cloudwatch describe-alarms` | List alarms (`--alarm-types LogAlarm` for log alarms) |
    | `aws cloudwatch describe-alarm-contributors` | Show breaching contributors of a multi-contributor alarm |
    | `aws cloudwatch put-alarm-mute-rule` | Create or update scheduled mute window |
    | `aws cloudwatch list-alarm-mute-rules` | List mute rules (`--statuses SCHEDULED ACTIVE EXPIRED`) |
    | `aws cloudwatch delete-alarm-mute-rule` | Delete mute rule (unmutes immediately) |
    | `aws cloudwatch set-alarm-state` | Manually set alarm state |
    | `aws cloudwatch delete-alarms` | Delete alarms |
    
    ### Logs Commands
    
    | Command | Description |
    |---------|-------------|
    | `aws logs create-log-group` | Create log group |
    | `aws logs put-log-events` | Write log events |
    | `aws logs filter-log-events` | Search log events |
    | `aws logs start-query` | Start Insights query |
    | `aws logs put-metric-filter` | Create metric filter |
    | `aws logs put-retention-policy` | Set log retention |
    
    ## Best Practices
    
    ### Metrics
    
    - **Use dimensions wisely** — too many creates metric explosion
    - **Aggregate before publishing** — batch custom metrics
    - **Use high-resolution metrics** (1-second) only when needed
    - **Set meaningful units** for custom metrics
    
    ### Alarms
    
    - **Use composite alarms** for complex conditions
    - **Set appropriate evaluation periods** to avoid flapping
    - **Include OK actions** to track recovery
    - **Use anomaly detection** for dynamic thresholds
    - **Use log alarms** instead of metric filter + metric alarm for query-based conditions; set `--treat-missing-data notBreaching` for sparse errors, `breaching` to detect logs that stop arriving
    - **Stagger log alarm schedules** — concurrent scheduled query executions per account are capped at 100
    - **Set `--warm-up-configuration`** on alarms created alongside new resources (1-2880 min) to avoid noise before metrics publish
    - **Use `--evaluation-window WallClockWindow={Timezone=...}`** for daily/weekly batch or backup alarms; keep the default sliding window for Auto Scaling
    - **Use mute rules for maintenance**, not `disable-alarm-actions`; `enable-alarm-actions` does not unmute an active mute rule
    
    ### Logs
    
    - **Set retention policies** — don't keep logs forever
    - **Use structured logging** (JSON) for better querying
    - **Create metric filters** for key events
    - **Use Contributor Insights** for top-N analysis
    
    ### Cost Optimization
    
    - **Delete unused dashboards**
    - **Reduce log retention** for non-critical logs
    - **Avoid high-resolution metrics** unless necessary
    - **Use log subscription filters** instead of polling
    
    ## Troubleshooting
    
    ### Missing Metrics
    
    **Causes:**
    - Service not publishing yet (wait 1-5 minutes)
    - Wrong namespace/dimensions
    - Detailed monitoring not enabled (EC2)
    
    **Debug:**
    
    ```bash
    # List metrics for a namespace
    aws cloudwatch list-metrics \
      --namespace AWS/Lambda \
      --dimensions Name=FunctionName,Value=MyFunction
    ```
    
    ### Alarm Stuck in INSUFFICIENT_DATA
    
    **Causes:**
    - Metric not being published
    - Dimensions mismatch
    - Evaluation period too short
    - Warm-up period still active (ends early once data fills the window unless `OnlyStartEvaluatingAfterWarmUpPeriodEnds=true`)
    - Log alarm: just created/updated (query, schedule, or log group changes reset state), scheduled query role lacks permissions (`EvaluationState` = `EVALUATION_ERROR`, see `StateReason`), or ingestion lag (shift window back with `EndTimeOffset`)
    
    **Debug:**
    
    ```bash
    # Check if metric has data
    aws cloudwatch get-metric-statistics \
      --namespace AWS/Lambda \
      --metric-name Invocations \
      --dimensions Name=FunctionName,Value=MyFunction \
      --start-time $(date -d '1 hour ago' -u +%Y-%m-%dT%H:%M:%SZ) \
      --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
      --period 60 \
      --statistics Sum
    
    # Log alarm: check evaluation state and reason
    aws cloudwatch describe-alarms --alarm-types LogAlarm --alarm-names HighErrorCount \
      --query 'LogAlarms[].[StateValue,EvaluationState,StateReason]'
    ```
    
    ### Log Events Not Appearing
    
    **Causes:**
    - IAM permissions missing
    - CloudWatch Logs agent not running
    - Log group doesn't exist
    
    **Debug:**
    
    ```bash
    # Check log streams
    aws logs describe-log-streams \
      --log-group-name /aws/lambda/MyFunction \
      --order-by LastEventTime \
      --descending \
      --limit 5
    ```
    
    ### High CloudWatch Costs
    
    **Check usage:**
    
    ```bash
    # Get PutLogEvents usage
    aws cloudwatch get-metric-statistics \
      --namespace AWS/Logs \
      --metric-name IncomingBytes \
      --dimensions Name=LogGroupName,Value=/aws/lambda/MyFunction \
      --start-time $(date -d '7 days ago' -u +%Y-%m-%dT%H:%M:%SZ) \
      --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
      --period 86400 \
      --statistics Sum
    ```
    
    ## References
    
    - [CloudWatch User Guide](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/)
    - [CloudWatch Logs User Guide](https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/)
    - [CloudWatch API Reference](https://docs.aws.amazon.com/AmazonCloudWatch/latest/APIReference/)
    - [CloudWatch CLI Reference](https://docs.aws.amazon.com/cli/latest/reference/cloudwatch/)
    - [Logs Insights Query Syntax](https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/CWL_QuerySyntax.html)
    - [boto3 CloudWatch](https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/cloudwatch.html)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related