Claude Cursor GitHub Copilot Skill

aws-serverless-production-readiness

Review AWS Lambda-centered serverless workloads for production readiness across execution roles, event sources, retries, DLQs/destinations, concurrency, idempotency, observability, deployment safety, performance, cost, and rollback. Prefer event-driven architecture for EventBridg

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_aws_aws-serverless-production-readiness-febe32a.zip · 6 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/aws/aws-serverless-production-readiness
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

AWS Serverless Production Readiness

Purpose

Act as the AWS serverless production-readiness reviewer who assumes retries, concurrency, and event semantics will punish vague design.

When to use

Use this skill for:

  • Lambda production readiness, performance, security, concurrency, or observability review
  • event-driven architecture using SQS, SNS, EventBridge, Step Functions, API Gateway, or DynamoDB streams
  • DLQ, retry, timeout, idempotency, or poison-message questions
  • serverless deployment, rollback, alias, versioning, or canary-release design

Lean operating rules

  • Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in references/official-sources.md; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing.
  • Separate confirmed facts from inference. If state was not queried or shown, say so.
  • Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
  • Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
  • Load references only when needed; do not pull all deep guidance into short answers.

References

Load these only when needed:

  • Workflow and output contract — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
  • Safety checklist — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
  • Official sources — use when grounding AWS service behavior or checking the detailed source list.
  • Lambda Event Production Readiness Guide — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.

Response minimum

Return, at minimum:

  • the scoped target and evidence level,
  • the main risks or control gaps,
  • the safest next actions,
  • validation or rollback notes where relevant,
  • the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
  • references
    • lambda-event-production-readiness.md 3.4 KB
      # Lambda Event Production Readiness Guide
      
      Use this reference for production readiness reviews of Lambda-centered serverless workloads, event source mappings, retries, DLQs/destinations, concurrency, idempotency, observability, API Gateway, EventBridge, SQS/SNS, Step Functions, and DynamoDB stream integrations.
      
      ## What people get wrong
      
      The lazy story is:
      
      > Lambda scales and retries automatically, so production readiness is mostly IAM and monitoring.
      
      Wrong. Serverless failures come from event semantics: duplicate delivery, poison messages, hot partitions, retry amplification, timeout mismatch, concurrency starvation, and silent data loss.
      
      Common bad assumptions:
      
      - At-least-once delivery can be ignored if the function is fast.
      - DLQ presence proves recoverability.
      - Reserved concurrency is only a cost control.
      - API Gateway timeout and Lambda timeout can be tuned independently without user impact.
      - EventBridge/SNS/SQS retries are harmless.
      - CloudWatch logs are enough observability.
      
      ## Serverless failure modes
      
      - Non-idempotent handlers double-charge, duplicate writes, or corrupt state during retry.
      - SQS visibility timeout, batch size, partial batch response, and function timeout are inconsistent.
      - Async Lambda destinations or DLQs are absent, unmonitored, or unreadable by responders.
      - Reserved/provisioned concurrency starves critical functions or allows noisy neighbors to dominate.
      - EventBridge rule pattern is too broad, causing fanout storms or unexpected consumers.
      - Step Functions retry/catch hides failed business states or creates unbounded downstream calls.
      
      ## Minimum safe workflow
      
      1. Map every trigger, event schema, retry policy, destination, DLQ, and downstream dependency.
      2. Confirm idempotency key, deduplication behavior, ordering needs, and poison-message handling.
      3. Check timeout and concurrency alignment across API Gateway, Lambda, SQS visibility, Step Functions, and downstream services.
      4. Verify observability: RED/USE metrics, structured logs, traces, alarms, DLQ depth, throttles, iterator age, and business KPIs.
      5. Review deployment safety: versions, aliases, CodeDeploy/SAM traffic shifting, rollback alarms, and previous version target.
      6. Identify cost and quota risks: concurrency, payload size, retention, logs, retries, and downstream throttling.
      7. Return readiness gaps with evidence level and approval-gated remediation.
      
      ## Verification targets
      
      - Lambda timeout, memory, ephemeral storage, architecture, environment, layers, versions, aliases, and concurrency settings
      - event source mapping: batch size, max batching window, filters, partial batch response, enabled state, and starting position
      - SQS/SNS/EventBridge retry, DLQ, destination, archive/replay, and fanout configuration
      - Step Functions state machine retries, catches, timeouts, task tokens, and compensation paths
      - IAM execution role, resource policies, KMS, VPC config, and secrets access
      - CloudWatch metrics/alarms for Errors, Throttles, Duration, IteratorAge, ConcurrentExecutions, DLQ depth, and custom business metrics
      
      ## When to push back
      
      Push back if the user asks to:
      
      - declare production-ready without idempotency and retry evidence
      - remove DLQs or alarms to simplify deployment
      - raise concurrency limits without downstream capacity proof
      - ignore duplicate/out-of-order event handling
      - log raw event payloads containing sensitive data
      - deploy alias/version changes without rollback target and alarms
      
    • official-sources.md 1.9 KB
      # Official sources
      
      Use this reference only when you need source grounding for AWS service behavior or the detailed source list.
      
      ## AWS documentation
      
      Use these as starting points, not as proof of the user's live AWS state:
      - https://docs.aws.amazon.com/lambda/latest/dg/durable-execution-sdk-retries.html
      - https://docs.aws.amazon.com/lambda/latest/dg/governance-observability.html
      - https://docs.aws.amazon.com/lambda/latest/dg/best-practices.html
      - https://docs.aws.amazon.com/serverless/latest/devguide/serverless-samples.html
      
      ## Grounding rule
      
      Official documentation explains AWS service behavior. It does not prove the user's current account, Region, quota, resource configuration, IAM boundary, pricing, entitlement, or operational state. Prefer read-only AWS MCP or CLI evidence, repository evidence, or sanitized user-provided evidence for current-state claims.
      
      ## Current MCP/documentation refresh (2026-06-02)
      
      Service facts from official docs:
      - Lambda durable function retry guidance covers step retries, invocation retries, backend retries, exponential backoff, CloudWatch monitoring, and retry best practices.
      - Lambda governance/observability guidance covers visibility into configurations, compliance, function boundaries through Security Hub CSPM, dashboards, tagging, and owner outreach.
      
      Sampled live evidence:
      - Read-only regional availability sampling reported Lambda, API Gateway, and Step Functions as `isAvailableIn` in `us-east-1`, `us-west-2`, `eu-west-1`, and `ap-southeast-1`.
      - Sampled APIs `Lambda+GetFunction` and `SFN+DescribeStateMachine` were reported `isAvailableIn` in those regions.
      
      Review implications:
      - Production readiness requires concurrency, timeout/memory, retry/DLQ/destination behavior, idempotency, observability, IAM, secrets, deployment/rollback, cost guardrails, and failure-mode tests.
      - Serverless service availability does not prove workload readiness or correct event-source semantics.
      
    • safety-checklist.md 1.4 KB
      # Safety checklist
      
      Use this reference before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
      
      ## Non-negotiables
      
      - Never ask users to paste secrets, access keys, session tokens, private keys, customer identifiers, or sensitive account data into chat.
      - Use read-only AWS MCP or read-only AWS CLI evidence for live state when available; otherwise use repository evidence, sanitized user evidence, or official documentation and label the evidence level.
      - Do not invent account IDs, ARNs, Regions, resource names, quotas, prices, or live configuration state.
      - Require explicit user approval before privileged, destructive, traffic-changing, cost-changing, or production-impacting actions.
      - Use current official AWS documentation for service behavior when the answer depends on AWS service details.
      - Keep remediation least-privilege, reversible, and scoped to the requested workload or account boundary.
      
      ## Stress checks
      
      - What can expose data?
      - What can escalate privilege?
      - What can break production or block rollback?
      - What can create unbounded cost?
      - What compliance or audit evidence is missing?
      - What rollback or validation path is unproven?
      
      ## Evidence labels
      
      Use `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. Documentation alone never proves the user's live AWS state.
      
    • workflow-and-output.md 2.1 KB
      # Workflow and output contract
      
      Use this reference only when performing the full review, implementation guidance, incident triage, or production-readiness pass.
      
      ## Review domains
      
      Check these areas before giving a verdict:
      - Execution role least privilege, environment variables, secrets, VPC access, and resource policies
      - Concurrency, timeouts, retries, event source mappings, DLQs/destinations, and idempotency
      - Metrics, logs, traces, alarms, dashboards, and failure-mode evidence
      - Deployment strategy, aliases/versions, rollback, cold starts, package risk, and cost controls
      
      ## Safe workflow
      
      1. **Frame scope**
         - Workload/account/Region/environment:
         - Business criticality and owner:
         - Data classification and compliance driver:
         - Required outcome:
         - Explicit non-goals:
      2. **Collect evidence**
         - Prefer read-only AWS MCP or read-only AWS CLI evidence for current-state claims when available.
         - Otherwise inspect repository IaC/config, sanitized user evidence, or official AWS docs.
         - Label each finding as `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`.
      3. **Stress-test risk**
         - What can expose data?
         - What can escalate privilege?
         - What can break production or block rollback?
         - What can create unbounded cost?
         - What evidence is missing?
      4. **Recommend the smallest safe action**
         - Prefer narrow scope, staged rollout, validation, and rollback.
         - If the safest action is to stop and gather evidence, say that plainly.
      
      ## Output contract
      
      Return this structure:
      ```markdown
      # AWS Serverless Production Readiness: <scope>
      ## Executive verdict
      - Status: READY / READY WITH RISKS / NOT READY / NEEDS EVIDENCE
      - Biggest risk:
      - Evidence level:
      ## Scope and assumptions
      - Confirmed:
      - Unknown:
      - Out of scope:
      ## Findings
      | Severity | Finding | Evidence | Why it matters | Minimum safe action |
      |---|---|---|---|---|
      ## Recommended actions
      1. <action> — owner: <owner>, validation: <check>, rollback: <rollback>
      ## Validation
      - Commands or checks:
      - Expected result:
      ## Residual risk
      - <risk or explicit none>
      ```
      
  • metadata.json 1.1 KB
    {
      "id": "aws-serverless-production-readiness",
      "name": "AWS Serverless Production Readiness",
      "type": "skill",
      "provider": "aws",
      "harnesses": [
        "codex",
        "claude-code",
        "cursor",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Review AWS Lambda and serverless workloads for IAM, concurrency, event sources, retries, DLQs, observability, secrets, performance, cost, and rollback readiness.",
      "source_type": "original",
      "official_docs": [
        "https://docs.aws.amazon.com/lambda/latest/dg/durable-execution-sdk-retries.html",
        "https://docs.aws.amazon.com/lambda/latest/dg/governance-observability.html",
        "https://docs.aws.amazon.com/lambda/latest/dg/best-practices.html",
        "https://docs.aws.amazon.com/serverless/latest/devguide/serverless-samples.html"
      ],
      "security_notes": "Do not approve serverless workloads that lack least-privilege execution roles, retry/DLQ semantics, concurrency controls, observability, idempotency, and rollback evidence.",
      "last_verified": "2026-06-02",
      "path": "skills/aws/aws-serverless-production-readiness",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.4"
    }
    
  • SKILL.md 2.8 KB
    ---
    name: aws-serverless-production-readiness
    description: Review AWS Lambda-centered serverless workloads for production readiness across execution roles, event sources, retries, DLQs/destinations, concurrency, idempotency, observability, deployment safety, performance, cost, and rollback. Prefer event-driven architecture for EventBridge/SNS/SQS/Step Functions system design, and DynamoDB/RDS skills for data-store performance.
    allowed-tools: Read Grep Glob
    metadata:
      author: "github: VincentChuWaiChow"
      version: "0.1.4"
      updated: "2026-06-02"
      category: platform
    ---
    
    # AWS Serverless Production Readiness
    
    ## Purpose
    
    Act as the AWS serverless production-readiness reviewer who assumes retries, concurrency, and event semantics will punish vague design.
    
    ## When to use
    
    Use this skill for:
    
    - Lambda production readiness, performance, security, concurrency, or observability review
    - event-driven architecture using SQS, SNS, EventBridge, Step Functions, API Gateway, or DynamoDB streams
    - DLQ, retry, timeout, idempotency, or poison-message questions
    - serverless deployment, rollback, alias, versioning, or canary-release design
    
    ## Lean operating rules
    
    - Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in `references/official-sources.md`; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing.
    - Separate confirmed facts from inference. If state was not queried or shown, say so.
    - Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
    - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
    - Load references only when needed; do not pull all deep guidance into short answers.
    
    ## References
    
    Load these only when needed:
    
    - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
    - [Safety checklist](references/safety-checklist.md) — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
    - [Official sources](references/official-sources.md) — use when grounding AWS service behavior or checking the detailed source list.
    - [Lambda Event Production Readiness Guide](references/lambda-event-production-readiness.md) — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.
    
    ## Response minimum
    
    Return, at minimum:
    
    - the scoped target and evidence level,
    - the main risks or control gaps,
    - the safest next actions,
    - validation or rollback notes where relevant,
    - the assumptions or blockers that prevent stronger conclusions.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related