aws-serverless-production-readiness
Review AWS Lambda-centered serverless workloads for production readiness across execution roles, event sources, retries, DLQs/destinations, concurrency, idempotency, observability, deployment safety, performance, cost, and rollback. Prefer event-driven architecture for EventBridg
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/aws/aws-serverless-production-readiness
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
AWS Serverless Production Readiness
Purpose
Act as the AWS serverless production-readiness reviewer who assumes retries, concurrency, and event semantics will punish vague design.
When to use
Use this skill for:
- Lambda production readiness, performance, security, concurrency, or observability review
- event-driven architecture using SQS, SNS, EventBridge, Step Functions, API Gateway, or DynamoDB streams
- DLQ, retry, timeout, idempotency, or poison-message questions
- serverless deployment, rollback, alias, versioning, or canary-release design
Lean operating rules
- Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in
references/official-sources.md; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - Separate confirmed facts from inference. If state was not queried or shown, say so.
- Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims.
- Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns.
- Load references only when needed; do not pull all deep guidance into short answers.
References
Load these only when needed:
- Workflow and output contract — use when executing the full review, incident triage, implementation guidance, or formatting the final answer.
- Safety checklist — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations.
- Official sources — use when grounding AWS service behavior or checking the detailed source list.
- Lambda Event Production Readiness Guide — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria.
Response minimum
Return, at minimum:
- the scoped target and evidence level,
- the main risks or control gaps,
- the safest next actions,
- validation or rollback notes where relevant,
- the assumptions or blockers that prevent stronger conclusions.
Files (vanguard-frontier-agentic)
-
references
-
lambda-event-production-readiness.md 3.4 KB
# Lambda Event Production Readiness Guide Use this reference for production readiness reviews of Lambda-centered serverless workloads, event source mappings, retries, DLQs/destinations, concurrency, idempotency, observability, API Gateway, EventBridge, SQS/SNS, Step Functions, and DynamoDB stream integrations. ## What people get wrong The lazy story is: > Lambda scales and retries automatically, so production readiness is mostly IAM and monitoring. Wrong. Serverless failures come from event semantics: duplicate delivery, poison messages, hot partitions, retry amplification, timeout mismatch, concurrency starvation, and silent data loss. Common bad assumptions: - At-least-once delivery can be ignored if the function is fast. - DLQ presence proves recoverability. - Reserved concurrency is only a cost control. - API Gateway timeout and Lambda timeout can be tuned independently without user impact. - EventBridge/SNS/SQS retries are harmless. - CloudWatch logs are enough observability. ## Serverless failure modes - Non-idempotent handlers double-charge, duplicate writes, or corrupt state during retry. - SQS visibility timeout, batch size, partial batch response, and function timeout are inconsistent. - Async Lambda destinations or DLQs are absent, unmonitored, or unreadable by responders. - Reserved/provisioned concurrency starves critical functions or allows noisy neighbors to dominate. - EventBridge rule pattern is too broad, causing fanout storms or unexpected consumers. - Step Functions retry/catch hides failed business states or creates unbounded downstream calls. ## Minimum safe workflow 1. Map every trigger, event schema, retry policy, destination, DLQ, and downstream dependency. 2. Confirm idempotency key, deduplication behavior, ordering needs, and poison-message handling. 3. Check timeout and concurrency alignment across API Gateway, Lambda, SQS visibility, Step Functions, and downstream services. 4. Verify observability: RED/USE metrics, structured logs, traces, alarms, DLQ depth, throttles, iterator age, and business KPIs. 5. Review deployment safety: versions, aliases, CodeDeploy/SAM traffic shifting, rollback alarms, and previous version target. 6. Identify cost and quota risks: concurrency, payload size, retention, logs, retries, and downstream throttling. 7. Return readiness gaps with evidence level and approval-gated remediation. ## Verification targets - Lambda timeout, memory, ephemeral storage, architecture, environment, layers, versions, aliases, and concurrency settings - event source mapping: batch size, max batching window, filters, partial batch response, enabled state, and starting position - SQS/SNS/EventBridge retry, DLQ, destination, archive/replay, and fanout configuration - Step Functions state machine retries, catches, timeouts, task tokens, and compensation paths - IAM execution role, resource policies, KMS, VPC config, and secrets access - CloudWatch metrics/alarms for Errors, Throttles, Duration, IteratorAge, ConcurrentExecutions, DLQ depth, and custom business metrics ## When to push back Push back if the user asks to: - declare production-ready without idempotency and retry evidence - remove DLQs or alarms to simplify deployment - raise concurrency limits without downstream capacity proof - ignore duplicate/out-of-order event handling - log raw event payloads containing sensitive data - deploy alias/version changes without rollback target and alarms -
official-sources.md 1.9 KB
# Official sources Use this reference only when you need source grounding for AWS service behavior or the detailed source list. ## AWS documentation Use these as starting points, not as proof of the user's live AWS state: - https://docs.aws.amazon.com/lambda/latest/dg/durable-execution-sdk-retries.html - https://docs.aws.amazon.com/lambda/latest/dg/governance-observability.html - https://docs.aws.amazon.com/lambda/latest/dg/best-practices.html - https://docs.aws.amazon.com/serverless/latest/devguide/serverless-samples.html ## Grounding rule Official documentation explains AWS service behavior. It does not prove the user's current account, Region, quota, resource configuration, IAM boundary, pricing, entitlement, or operational state. Prefer read-only AWS MCP or CLI evidence, repository evidence, or sanitized user-provided evidence for current-state claims. ## Current MCP/documentation refresh (2026-06-02) Service facts from official docs: - Lambda durable function retry guidance covers step retries, invocation retries, backend retries, exponential backoff, CloudWatch monitoring, and retry best practices. - Lambda governance/observability guidance covers visibility into configurations, compliance, function boundaries through Security Hub CSPM, dashboards, tagging, and owner outreach. Sampled live evidence: - Read-only regional availability sampling reported Lambda, API Gateway, and Step Functions as `isAvailableIn` in `us-east-1`, `us-west-2`, `eu-west-1`, and `ap-southeast-1`. - Sampled APIs `Lambda+GetFunction` and `SFN+DescribeStateMachine` were reported `isAvailableIn` in those regions. Review implications: - Production readiness requires concurrency, timeout/memory, retry/DLQ/destination behavior, idempotency, observability, IAM, secrets, deployment/rollback, cost guardrails, and failure-mode tests. - Serverless service availability does not prove workload readiness or correct event-source semantics. -
safety-checklist.md 1.4 KB
# Safety checklist Use this reference before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations. ## Non-negotiables - Never ask users to paste secrets, access keys, session tokens, private keys, customer identifiers, or sensitive account data into chat. - Use read-only AWS MCP or read-only AWS CLI evidence for live state when available; otherwise use repository evidence, sanitized user evidence, or official documentation and label the evidence level. - Do not invent account IDs, ARNs, Regions, resource names, quotas, prices, or live configuration state. - Require explicit user approval before privileged, destructive, traffic-changing, cost-changing, or production-impacting actions. - Use current official AWS documentation for service behavior when the answer depends on AWS service details. - Keep remediation least-privilege, reversible, and scoped to the requested workload or account boundary. ## Stress checks - What can expose data? - What can escalate privilege? - What can break production or block rollback? - What can create unbounded cost? - What compliance or audit evidence is missing? - What rollback or validation path is unproven? ## Evidence labels Use `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. Documentation alone never proves the user's live AWS state. -
workflow-and-output.md 2.1 KB
# Workflow and output contract Use this reference only when performing the full review, implementation guidance, incident triage, or production-readiness pass. ## Review domains Check these areas before giving a verdict: - Execution role least privilege, environment variables, secrets, VPC access, and resource policies - Concurrency, timeouts, retries, event source mappings, DLQs/destinations, and idempotency - Metrics, logs, traces, alarms, dashboards, and failure-mode evidence - Deployment strategy, aliases/versions, rollback, cold starts, package risk, and cost controls ## Safe workflow 1. **Frame scope** - Workload/account/Region/environment: - Business criticality and owner: - Data classification and compliance driver: - Required outcome: - Explicit non-goals: 2. **Collect evidence** - Prefer read-only AWS MCP or read-only AWS CLI evidence for current-state claims when available. - Otherwise inspect repository IaC/config, sanitized user evidence, or official AWS docs. - Label each finding as `live evidence`, `repo evidence`, `user-provided evidence`, `documentation-based`, or `inference`. 3. **Stress-test risk** - What can expose data? - What can escalate privilege? - What can break production or block rollback? - What can create unbounded cost? - What evidence is missing? 4. **Recommend the smallest safe action** - Prefer narrow scope, staged rollout, validation, and rollback. - If the safest action is to stop and gather evidence, say that plainly. ## Output contract Return this structure: ```markdown # AWS Serverless Production Readiness: <scope> ## Executive verdict - Status: READY / READY WITH RISKS / NOT READY / NEEDS EVIDENCE - Biggest risk: - Evidence level: ## Scope and assumptions - Confirmed: - Unknown: - Out of scope: ## Findings | Severity | Finding | Evidence | Why it matters | Minimum safe action | |---|---|---|---|---| ## Recommended actions 1. <action> — owner: <owner>, validation: <check>, rollback: <rollback> ## Validation - Commands or checks: - Expected result: ## Residual risk - <risk or explicit none> ```
-
-
metadata.json 1.1 KB
{ "id": "aws-serverless-production-readiness", "name": "AWS Serverless Production Readiness", "type": "skill", "provider": "aws", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Review AWS Lambda and serverless workloads for IAM, concurrency, event sources, retries, DLQs, observability, secrets, performance, cost, and rollback readiness.", "source_type": "original", "official_docs": [ "https://docs.aws.amazon.com/lambda/latest/dg/durable-execution-sdk-retries.html", "https://docs.aws.amazon.com/lambda/latest/dg/governance-observability.html", "https://docs.aws.amazon.com/lambda/latest/dg/best-practices.html", "https://docs.aws.amazon.com/serverless/latest/devguide/serverless-samples.html" ], "security_notes": "Do not approve serverless workloads that lack least-privilege execution roles, retry/DLQ semantics, concurrency controls, observability, idempotency, and rollback evidence.", "last_verified": "2026-06-02", "path": "skills/aws/aws-serverless-production-readiness", "author": "github: VincentChuWaiChow", "version": "0.1.4" } -
SKILL.md 2.8 KB
--- name: aws-serverless-production-readiness description: Review AWS Lambda-centered serverless workloads for production readiness across execution roles, event sources, retries, DLQs/destinations, concurrency, idempotency, observability, deployment safety, performance, cost, and rollback. Prefer event-driven architecture for EventBridge/SNS/SQS/Step Functions system design, and DynamoDB/RDS skills for data-store performance. allowed-tools: Read Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.4" updated: "2026-06-02" category: platform --- # AWS Serverless Production Readiness ## Purpose Act as the AWS serverless production-readiness reviewer who assumes retries, concurrency, and event semantics will punish vague design. ## When to use Use this skill for: - Lambda production readiness, performance, security, concurrency, or observability review - event-driven architecture using SQS, SNS, EventBridge, Step Functions, API Gateway, or DynamoDB streams - DLQ, retry, timeout, idempotency, or poison-message questions - serverless deployment, rollback, alias, versioning, or canary-release design ## Lean operating rules - Prefer current AWS documentation tools for service behavior. Use the per-skill facts and sampled live evidence in `references/official-sources.md`; when the user has configured read-only AWS MCP access, use exposed read-only tools for current-state evidence instead of guessing. - Separate confirmed facts from inference. If state was not queried or shown, say so. - Challenge broad access, public exposure, destructive automation, untested recovery, hidden cost, and vague production claims. - Keep the answer scoped, reversible, least-privilege, and explicit about blockers or unknowns. - Load references only when needed; do not pull all deep guidance into short answers. ## References Load these only when needed: - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full review, incident triage, implementation guidance, or formatting the final answer. - [Safety checklist](references/safety-checklist.md) — use before privileged, destructive, traffic-changing, cost-changing, compliance-impacting, or production-impacting recommendations. - [Official sources](references/official-sources.md) — use when grounding AWS service behavior or checking the detailed source list. - [Lambda Event Production Readiness Guide](references/lambda-event-production-readiness.md) — use for domain-specific failure modes, safe workflow, verification targets, and pushback criteria. ## Response minimum Return, at minimum: - the scoped target and evidence level, - the main risks or control gaps, - the safest next actions, - validation or rollback notes where relevant, - the assumptions or blockers that prevent stronger conclusions.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.