aws-essentials
Use when standing up the core AWS surface a small product needs: hardening a fresh account, a private S3 bucket, encrypted RDS Postgres, ECS Fargate vs EC2, CloudFront + OAC, or scoping an IAM policy to least privilege. NOT the CI pipeline that ships the container (that is `deplo
Install
npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/aws-essentials
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
git clone https://github.com/ericrisco/rsc-harness.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
AWS essentials — the core surface a small product needs, secured from the first command
Stand up the foundational AWS services a one-app product actually uses — IAM, S3, ECS Fargate, RDS, CloudFront — with the security defaults that prevent the incidents (public buckets, god-mode app roles, long-lived keys, unencrypted databases). Pick the right tier for small (one app, modest traffic, two engineers), provision it correctly, and wire it without foot-guns.
account hardening → IAM (roles + scoped policies) → S3 (private) / RDS (encrypted) / ECS Fargate → CloudFront (OAC) → infra exists, wired, least-privilege
Service decision table
| Need | Use | Use instead if |
|---|---|---|
| Object/file storage (uploads, assets, backups) | S3 (private bucket) | — |
| Relational data (users, orders, anything with joins) | RDS (Postgres/MySQL) | key-value / serverless access pattern → dynamodb skill |
| Long-running container/API | ECS Fargate | steady ~70%+ CPU 24/7 → EC2 launch type with Savings Plans/Spot; GPU or >120 GB RAM → EC2 |
| Static site / SPA + public assets | S3 + CloudFront | edge functions / global KV → ../cloudflare/SKILL.md |
| Tiny app, no real AWS need yet | be honest → ../vercel/SKILL.md or ../deployment/SKILL.md |
you genuinely need AWS primitives → stay here |
Fargate cold start is ~30–60 s; for spiky/variable small-product load its operational simplicity (no host patching, per-second billing, strong task isolation) wins. EC2 launch type only earns its host-management cost at sustained high utilization. (ECS Managed Instances, Sept 2025, is a newer hybrid — out of scope for a first setup.)
Account zero-day hardening checklist
Do this once, before anything else. Each line has a reason; skip none.
- Enable MFA on the root user — prefer a passkey / security key (phishing-resistant). Root with no MFA is the single highest-blast-radius account.
- Stop using root for daily work — root is for the handful of root-only tasks (close account, change support plan). Everything else uses an IAM identity.
- Create an admin identity via IAM Identity Center (or an assumable admin role). Humans log in to a role with temporary creds, not a static user.
- Delete any root access keys — root should have zero access keys. If one exists, it is a liability with no upside.
- No long-lived IAM-user access keys for apps or CI — apps use task roles, CI uses OIDC (see
../deployment/SKILL.md). - Set your home region and create resources there consistently (one exception below: ACM certs for CloudFront must be in
us-east-1). - Create a billing/cost budget alarm — a misconfigured resource should page you, not surprise you on the invoice.
IAM — least privilege without guessing
Two principal types. IAM users = long-lived humans/keys; avoid them for workloads. Roles
= an identity something assumes to get temporary credentials — this is what ECS tasks, Lambda,
CI, and federated humans use. Default to roles: a temporary credential beats a long-lived AKIA…
key that lives forever in a .env. (IAM governs who may call AWS APIs; access control inside
your app code is ../secure-coding/SKILL.md.)
A policy is a JSON document. The four parts that matter:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::acme-uploads/users/*",
"Condition": { "StringEquals": { "aws:SecureTransport": "true" } }
}]
}
Effect (Allow/Deny) · Action (which API calls) · Resource (which ARNs) · Condition
(extra constraints). The whole game is keeping Action and Resource narrow.
The workflow — start broad, then tighten (do not hand-author from zero):
- Attach the closest AWS managed policy to get the app working.
- Let it run, then use IAM Access Analyzer → generate policy from CloudTrail activity to produce a fine-grained policy from what it actually called.
- Replace the managed policy with the generated one.
- Validate with Access Analyzer (runs 100+ policy checks) and review findings.
- Periodically prune with last-accessed data — remove permissions nothing has used.
// Bad — one leak owns the account
{ "Effect": "Allow", "Action": "*", "Resource": "*" }
// Good — exactly what this service does, on exactly its resources
{ "Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::acme-uploads/users/*" }
An ECS task needs a trust policy (who may assume the role) plus a permission policy (what it may do). Trust policy for a task role:
{ "Version": "2012-10-17",
"Statement": [{ "Effect": "Allow",
"Principal": { "Service": "ecs-tasks.amazonaws.com" },
"Action": "sts:AssumeRole" }] }
Policy JSON anatomy, condition keys, Access Analyzer CLI flow, and copy-ready scoped templates
(S3 one-prefix R/W, read one Secrets Manager secret, write CloudWatch logs, ECS trust) →
references/iam-least-privilege.md.
S3 — private object storage
Create a bucket. The defaults are already what you want:
aws s3api create-bucket \
--bucket acme-uploads \
--region eu-west-1 \
--create-bucket-configuration LocationConstraint=eu-west-1
# Since Apr 2023, this bucket is already: Block Public Access ON (all four),
# Object Ownership = bucket-owner-enforced (ACLs disabled), SSE-S3 on every object.
Keep all of that. Do not re-enable ACLs; do not turn off Block Public Access. Grant access two ways instead: a bucket policy (resource-side, e.g. allow one CloudFront distribution) or an IAM identity policy (subject-side, e.g. the task role above). For browser uploads, hand the client a presigned URL so the app never proxies the bytes and the bucket stays private:
aws s3 presign s3://acme-uploads/users/123/avatar.png --expires-in 900
Bad: set bucket to public-read so the <img> tags work
Good: bucket stays private → presigned URLs for direct upload/download,
and CloudFront + OAC for public-read web content (see below)
Compute — ECS Fargate first
Run the container on Fargate (rationale in the decision table). The mistake that costs an afternoon every time:
Task role vs execution role — they are different.
- Execution role: lets ECS itself pull the image from ECR and push logs to CloudWatch. Start from the managed
AmazonECSTaskExecutionRolePolicy.- Task role: the identity your application code assumes at runtime to call AWS (read the S3 bucket, read a secret). This is where your scoped least-privilege policy goes. Putting app permissions on the execution role (or vice-versa) is the classic "works in console, 403 at runtime" bug.
Network layout: tasks in private subnets, a load balancer (ALB) in public subnets, egress via
NAT. The DB and tasks never get public IPs. EC2 launch type only if you hit the steady-utilization
or hardware thresholds above. Full task-def + service CLI path lives in ../deployment/SKILL.md
(that skill owns the ship step); this skill owns the roles and networking it runs on.
RDS — managed relational DB
aws rds create-db-instance \
--db-instance-identifier acme-prod \
--engine postgres \
--db-instance-class db.t4g.small \
--allocated-storage 20 \
--storage-encrypted --kms-key-id <your-rds-cmk> \
--multi-az \
--no-publicly-accessible \
--master-username acme --manage-master-user-password \
--vpc-security-group-ids sg-app-db
Encrypt at create time — you cannot encrypt an existing instance in place. Storage encryption (AES-256 via KMS) must be set at creation; it then covers backups, read replicas, and snapshots. To fix an unencrypted instance you must snapshot → copy-snapshot with encryption → restore (Multi-AZ clusters can't even do that directly). Prefer a customer-managed KMS key dedicated to RDS.
Two more non-negotiables: the DB security group references the app's security group, never
0.0.0.0/0 (a DB open to the internet is a breach, not a convenience); credentials live in
Secrets Manager with managed rotation (--manage-master-user-password above), never in task
env vars. Use --multi-az for production HA. Full recipe (SG wiring, Secrets Manager rotation,
connecting from ECS) → references/rds-cloudfront-recipes.md. Schema, indexes and query tuning
once the instance exists → ../postgresdb/SKILL.md.
CloudFront + OAC — public web content, private bucket
To serve S3 content publicly, do not make the bucket public. Put CloudFront in front and grant it via Origin Access Control (OAC) — the modern replacement for the legacy OAI:
- OAC uses short-term, rotated credentials and a resource-based bucket policy scoped to the distribution ARN; CloudFront→S3 is always HTTPS with "Sign requests" (the default).
- OAC supports SSE-KMS origins and all regions. OAI is legacy — never reach for it.
- The bucket keeps Block Public Access on; you grant only the distribution, by bucket policy.
- Set the viewer protocol policy to redirect-to-HTTPS; ACM cert for a custom domain must be
in
us-east-1.
Full CLI: create OAC → distribution → S3 bucket policy JSON → invalidations → custom domain →
references/rds-cloudfront-recipes.md.
Anti-patterns
| Anti-pattern | Why it bites | Fix |
|---|---|---|
| Public-read S3 bucket | Anyone enumerates/downloads everything; classic breach headline | Keep Block Public Access on; presigned URLs or CloudFront+OAC |
| Re-enabling S3 ACLs | Brings back the confused-deputy/ownership mess April-2023 defaults removed | Leave bucket-owner-enforced; use bucket/IAM policies |
AdministratorAccess on an app/task role |
One leaked task credential = full account compromise | Scope to the exact actions+ARNs the service uses |
"Action": "*", "Resource": "*" policy |
Same blast radius, just hand-written | Generate from CloudTrail via Access Analyzer; validate |
IAM-user access keys in app/.env/commit |
Long-lived, never rotated, leak forever | Task role (app) / OIDC (CI) — temporary creds |
| Unencrypted RDS | Can't encrypt later without snapshot-copy-restore downtime | --storage-encrypted at create, customer-managed KMS key |
DB security group open to 0.0.0.0/0 |
Database directly reachable from the internet | SG references the app SG only; --no-publicly-accessible |
| CloudFront with OAI | Legacy; misses SSE-KMS, weaker credential model | Use OAC, bucket policy scoped to the distribution ARN |
| Secrets in task env vars | Leak via logs, console, task definition history | Secrets Manager + managed rotation, injected at runtime |
| Root user for daily ops | Highest blast radius, no per-action attribution | Root only for root-only tasks; admin via Identity Center |
| No MFA on root | One phished password = total account loss | Passkey/security-key MFA on root and every human |
| Confusing task role and execution role | App gets 403 at runtime, or ECS can't pull the image | Execution = pull image/logs; task = app's runtime perms |
Files (rsc-harness)
-
evals
-
cases.yaml 2.9 KB
skill: aws-essentials should_trigger: - prompt: "Set up an S3 bucket for user profile uploads on AWS" why: Core S3 provisioning — the bucket-defaults + presigned-URL path this skill owns. - prompt: "This IAM role has AdministratorAccess, tighten it to least privilege" why: Cloud IAM scoping. Non-obvious that this is aws-essentials and not secure-coding — it is the cloud-identity surface, not app-code access control. - prompt: "Spin up an encrypted Postgres on RDS with Multi-AZ" why: RDS provisioning, including the encrypt-at-create gotcha this skill warns about. - prompt: "monta CloudFront delante del meu bucket S3 privat" why: Catalan phrasing for the CloudFront+OAC-over-private-S3 recipe. - prompt: "Should my container run on ECS Fargate or EC2 for a low-traffic app?" why: Compute decision. Non-obvious — the word AWS never appears, but ECS implies it and the Fargate-vs-EC2 tradeoff is core here. - prompt: "My S3 bucket is public and I don't know why — make it private and still serve the images" why: Symptom phrasing; routes to keeping Block Public Access on + CloudFront/OAC instead of public-read. - prompt: "Lock down the security group on our RDS instance, it's open to the world" why: DB-SG-references-app-SG hardening, a named anti-pattern in this skill. should_not_trigger: - prompt: "Write the Dockerfile and GitHub Actions workflow to deploy to ECS" route_to: deployment why: Containerization + CI pipeline (incl. OIDC to ECR), not infra provisioning. This skill ends where the container starts shipping. - prompt: "Review this login handler for broken access control" route_to: secure-coding why: App-code OWASP review, not cloud IAM least-privilege. - prompt: "Model a single-table DynamoDB schema for my app" route_to: dynamodb why: NoSQL data modeling, not the AWS core-setup surface (this skill picks RDS for relational and points at dynamodb otherwise). - prompt: "Optimize this slow Postgres query and add the right indexes" route_to: postgresdb why: Query/schema tuning on an existing DB, not RDS provisioning. - prompt: "Set up Cloudflare Workers and a CDN for my static site" route_to: cloudflare why: Different provider's edge/CDN, not AWS CloudFront. capability: - scenario: "Provision storage and a CDN for a small product's user-uploaded images on AWS, with least privilege and no shortcuts." must_include: - Private S3 bucket with Block Public Access kept ON (no public-read ACL, no re-enabled ACLs). - Access via IAM/bucket policy + presigned URLs or CloudFront — never a public bucket. - CloudFront uses OAC (not the legacy OAI) and serves over HTTPS. - IAM policy scoped to the specific bucket ARN + specific actions (no "Action":"*" on "Resource":"*"). - No long-lived access keys — a task role / temporary credentials instead. - Acknowledges encryption at rest (SSE-S3 default on the bucket). -
README.md 747 B
# Evals — aws-essentials These cases are LLM routing and quality checks, not executable AWS calls — nothing here touches a real account or needs credentials. Run them through the repo's eval harness: `should_trigger` and `should_not_trigger` feed the skill's `description` + body to the router and assert it selects (or correctly declines, routing to the named real sibling) this skill; `capability` prompts the agent with the scenario and grades the produced answer against the `must_include` rubric (private bucket, OAC-not-OAI, scoped IAM, no long-lived keys, encryption acknowledged). The static linter `scripts/verify.sh` is separate and runs standalone over a directory of policy/config files — it needs no harness and no AWS access.
-
-
references
-
iam-least-privilege.md 4.8 KB
# IAM least-privilege — anatomy, workflow, and copy-ready templates Depth offloaded from `SKILL.md`. Everything here keeps `Action` and `Resource` narrow and prefers temporary credentials over long-lived keys. ## Policy JSON anatomy ```json { "Version": "2012-10-17", "Statement": [ { "Sid": "ReadWriteOwnPrefix", "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"], "Resource": "arn:aws:s3:::acme-uploads/users/*", "Condition": { "Bool": { "aws:SecureTransport": "true" } } } ] } ``` - `Version` is always the literal `2012-10-17` (a policy-language date, not "use the latest"). - `Sid` is an optional human label — use it; future-you reads policies more than writes them. - An explicit `Deny` always wins over any `Allow`. Use `Deny` for guardrails, not for the everyday "what may this role do" — that should be a tight `Allow`. ## Condition keys worth knowing | Key | Use | Example | |---|---|---| | `aws:SecureTransport` | force TLS | `"Bool": {"aws:SecureTransport": "true"}` | | `aws:SourceArn` | confused-deputy guard on resource policies | restrict S3 bucket policy to one CloudFront distribution ARN | | `aws:PrincipalTag/team` | attribute-based access (ABAC) | `"StringEquals": {"aws:PrincipalTag/team": "payments"}` | | `s3:prefix` | limit which keys a `ListBucket` can see | `"StringLike": {"s3:prefix": ["users/${aws:userid}/*"]}` | ## The tighten-with-Access-Analyzer flow Hand-authoring a minimal policy from scratch means guessing every API call a service makes — you will be wrong and either over-grant or break it. Let CloudTrail tell you the truth. ```bash # 1. Generate a fine-grained policy from what the role ACTUALLY called (CloudTrail-backed). aws accessanalyzer start-policy-generation \ --policy-generation-details '{"principalArn":"arn:aws:iam::123456789012:role/acme-task"}' \ --cloud-trail-details '{ "trails":[{"cloudTrailArn":"arn:aws:cloudtrail:eu-west-1:123456789012:trail/acme","allRegions":true}], "accessRole":"arn:aws:iam::123456789012:role/AccessAnalyzerCT", "startTime":"2026-05-01T00:00:00Z" }' aws accessanalyzer get-generated-policy --job-id <job-id> # poll, then copy the JSON # 2. Validate any policy against 100+ checks before you attach it. aws accessanalyzer validate-policy \ --policy-type IDENTITY_POLICY \ --policy-document file://acme-task-policy.json # Review findings: SECURITY_WARNING / ERROR / SUGGESTION. Fix before attaching. # 3. Periodically prune: which permissions has nobody used? aws iam generate-service-last-accessed-details --arn arn:aws:iam::123456789012:role/acme-task ``` Replace the broad managed policy you started with by the generated, validated one. Re-run the last-accessed prune every quarter. ## ECS task: two roles, two policies ```json // Trust policy — who may assume this role (same for task and execution role) { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Principal": { "Service": "ecs-tasks.amazonaws.com" }, "Action": "sts:AssumeRole" }] } ``` - **Execution role** permission policy: start from the AWS managed `AmazonECSTaskExecutionRolePolicy` (pull from ECR + write CloudWatch logs). Add `secretsmanager:GetSecretValue` *here* only for secrets injected by ECS at container start. - **Task role** permission policy: your application's runtime grants — the scoped templates below. ## Copy-ready scoped templates **S3 — read/write exactly one prefix:** ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:GetObject","s3:PutObject","s3:DeleteObject"], "Resource": "arn:aws:s3:::acme-uploads/users/*" }, { "Effect": "Allow", "Action": "s3:ListBucket", "Resource": "arn:aws:s3:::acme-uploads", "Condition": { "StringLike": { "s3:prefix": ["users/*"] } } } ] } ``` **Secrets Manager — read exactly one secret:** ```json { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Action": "secretsmanager:GetSecretValue", "Resource": "arn:aws:secretsmanager:eu-west-1:123456789012:secret:acme/prod/db-*" }] } ``` **CloudWatch Logs — write the app's own log group:** ```json { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Action": ["logs:CreateLogStream","logs:PutLogEvents"], "Resource": "arn:aws:logs:eu-west-1:123456789012:log-group:/ecs/acme:*" }] } ``` **Trust policy for human admin via federation** (Identity Center handles this for you; shown for a self-managed assumable role): ```json { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam::123456789012:root" }, "Action": "sts:AssumeRole", "Condition": { "Bool": { "aws:MultiFactorAuthPresent": "true" } } }] } ``` Note the MFA condition: an assumable role with no MFA requirement is barely better than a static key. Require `aws:MultiFactorAuthPresent` on any human-assumed role. -
rds-cloudfront-recipes.md 4.7 KB
# RDS and CloudFront — end-to-end recipes Depth offloaded from `SKILL.md`. Two complete paths: an encrypted Multi-AZ Postgres wired to an app, and a public CloudFront distribution over a private S3 origin via OAC. ## RDS — encrypted Multi-AZ Postgres, wired to ECS ### 1. Security groups — the DB SG references the app SG, never `0.0.0.0/0` ```bash # App tasks' SG already exists: sg-app. Create the DB SG and allow ONLY the app SG on 5432. aws ec2 create-security-group --group-name acme-db --description "RDS ingress from app only" \ --vpc-id vpc-0abc --query GroupId --output text # -> sg-app-db aws ec2 authorize-security-group-ingress \ --group-id sg-app-db \ --protocol tcp --port 5432 \ --source-group sg-app # source is the SG, not a CIDR — never 0.0.0.0/0 ``` ### 2. Create the instance — encrypted at create time, Multi-AZ, not public ```bash aws rds create-db-instance \ --db-instance-identifier acme-prod \ --engine postgres --engine-version 16 \ --db-instance-class db.t4g.small \ --allocated-storage 20 --storage-type gp3 \ --storage-encrypted --kms-key-id alias/acme-rds \ --multi-az \ --no-publicly-accessible \ --vpc-security-group-ids sg-app-db \ --db-subnet-group-name acme-private \ --master-username acme \ --manage-master-user-password \ --backup-retention-period 7 ``` - `--storage-encrypted` **must** be set now. You cannot encrypt an existing instance in place; the fix is snapshot → `copy-db-snapshot` with `--kms-key-id` → `restore-db-instance-from-db-snapshot`. Multi-AZ *clusters* can't even do that directly. Encryption covers storage, backups, replicas, and snapshots. - `--kms-key-id alias/acme-rds` uses a customer-managed key dedicated to RDS (preferred over the AWS-managed default). - `--manage-master-user-password` puts the master password in Secrets Manager — no plaintext. ### 3. Secrets Manager — rotation + app retrieval ```bash # Find the managed secret ARN RDS created: aws rds describe-db-instances --db-instance-identifier acme-prod \ --query 'DBInstances[0].MasterUserSecret.SecretArn' --output text # Turn on automatic rotation (RDS provides the rotation Lambda for managed secrets): aws secretsmanager rotate-secret --secret-id <arn> \ --rotation-rules '{"AutomaticallyAfterDays": 30}' ``` The ECS **task role** gets `secretsmanager:GetSecretValue` on that exact secret ARN (template in `iam-least-privilege.md`). The app reads the secret at startup — never bake the password into a task-definition env var (it leaks via task-definition history and logs). Connect over TLS. ## CloudFront + OAC over a private S3 origin ### 1. Create the Origin Access Control ```bash aws cloudfront create-origin-access-control --origin-access-control-config '{ "Name": "acme-site-oac", "OriginAccessControlOriginType": "s3", "SigningBehavior": "always", "SigningProtocol": "sigv4" }' # -> note the OAC Id ``` `"SigningBehavior": "always"` is the recommended "Sign requests" default. Never create an `origin-access-identity` (OAI) — it is legacy. ### 2. Create the distribution pointing at the bucket's regional domain, with the OAC attached Key fields in the distribution config: origin `DomainName` = `acme-site.s3.eu-west-1.amazonaws.com`, `OriginAccessControlId` = the id above, `S3OriginConfig.OriginAccessIdentity` = empty string, and the default cache behavior `ViewerProtocolPolicy` = `redirect-to-https`. ```bash aws cloudfront create-distribution --distribution-config file://dist-config.json # After creation, note the distribution ARN: arn:aws:cloudfront::123456789012:distribution/E123 ``` ### 3. Bucket policy — grant ONLY this distribution, bucket stays private Block Public Access stays **on**. Access is granted purely by this resource policy, scoped to the distribution ARN via `aws:SourceArn` (confused-deputy guard): ```json { "Version": "2012-10-17", "Statement": [{ "Sid": "AllowCloudFrontOACRead", "Effect": "Allow", "Principal": { "Service": "cloudfront.amazonaws.com" }, "Action": "s3:GetObject", "Resource": "arn:aws:s3:::acme-site/*", "Condition": { "StringEquals": { "AWS:SourceArn": "arn:aws:cloudfront::123456789012:distribution/E123" } } }] } ``` ```bash aws s3api put-bucket-policy --bucket acme-site --policy file://bucket-policy.json ``` ### 4. Invalidations and custom domain ```bash # Bust the cache after a deploy: aws cloudfront create-invalidation --distribution-id E123 --paths "/*" ``` For a custom domain, request the **ACM certificate in `us-east-1`** (CloudFront only reads certs from there, regardless of where your bucket and app live), validate it via DNS, then set the distribution's `Aliases` + `ViewerCertificate.ACMCertificateArn`. Point the domain at the distribution with a DNS alias/`CNAME`.
-
-
scripts
-
verify.sh 4.5 KB
#!/usr/bin/env bash # verify.sh — read-only static lint for AWS IAM/policy/config artifacts. # # Mirrors the SKILL.md anti-patterns table so the advice is enforceable. It scans # JSON / .tf / .yaml / .yml / .sh / .env-ish files under TARGET (default ".") for # dangerous patterns. It is a LINT — no AWS API calls, no credentials, deterministic, # CI-safe. Read-only: it never writes or mutates anything. # # Rules: # 1. Full-admin policy: "Action":"*" together with "Resource":"*" # 2. AdministratorAccess attached/referenced (god-mode managed policy) # 3. Public S3 bucket policy: Effect Allow with "Principal":"*" (or {"AWS":"*"}) # 4. Legacy OAI: origin-access-identity / OriginAccessIdentity (non-empty) # 5. Long-lived keys: AKIA... access-key id, or aws_secret_access_key literal # 6. Open DB ingress: 0.0.0.0/0 on or near port 5432 / 3306 # # Exits 1 with file:line + rule on any hit. Exits 0 on a clean OR empty target. # Usage: verify.sh [TARGET_DIR_OR_FILE] set -uo pipefail TARGET="${1:-.}" fail=0 hit() { printf 'FAIL [%s] %s:%s — %s\n' "$1" "$2" "$3" "$4" >&2; fail=1; } note() { printf '%s\n' "$1"; } if [ ! -e "$TARGET" ]; then note "verify: target does not exist: $TARGET — nothing to check." exit 0 fi # Collect candidate files. No matches => clean/empty => exit 0. files=() if [ -f "$TARGET" ]; then files=("$TARGET") else while IFS= read -r f; do files+=("$f") done < <(find "$TARGET" -type f \ \( -name '*.json' -o -name '*.tf' -o -name '*.yaml' -o -name '*.yml' \ -o -name '*.sh' -o -name '*.env' -o -name '*.tfvars' \) \ -not -path '*/.git/*' 2>/dev/null) fi if [ "${#files[@]}" -eq 0 ]; then note "verify: no AWS policy/config files found under $TARGET — nothing to check." exit 0 fi for f in "${files[@]}"; do # Strip CR so Windows-edited files match cleanly. content=$(tr -d '\r' < "$f") # --- Rule 1: full-admin "*"/"*" (file-level: both appear in the same file) --- if printf '%s' "$content" | grep -Eq '"Action"[[:space:]]*:[[:space:]]*"\*"' \ && printf '%s' "$content" | grep -Eq '"Resource"[[:space:]]*:[[:space:]]*"\*"'; then ln=$(grep -nE '"Action"[[:space:]]*:[[:space:]]*"\*"' "$f" | head -n1 | cut -d: -f1) hit "full-admin" "$f" "${ln:-?}" 'policy grants Action "*" on Resource "*" — scope to specific actions+ARNs' fi # --- Rule 2: AdministratorAccess --- while IFS=: read -r ln _; do [ -n "$ln" ] && hit "admin-access" "$f" "$ln" 'AdministratorAccess referenced — scope an app/task role to least privilege' done < <(grep -nE 'AdministratorAccess' "$f" 2>/dev/null) # --- Rule 3: public S3 / resource policy (Principal "*") --- while IFS=: read -r ln _; do [ -n "$ln" ] && hit "public-principal" "$f" "$ln" 'resource policy with Principal "*" — bucket/resource is public; scope to a specific ARN' done < <(grep -nE '"Principal"[[:space:]]*:[[:space:]]*("\*"|\{[[:space:]]*"AWS"[[:space:]]*:[[:space:]]*"\*")' "$f" 2>/dev/null) # --- Rule 4: legacy OAI (ignore the empty-string OAC form "OriginAccessIdentity":"") --- while IFS=: read -r ln rest; do [ -z "$ln" ] && continue # Skip the legitimate empty OAC form. printf '%s' "$rest" | grep -Eq 'OriginAccessIdentity"[[:space:]]*:[[:space:]]*""' && continue hit "legacy-oai" "$f" "$ln" 'origin-access-identity (OAI) is legacy — use Origin Access Control (OAC)' done < <(grep -nE 'origin-access-identity|OriginAccessIdentity"[[:space:]]*:[[:space:]]*"[^"]+|create-cloud-front-origin-access-identity' "$f" 2>/dev/null) # --- Rule 5: long-lived access keys --- while IFS=: read -r ln _; do [ -n "$ln" ] && hit "long-lived-key" "$f" "$ln" 'looks like an AWS access key id (AKIA…) — use a role / temporary credentials' done < <(grep -nE '\bAKIA[0-9A-Z]{16}\b' "$f" 2>/dev/null) while IFS=: read -r ln _; do [ -n "$ln" ] && hit "long-lived-key" "$f" "$ln" 'aws_secret_access_key literal — secrets belong in Secrets Manager / OIDC, not code' done < <(grep -niE 'aws_secret_access_key[[:space:]]*[:=]' "$f" 2>/dev/null) # --- Rule 6: open DB ingress (0.0.0.0/0 near a DB port) --- while IFS=: read -r ln _; do [ -n "$ln" ] && hit "open-db-sg" "$f" "$ln" 'DB port (5432/3306) ingress from 0.0.0.0/0 — reference the app security group, never the internet' done < <(grep -nE '(5432|3306).*0\.0\.0\.0/0|0\.0\.0\.0/0.*(5432|3306)' "$f" 2>/dev/null) done if [ "$fail" -ne 0 ]; then note "verify: AWS artifact lint FAILED — fix the issues above." exit 1 fi note "verify: all scanned AWS policy/config files pass the lint." exit 0
-
-
SKILL.md 11.5 KB
--- name: aws-essentials description: "Use when standing up the core AWS surface a small product needs: hardening a fresh account, a private S3 bucket, encrypted RDS Postgres, ECS Fargate vs EC2, CloudFront + OAC, or scoping an IAM policy to least privilege. NOT the CI pipeline that ships the container (that is `deployment`), NOT app-code access-control review (that is `secure-coding`)." tags: [aws, cloud, iam, s3, infrastructure] recommends: [deployment, secure-coding, dynamodb, postgresdb] origin: risco --- # AWS essentials — the core surface a small product needs, secured from the first command Stand up the foundational AWS services a one-app product actually uses — IAM, S3, ECS Fargate, RDS, CloudFront — with the security defaults that prevent the incidents (public buckets, god-mode app roles, long-lived keys, unencrypted databases). Pick the right tier for *small* (one app, modest traffic, two engineers), provision it correctly, and wire it without foot-guns. ```text account hardening → IAM (roles + scoped policies) → S3 (private) / RDS (encrypted) / ECS Fargate → CloudFront (OAC) → infra exists, wired, least-privilege ``` ## Service decision table | Need | Use | Use instead if | |------|-----|----------------| | Object/file storage (uploads, assets, backups) | **S3** (private bucket) | — | | Relational data (users, orders, anything with joins) | **RDS** (Postgres/MySQL) | key-value / serverless access pattern → `dynamodb` skill | | Long-running container/API | **ECS Fargate** | steady ~70%+ CPU 24/7 → EC2 launch type with Savings Plans/Spot; GPU or >120 GB RAM → EC2 | | Static site / SPA + public assets | **S3 + CloudFront** | edge functions / global KV → `../cloudflare/SKILL.md` | | Tiny app, no real AWS need yet | be honest → `../vercel/SKILL.md` or `../deployment/SKILL.md` | you genuinely need AWS primitives → stay here | Fargate cold start is ~30–60 s; for spiky/variable small-product load its operational simplicity (no host patching, per-second billing, strong task isolation) wins. EC2 launch type only earns its host-management cost at sustained high utilization. (ECS Managed Instances, Sept 2025, is a newer hybrid — out of scope for a first setup.) ## Account zero-day hardening checklist Do this once, before anything else. Each line has a reason; skip none. - [ ] **Enable MFA on the root user** — prefer a passkey / security key (phishing-resistant). Root with no MFA is the single highest-blast-radius account. - [ ] **Stop using root for daily work** — root is for the handful of root-only tasks (close account, change support plan). Everything else uses an IAM identity. - [ ] **Create an admin identity via IAM Identity Center** (or an assumable admin role). Humans log in to a role with temporary creds, not a static user. - [ ] **Delete any root access keys** — root should have zero access keys. If one exists, it is a liability with no upside. - [ ] **No long-lived IAM-user access keys for apps or CI** — apps use task roles, CI uses OIDC (see `../deployment/SKILL.md`). - [ ] **Set your home region** and create resources there consistently (one exception below: ACM certs for CloudFront must be in `us-east-1`). - [ ] **Create a billing/cost budget alarm** — a misconfigured resource should page you, not surprise you on the invoice. ## IAM — least privilege without guessing Two principal types. **IAM users** = long-lived humans/keys; avoid them for workloads. **Roles** = an identity something *assumes* to get temporary credentials — this is what ECS tasks, Lambda, CI, and federated humans use. Default to roles: a temporary credential beats a long-lived `AKIA…` key that lives forever in a `.env`. (IAM governs who may call AWS APIs; access control *inside* your app code is `../secure-coding/SKILL.md`.) A policy is a JSON document. The four parts that matter: ```json { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject"], "Resource": "arn:aws:s3:::acme-uploads/users/*", "Condition": { "StringEquals": { "aws:SecureTransport": "true" } } }] } ``` `Effect` (Allow/Deny) · `Action` (which API calls) · `Resource` (which ARNs) · `Condition` (extra constraints). The whole game is keeping `Action` and `Resource` narrow. **The workflow — start broad, then tighten (do not hand-author from zero):** 1. Attach the closest **AWS managed policy** to get the app working. 2. Let it run, then use **IAM Access Analyzer → generate policy from CloudTrail activity** to produce a fine-grained policy from what it *actually* called. 3. Replace the managed policy with the generated one. 4. **Validate** with Access Analyzer (runs 100+ policy checks) and review findings. 5. Periodically prune with **last-accessed data** — remove permissions nothing has used. ```jsonc // Bad — one leak owns the account { "Effect": "Allow", "Action": "*", "Resource": "*" } // Good — exactly what this service does, on exactly its resources { "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject"], "Resource": "arn:aws:s3:::acme-uploads/users/*" } ``` An ECS task needs a **trust policy** (who may assume the role) plus a **permission policy** (what it may do). Trust policy for a task role: ```json { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Principal": { "Service": "ecs-tasks.amazonaws.com" }, "Action": "sts:AssumeRole" }] } ``` Policy JSON anatomy, condition keys, Access Analyzer CLI flow, and copy-ready scoped templates (S3 one-prefix R/W, read one Secrets Manager secret, write CloudWatch logs, ECS trust) → `references/iam-least-privilege.md`. ## S3 — private object storage Create a bucket. The defaults are already what you want: ```bash aws s3api create-bucket \ --bucket acme-uploads \ --region eu-west-1 \ --create-bucket-configuration LocationConstraint=eu-west-1 # Since Apr 2023, this bucket is already: Block Public Access ON (all four), # Object Ownership = bucket-owner-enforced (ACLs disabled), SSE-S3 on every object. ``` **Keep all of that.** Do not re-enable ACLs; do not turn off Block Public Access. Grant access two ways instead: a **bucket policy** (resource-side, e.g. allow one CloudFront distribution) or an **IAM identity policy** (subject-side, e.g. the task role above). For browser uploads, hand the client a **presigned URL** so the app never proxies the bytes and the bucket stays private: ```bash aws s3 presign s3://acme-uploads/users/123/avatar.png --expires-in 900 ``` ```text Bad: set bucket to public-read so the <img> tags work Good: bucket stays private → presigned URLs for direct upload/download, and CloudFront + OAC for public-read web content (see below) ``` ## Compute — ECS Fargate first Run the container on Fargate (rationale in the decision table). The mistake that costs an afternoon every time: > **Task role vs execution role — they are different.** > - **Execution role**: lets *ECS itself* pull the image from ECR and push logs to CloudWatch. Start from the managed `AmazonECSTaskExecutionRolePolicy`. > - **Task role**: the identity *your application code* assumes at runtime to call AWS (read the S3 bucket, read a secret). This is where your scoped least-privilege policy goes. > Putting app permissions on the execution role (or vice-versa) is the classic "works in console, 403 at runtime" bug. Network layout: **tasks in private subnets**, a load balancer (ALB) in public subnets, egress via NAT. The DB and tasks never get public IPs. EC2 launch type only if you hit the steady-utilization or hardware thresholds above. Full task-def + service CLI path lives in `../deployment/SKILL.md` (that skill owns the ship step); this skill owns the roles and networking it runs on. ## RDS — managed relational DB ```bash aws rds create-db-instance \ --db-instance-identifier acme-prod \ --engine postgres \ --db-instance-class db.t4g.small \ --allocated-storage 20 \ --storage-encrypted --kms-key-id <your-rds-cmk> \ --multi-az \ --no-publicly-accessible \ --master-username acme --manage-master-user-password \ --vpc-security-group-ids sg-app-db ``` > **Encrypt at create time — you cannot encrypt an existing instance in place.** Storage > encryption (AES-256 via KMS) must be set at creation; it then covers backups, read replicas, > and snapshots. To fix an unencrypted instance you must snapshot → copy-snapshot *with* > encryption → restore (Multi-AZ *clusters* can't even do that directly). Prefer a > customer-managed KMS key dedicated to RDS. Two more non-negotiables: the DB security group **references the app's security group**, never `0.0.0.0/0` (a DB open to the internet is a breach, not a convenience); credentials live in **Secrets Manager** with managed rotation (`--manage-master-user-password` above), never in task env vars. Use `--multi-az` for production HA. Full recipe (SG wiring, Secrets Manager rotation, connecting from ECS) → `references/rds-cloudfront-recipes.md`. Schema, indexes and query tuning once the instance exists → `../postgresdb/SKILL.md`. ## CloudFront + OAC — public web content, private bucket To serve S3 content publicly, do **not** make the bucket public. Put CloudFront in front and grant it via **Origin Access Control (OAC)** — the modern replacement for the legacy OAI: - OAC uses short-term, rotated credentials and a resource-based bucket policy scoped to the distribution ARN; CloudFront→S3 is always HTTPS with "Sign requests" (the default). - OAC supports SSE-KMS origins and all regions. **OAI is legacy — never reach for it.** - The bucket keeps Block Public Access **on**; you grant only the distribution, by bucket policy. - Set the viewer protocol policy to **redirect-to-HTTPS**; ACM cert for a custom domain must be in **`us-east-1`**. Full CLI: create OAC → distribution → S3 bucket policy JSON → invalidations → custom domain → `references/rds-cloudfront-recipes.md`. ## Anti-patterns | Anti-pattern | Why it bites | Fix | |---|---|---| | Public-read S3 bucket | Anyone enumerates/downloads everything; classic breach headline | Keep Block Public Access on; presigned URLs or CloudFront+OAC | | Re-enabling S3 ACLs | Brings back the confused-deputy/ownership mess April-2023 defaults removed | Leave bucket-owner-enforced; use bucket/IAM policies | | `AdministratorAccess` on an app/task role | One leaked task credential = full account compromise | Scope to the exact actions+ARNs the service uses | | `"Action": "*", "Resource": "*"` policy | Same blast radius, just hand-written | Generate from CloudTrail via Access Analyzer; validate | | IAM-user access keys in app/`.env`/commit | Long-lived, never rotated, leak forever | Task role (app) / OIDC (CI) — temporary creds | | Unencrypted RDS | Can't encrypt later without snapshot-copy-restore downtime | `--storage-encrypted` at create, customer-managed KMS key | | DB security group open to `0.0.0.0/0` | Database directly reachable from the internet | SG references the app SG only; `--no-publicly-accessible` | | CloudFront with OAI | Legacy; misses SSE-KMS, weaker credential model | Use OAC, bucket policy scoped to the distribution ARN | | Secrets in task env vars | Leak via logs, console, task definition history | Secrets Manager + managed rotation, injected at runtime | | Root user for daily ops | Highest blast radius, no per-action attribution | Root only for root-only tasks; admin via Identity Center | | No MFA on root | One phished password = total account loss | Passkey/security-key MFA on root and every human | | Confusing task role and execution role | App gets 403 at runtime, or ECS can't pull the image | Execution = pull image/logs; task = app's runtime perms |
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.