Claude Skill

backups

Use when designing or auditing a backup-and-restore program that must survive a real disaster: setting defensible RPO/RTO targets, laying out 3-2-1-1-0 copies that are offsite and immutable, wiring point-in-time recovery, and proving restores work on a schedule. NOT tuning Postgr

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download ericrisco-rsc-harness-skills_backups-953fef5.zip · 12 KB
Part of ericrisco/rsc-harness — 46 skills

Install

skills CLI npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/backups
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
Git git clone https://github.com/ericrisco/rsc-harness.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Backups & disaster recovery

Start here

One question decides whether you have backups or a hope: if you restored right now, would it work — and how do you know?

If the answer is "we have a nightly job and it hasn't errored," you do not have backups. You have a job. A backup you have never restored is an untested assertion about the future. This skill exists to turn that assertion into evidence: pick the numbers, lay out the copies so ransomware can't reach them, and run a real restore on a schedule so the answer becomes "yes, we restored it last Tuesday in 41 minutes."

This is the cross-engine strategy + verification skill. Engine-specific SQL knobs (the actual archive_command, VACUUM, schema migrations) belong to ../postgresdb/SKILL.md. Encryption-key management as a general security control belongs to ../secure-coding/SKILL.md.

1. Decide RPO and RTO first — numbers, not adjectives

Everything downstream derives from two numbers. Never start from "what does the tool do."

  • RPO = maximum acceptable data loss, the gap between the disaster and your last good copy. It sets backup frequency. "RPO 1 hour" means a copy at least hourly.
  • RTO = maximum acceptable downtime until you're serving again. It sets topology: a snapshot restore is hours; a warm standby you promote is minutes.

Get a real number from whoever owns the revenue, not "as little as possible." Then map it:

RPO target Required cadence + mechanism
24 h Nightly full/snapshot is enough
~5 min Continuous change-log shipping: WAL / binlog / transaction-log → PITR
~0 Synchronous replica plus PITR (the replica covers hardware loss, PITR covers the bad DELETE)
RTO target Required topology
Hours Restore from snapshot / object-storage backup is fine
Minutes Warm standby or read-replica you promote; pre-provisioned restore target
Seconds Multi-AZ / failover cluster (and you still need backups for logical corruption)

Rule: a replica is not a backup — it faithfully replicates your DROP TABLE in milliseconds. You need point-in-time recovery to step back before the mistake.

2. The 3-2-1-1-0 layout

The classic 3-2-1 rule grew two digits because ransomware now targets the backup infrastructure itself in ~96% of attacks — an attacker who can delete your backups has no reason to fear your backups. The current rule:

  • 3 copies of the data (production + 2 backups). Why: one extra copy still dies with a correlated failure.
  • 2 different media / storage types. Why: a single storage class has a single failure mode.
  • 1 offsite. Why: fire, flood, region outage, or a billing-locked cloud account kills everything in one place. Offsite ≠ another folder on the same host.
  • 1 immutable or offline. Why: a mutable copy an attacker (or a buggy script) can delete is not a safety net. See §5.
  • 0 verification errors. Why: an unverified backup is Schrödinger's backup — both good and corrupt until you restore it. See §6–7.

Map it concretely: production Postgres (copy 1) → pgBackRest repo on local disk, different media (copy 2) → same repo replicated to an S3 bucket in another region with Object Lock (offsite + immutable) → pgbackrest verify + a scheduled test restore (the 0).

3. Pick the mechanism per data store

PITR availability is the column that decides whether you can hit a sub-hour RPO.

Data store Tool PITR? The one gotcha
Managed DB (RDS, Cloud SQL, Supabase, Neon) Built-in automated backups Yes Set retention to your real window; know the ceiling (RDS caps at 35 days). Add a cross-region/cross-account copy you control.
Self-hosted PostgreSQL pgBackRest or WAL-G + WAL archiving Yes Single-threaded archive_command falling behind = WAL storm (see §4). Use archive-async=y.
MySQL / MariaDB Percona XtraBackup + binlog Yes (binlog) mysqldump alone has no PITR — you also need --single-transaction and binlog shipping.
Redis RDB snapshot + AOF Partial Cache-first caveat: if Redis is just a cache, a backup may be pointless; if it's a system of record, enable AOF everysec and treat it like a DB.
Files / app data + object storage restic or BorgBackup → Object-Lock bucket n/a (versions) Both dedup + AES-256 encrypt. restic = concurrent shared repos + faster restore; Borg = smaller repos but one exclusive lock per repo.
App config / secrets Versioned + encrypted store n/a Back these up too, or your restored DB has nothing to connect to. Key must live somewhere the disaster doesn't.

Concrete copy-paste config (pgBackRest stanza, RDS CLI, XtraBackup, restic/Borg) lives in references/engine-recipes.md.

4. PITR, generically

The model is the same for every engine that supports it: a base/full backup + a continuous change log replayed forward to a target time.

base backup (T0) ──► change log: WAL / binlog / txn-log ──► replay to "2026-06-02 13:59:00"

You restore the base, then replay the log up to one second before the bad event. That second is why PITR beats snapshots for logical corruption.

  • Amazon RDS: a daily automated snapshot plus transaction logs shipped to S3 every 5 minutes, restorable to any second within a retention window of up to 35 days. Snapshots after the first are incremental. Trap: AWS Backup does not support a PITR restore into another region — you can copy the backup cross-region, but the restore-to-a-point-in-time happens in the original region. Plan failover accordingly.
  • Self-hosted Postgres traps (both kill PITR silently):
    • WAL storm: a single-threaded archive_command can't keep up under write load, pg_wal/ fills the disk, and Postgres stops accepting writes. Use async/parallel archiving (archive-async=y, --process-max tuned to disk count — it's I/O-bound, not CPU-bound).
    • Retention trap: expiring WAL too aggressively means the chain to your target time is gone and PITR fails. Retain WAL at least as long as your oldest restorable full backup.

Engine commands are in references/engine-recipes.md. For the Postgres-internals side of this (writing the archive_command, tuning), use ../postgresdb/SKILL.md.

5. Offsite + immutability

  • Immutable = Object Lock / WORM on the backup bucket: copies cannot be edited or deleted until the retention window expires, even by a root key. Enable it on a dedicated backup bucket.
  • Retention ≥ 90 days. A 30-day default is often too short: malware commonly dwells for weeks before triggering, so a 30-day lock can expire on the very copies that predate the infection. 90+ days outlives typical dwell time.
  • Offsite means a different blast radius, not a different folder. Different region, and ideally a different account/project with separate credentials — so a compromised application key cannot reach in and delete the backups. The app that writes backups should not hold the key that can delete them.

6. The tested restore — this is the part everyone skips

The number-one reason PITR fails in a real disaster is that it was never tested. Schedule restores like you schedule backups.

Test Cadence Scope
File / single-object restore Monthly Pull one file/table back, in an isolated env, confirm integrity
Application recovery Quarterly Stand up the app against the restored data, run smoke queries
Full-environment failover Annually Rebuild the whole stack from backups in an isolated account/region

Rules:

  • Always restore into an isolated environment — never over production, never sharing its credentials.
  • Record the actual recovery time every run and update your RTO to that reality. An RTO of "4 hours" that you've never measured is fiction.
  • A restore test that "looks fine" isn't done until you've run a validation query / row-count / checksum that proves the data is correct, not just present.

The fill-in-the-blank restore runbook, the scheduled-test checklist, and the actual-RTO log format are in references/restore-runbook.md.

7. Verify & monitor

  • Integrity-verify the backups themselves: restic check / borg check / pgbackrest verify re-read and checksum chunks. A backup that won't pass check won't restore.
  • Alert on three conditions: a backup job failed, the newest good backup is older than your RPO (backup-age check — silence is the dangerous failure mode), and a restore test is overdue.
  • Wiring those alerts into a metrics/paging stack is a monitoring concern — own the what to alert on here, hand the how to ../monitoring/SKILL.md.

Anti-patterns

Anti-pattern Why it bites Do instead
Backups on the same host/disk as the source One disk/host failure takes the source and the backup together Offsite copy, different blast radius (§2)
Never test-restored The restore fails for the first time during the disaster Monthly/quarterly/annual scheduled restores (§6)
"The replica is our backup" It replicates your DROP TABLE faithfully Replica for RTO, PITR for the bad write (§1)
Mutable bucket the app key can delete Ransomware/compromised key wipes backups too Object Lock + separate credentials (§5)
30-day retention only Malware dwell time outlasts the lock Retention ≥ 90 days (§5)
pg_dump/dump to /tmp Reboot or full disk silently loses it; no PITR Dedicated repo + WAL/binlog shipping (§3–4)
RPO promised, cadence can't meet it "RPO 5 min" with a nightly job = up to 24 h loss Derive cadence from RPO first (§1)
Single-threaded archive_command under load WAL storm fills pg_wal, writes stop Async/parallel archiving (§4)
Encrypted backup, key only in the vault that's gone Backup is recoverable but unreadable Store the key in a separate blast radius (§3)
Trusting managed snapshots without knowing the ceiling RDS caps at 35 days; longer needs an exported copy Know the retention ceiling, export beyond it (§3)
Counting "job succeeded" as success The job wrote a corrupt/empty archive check/verify + a real test restore (§6–7)

scripts/verify.sh <path-to-policy-or-runbook> lints a produced backup-policy/runbook artifact for the five pillars (RPO, RTO, offsite, immutability, scheduled-restore + verification). It is a completeness lint, not a backup executor.

Files (rsc-harness)
  • evals
    • cases.yaml 3.3 KB
      skill: backups
      
      should_trigger:
        - prompt: "We don't really have backups beyond a cron job nobody has checked. Design something real."
          why: "Designing a backup-and-restore program from scratch is the core purpose."
        - prompt: "What RPO and RTO should we be targeting for our main database?"
          why: "Picking defensible RPO/RTO targets and deriving cadence/topology is section 1 of the skill."
        - prompt: "Set up point-in-time recovery for our self-hosted Postgres so we can roll back to before a bad delete."
          why: "PITR strategy (base + WAL replay to a target time) is owned here; engine internals route to postgresdb but the strategy is here."
        - prompt: "Our backups live on the same server as the database — is that actually fine?"
          why: "Non-obvious: sounds like infra trivia, but it is a direct 3-2-1 / blast-radius violation this skill exists to catch."
        - prompt: "Com provo que les còpies de seguretat funcionen de debò abans que passi un desastre?"
          why: "Catalan for 'how do I prove the backups really work before a disaster' — the tested-restore discipline, the load-bearing part of the skill."
        - prompt: "Are our backups any good? Audit what we have."
          why: "Auditing an existing setup against the 3-2-1-1-0 pillars and tested-restore gap is an explicit use case."
      
      should_not_trigger:
        - prompt: "Write the archive_command and tune autovacuum for our Postgres instance."
          route_to: postgresdb
          why: "Engine-specific internals (archive_command, VACUUM tuning) are postgresdb's domain, not cross-engine backup strategy."
        - prompt: "Plan an expand-contract schema migration with a safe rollback path."
          route_to: db-migrations
          why: "Reversible DDL / expand-contract rollout is a schema-migration concern owned by db-migrations, not backup-and-restore."
        - prompt: "Page me when a backup job fails and build a dashboard of backup status."
          route_to: monitoring
          why: "The skill decides WHAT to alert on; wiring the actual alerting/dashboards/paging is monitoring's domain."
        - prompt: "Rotate the KMS key we use for encryption at rest and manage key access."
          route_to: secure-coding
          why: "Key management as a general security control belongs to secure-coding, not backup strategy."
        - prompt: "Set up cross-region read replicas to reduce read latency for European users."
          route_to: postgresdb
          why: "Replication for performance/availability topology is a database concern; a replica is explicitly not a backup."
      
      capability:
        - scenario: "User says: 'We run a nightly mysqldump to /backups on the same box as the MySQL server, no offsite copy, we've never restored it, and we tell customers our RPO is 1 hour.' Audit it."
          must_include:
            - "Flags the single-location violation of 3-2-1 (backup on the same host/blast radius as the source)"
            - "Flags the absence of an offsite and immutable/Object-Lock copy (3-2-1-1-0 gaps)"
            - "Flags that the restore has never been tested — the backup is unverified"
            - "Flags the RPO mismatch: a nightly dump cannot deliver a 1-hour RPO (up to ~24h data loss); mysqldump alone has no PITR"
            - "Proposes a concrete fix: offsite immutable copy with retention >=90 days, separate credentials, and binlog/PITR to meet the RPO"
            - "Proposes a scheduled restore-test cadence (monthly file / quarterly app / annual failover) and recording the actual RTO"
      
    • README.md 850 B
      # Evals for the `backups` skill
      
      These cases are routing and capability checks, not an executable test suite — there is no runner to invoke. Read `cases.yaml` against the skill: for each `should_trigger` prompt, confirm the SKILL.md frontmatter `description` and body would plausibly fire the skill; for each `should_not_trigger` prompt, confirm the named `route_to` sibling is the better home and that the description's `NOT ... (that is postgresdb)` boundary holds. For the `capability` case, run the scenario mentally (or hand it to an agent loaded with SKILL.md) and check the response covers every line in `must_include` — the 3-2-1-1-0 gaps, the untested-restore flag, the RPO-vs-cadence mismatch, and the concrete fix with a scheduled restore-test cadence. Treat a missing rubric item as a failing eval and patch the body, not the rubric.
      
  • references
    • engine-recipes.md 4.7 KB
      # Engine recipes
      
      Copy-paste starting points per data store. Adapt paths, regions, and retention to your RPO/RTO. Each block ends with the one gotcha that bites in production.
      
      ## Self-hosted PostgreSQL — pgBackRest
      
      Current standard: full weekly + differential daily, zstd compression, parallel by disk count, async archiving.
      
      ```ini
      # /etc/pgbackrest/pgbackrest.conf
      [global]
      repo1-path=/var/lib/pgbackrest
      repo1-retention-full=4
      repo1-type=s3
      repo1-s3-bucket=my-pgbackrest-backups
      repo1-s3-region=eu-west-1
      compress-type=zst
      process-max=4          # tune to DISK count, not CPU — backup is I/O-bound
      archive-async=y        # avoids the WAL storm under write load
      
      [main]
      pg1-path=/var/lib/postgresql/18/main
      ```
      
      ```bash
      pgbackrest --stanza=main stanza-create
      pgbackrest --stanza=main --type=full backup       # weekly
      pgbackrest --stanza=main --type=diff backup        # daily
      pgbackrest --stanza=main check                     # archive + backup reachable
      pgbackrest --stanza=main verify                    # re-checksum the repo (the "0")
      ```
      
      PITR restore to a target time (restore into an ISOLATED target, never over prod):
      
      ```bash
      pgbackrest --stanza=main --type=time \
        --target="2026-06-02 13:59:00+00" \
        --target-action=promote restore
      ```
      
      Gotcha: `repo1-retention-full` controls how much WAL is kept. Expire it too aggressively and the chain to your target time is gone — PITR silently fails. Keep WAL at least as long as the oldest full you might restore.
      
      WAL-G is the cloud-native alternative (env-driven, smaller footprint):
      
      ```bash
      export WALG_S3_PREFIX=s3://my-bucket/wal-g
      export AWS_REGION=eu-west-1
      wal-g backup-push /var/lib/postgresql/18/main
      wal-g backup-fetch /var/lib/postgresql/18/restore LATEST
      # then recovery_target_time in postgresql.conf + restore_command = 'wal-g wal-fetch %f %p'
      ```
      
      ## Amazon RDS — PITR via AWS CLI
      
      ```bash
      aws rds restore-db-instance-to-point-in-time \
        --source-db-instance-identifier prod-db \
        --target-db-instance-identifier prod-db-restore-test \
        --restore-time 2026-06-02T13:59:00Z \
        --db-subnet-group-name isolated-restore-subnet
      ```
      
      Gotcha: this restores into the **same region**. AWS Backup can *copy* a backup cross-region but does **not** support a PITR restore in another region. Retention ceiling is 35 days; for longer, export a manual snapshot.
      
      ## MySQL / MariaDB — XtraBackup + binlog
      
      ```bash
      # Physical hot backup
      xtrabackup --backup --target-dir=/backup/base --user=backup --password=...
      xtrabackup --prepare --target-dir=/backup/base
      
      # PITR = restore base, then replay binlog to a stop point
      mysqlbinlog --stop-datetime="2026-06-02 13:59:00" \
        /var/lib/mysql/binlog.000042 | mysql -u root -p
      ```
      
      Gotcha: `mysqldump` alone gives no PITR. If you must use it, add `--single-transaction --source-data=2` and ship binlogs separately, or you can only restore to the dump instant. (`--master-data` is the deprecated alias — replaced by `--source-data` since 8.0.26.)
      
      ## Redis — RDB + AOF
      
      ```conf
      # redis.conf — only if Redis is a system of record, not just a cache
      save 900 1
      appendonly yes
      appendfsync everysec   # ~1s RPO; 'always' is durable but slow
      ```
      
      Gotcha: if Redis is purely a cache, backing it up is usually wasted effort — rebuild from the source of truth instead. Decide which Redis you have first.
      
      ## Files / app data — restic to an Object-Lock bucket
      
      ```bash
      export RESTIC_REPOSITORY=s3:s3.amazonaws.com/my-immutable-backups/app
      export RESTIC_PASSWORD_FILE=/etc/restic/pass   # key lives OUTSIDE this bucket
      restic init
      restic backup /srv/app/data --exclude-file=/etc/restic/excludes
      restic check --read-data-subset=10%            # integrity verify
      restic restore latest --target /restore/test   # into isolated path
      ```
      
      Borg equivalent (smaller repos, one exclusive lock per repo):
      
      ```bash
      borg init --encryption=repokey-blake2 /backup/repo
      borg create /backup/repo::'{hostname}-{now}' /srv/app/data
      borg check /backup/repo
      borg extract /backup/repo::archive-name        # restore
      ```
      
      Gotcha: both encrypt with AES-256, but the key/password must live in a *different* blast radius than the data. An encrypted backup whose key died with the source is unrecoverable.
      
      ## S3 Object Lock (the immutability layer)
      
      ```bash
      aws s3api create-bucket --bucket my-immutable-backups --object-lock-enabled-for-bucket \
        --region eu-west-1 --create-bucket-configuration LocationConstraint=eu-west-1
      aws s3api put-object-lock-configuration --bucket my-immutable-backups \
        --object-lock-configuration '{"ObjectLockEnabled":"ENABLED","Rule":{"DefaultRetention":{"Mode":"COMPLIANCE","Days":90}}}'
      ```
      
      `COMPLIANCE` mode cannot be shortened or bypassed even by root — that is the point. Use a separate IAM principal for writing backups vs. one (rarely used) for lifecycle administration.
      
    • restore-runbook.md 3 KB
      # Restore runbook + scheduled-test discipline
      
      A backup you can't restore under pressure is worthless. Fill this in *before* the disaster, and rehearse it on the cadence below.
      
      ## Restore runbook (fill-in template)
      
      ```text
      SERVICE / DATA STORE: ____________________
      RUNBOOK OWNER:        ____________________   LAST REHEARSED: __________
      
      PRECONDITIONS
      - [ ] Isolated restore target provisioned (NOT production, separate credentials)
      - [ ] Backup source identified: repo / bucket = __________________
      - [ ] Decryption key available from: __________ (must NOT be the lost system)
      - [ ] Target recovery time / point: __________________
      
      STEPS
      1. Provision the isolated target ............... cmd: ______________
      2. Pull the base/full backup ................... cmd: ______________
      3. Replay change log to target time (PITR) ..... cmd: ______________
      4. Bring the store up in isolation ............. cmd: ______________
      5. VALIDATE (do not skip):
         - row counts vs. expected ................... query: ___________
         - latest known-good record present ......... query: ___________
         - app smoke test against restored data ..... ______________
      6. Cut over (only if this is a real recovery) .. ______________
      
      SIGN-OFF
      - Restore succeeded: Y / N
      - Data validated correct (not just present): Y / N
      - Actual recovery time: __________  (start ____ → serving ____)
      - Issues / follow-ups: ______________________________
      ```
      
      ## Scheduled restore-test checklist
      
      Run on the cadence from SKILL.md §6. Tick every box or the test does not count.
      
      ```text
      [ ] Restore ran in an ISOLATED environment (no prod credentials, no shared bucket write access)
      [ ] Restored to a specific point in time (not just "latest"), to exercise PITR
      [ ] Ran integrity verify first (restic/borg check, pgbackrest verify)
      [ ] Ran a validation query proving data is CORRECT, not merely present
      [ ] Recorded the ACTUAL recovery time (see log below)
      [ ] Compared actual time against the stated RTO; flagged if RTO is now fiction
      [ ] Logged any step that was slower/harder than the runbook claims; updated the runbook
      [ ] Torn down the isolated target afterwards (cost + stale-data hygiene)
      ```
      
      Cadence reminder:
      - Monthly: file/single-object/single-table restore.
      - Quarterly: full application recovery (app stands up against restored data).
      - Annually: full-environment failover in an isolated account/region.
      
      ## Actual-RTO log
      
      Keep this table in the repo. The RTO you publish should be the worst recent *measured* number, not an estimate.
      
      | Date | Test type | Target point | Restored OK? | Validated correct? | Actual RTO | Notes |
      |------|-----------|--------------|--------------|--------------------|------------|-------|
      | 2026-06-01 | file | latest-1d | Y | Y | 00:12 | clean |
      | 2026-06-30 | app recovery | 2026-06-30 02:00 | Y | Y | 00:48 | DNS step slow |
      | 2026-09-30 | full failover | 2026-09-29 23:00 | Y | Y | 03:41 | RTO target was 2h — fix |
      
      If a row says "Restored OK? N" or "Validated correct? N", that is an incident, not a footnote. Treat it like a production outage you got to schedule.
      
  • scripts
    • verify.sh 2.4 KB
      #!/usr/bin/env bash
      # verify.sh — completeness lint for a backup-policy / restore-runbook artifact.
      #
      # Read-only. It does NOT run, create, or touch any backup. It only greps a
      # produced artifact for the five pillars this skill insists on:
      #   1. an RPO value           4. immutability (Object Lock / WORM)
      #   2. an RTO value           5. a scheduled restore-test cadence + a verification step
      #   3. an offsite copy
      #
      # Usage:
      #   scripts/verify.sh <path-to-policy-or-runbook.md>
      #
      # Exit codes:
      #   0  all pillars present, OR no/empty target given (nothing to lint = no false failure)
      #   1  the target exists with content but is missing one or more pillars
      #   2  the target path was given but does not exist / is unreadable
      
      set -euo pipefail
      
      target="${1:-}"
      
      # No target, or an empty target → nothing to lint. Clean exit, never a false failure.
      if [[ -z "${target}" ]]; then
        echo "verify.sh: no artifact path given — nothing to lint (ok)"
        exit 0
      fi
      
      if [[ ! -e "${target}" ]]; then
        echo "verify.sh: '${target}' does not exist" >&2
        exit 2
      fi
      
      if [[ ! -r "${target}" ]]; then
        echo "verify.sh: '${target}' is not readable" >&2
        exit 2
      fi
      
      if [[ ! -s "${target}" ]]; then
        echo "verify.sh: '${target}' is empty — nothing to lint (ok)"
        exit 0
      fi
      
      # Case-insensitive presence check helper.
      has() { grep -Eiq -- "$1" "${target}"; }
      
      missing=()
      
      # Pillar 1: RPO — the acronym, with a number somewhere near it is ideal but the
      # label alone is the minimum signal.
      has 'rpo' || missing+=("RPO target (recovery point objective)")
      
      # Pillar 2: RTO.
      has 'rto' || missing+=("RTO target (recovery time objective)")
      
      # Pillar 3: an offsite copy.
      has 'off[ -]?site|different (region|account)|cross[ -]region' \
        || missing+=("offsite copy (different region/account)")
      
      # Pillar 4: immutability.
      has 'immutab|object[ -]?lock|worm' \
        || missing+=("immutability (Object Lock / WORM)")
      
      # Pillar 5a: a scheduled restore-test cadence.
      has 'restore[ -](test|drill)|test(ed)? restore|monthly|quarterly|annual' \
        || missing+=("scheduled restore-test cadence")
      
      # Pillar 5b: a verification / integrity step.
      has 'verif|integrity|checksum|\bcheck\b|0 (verification )?errors' \
        || missing+=("verification / integrity step")
      
      if [[ ${#missing[@]} -eq 0 ]]; then
        echo "verify.sh: '${target}' covers all five pillars — ok"
        exit 0
      fi
      
      echo "verify.sh: '${target}' is missing the following pillar(s):" >&2
      for m in "${missing[@]}"; do
        echo "  - ${m}" >&2
      done
      exit 1
      
  • SKILL.md 10.8 KB
    ---
    name: backups
    description: "Use when designing or auditing a backup-and-restore program that must survive a real disaster: setting defensible RPO/RTO targets, laying out 3-2-1-1-0 copies that are offsite and immutable, wiring point-in-time recovery, and proving restores work on a schedule. NOT tuning Postgres internals or writing the archive_command (that is `postgresdb`)."
    tags: [backups, disaster-recovery, pitr, rpo-rto, restore]
    recommends: [postgresdb, secure-coding, monitoring]
    origin: risco
    ---
    
    # Backups & disaster recovery
    
    ## Start here
    
    One question decides whether you have backups or a hope: **if you restored right now, would it work — and how do you know?**
    
    If the answer is "we have a nightly job and it hasn't errored," you do not have backups. You have a job. A backup you have never restored is an untested assertion about the future. This skill exists to turn that assertion into evidence: pick the numbers, lay out the copies so ransomware can't reach them, and run a *real* restore on a schedule so the answer becomes "yes, we restored it last Tuesday in 41 minutes."
    
    This is the cross-engine **strategy + verification** skill. Engine-specific SQL knobs (the actual `archive_command`, VACUUM, schema migrations) belong to `../postgresdb/SKILL.md`. Encryption-key management as a general security control belongs to `../secure-coding/SKILL.md`.
    
    ## 1. Decide RPO and RTO first — numbers, not adjectives
    
    Everything downstream derives from two numbers. Never start from "what does the tool do."
    
    - **RPO** = maximum acceptable *data loss*, the gap between the disaster and your last good copy. It sets **backup frequency**. "RPO 1 hour" means a copy at least hourly.
    - **RTO** = maximum acceptable *downtime* until you're serving again. It sets **topology**: a snapshot restore is hours; a warm standby you promote is minutes.
    
    Get a real number from whoever owns the revenue, not "as little as possible." Then map it:
    
    | RPO target | Required cadence + mechanism |
    |---|---|
    | 24 h | Nightly full/snapshot is enough |
    | ~5 min | Continuous change-log shipping: WAL / binlog / transaction-log → PITR |
    | ~0 | Synchronous replica **plus** PITR (the replica covers hardware loss, PITR covers the bad `DELETE`) |
    
    | RTO target | Required topology |
    |---|---|
    | Hours | Restore from snapshot / object-storage backup is fine |
    | Minutes | Warm standby or read-replica you promote; pre-provisioned restore target |
    | Seconds | Multi-AZ / failover cluster (and you still need backups for logical corruption) |
    
    Rule: a replica is **not** a backup — it faithfully replicates your `DROP TABLE` in milliseconds. You need point-in-time recovery to step *back before* the mistake.
    
    ## 2. The 3-2-1-1-0 layout
    
    The classic 3-2-1 rule grew two digits because ransomware now targets the backup infrastructure itself in ~96% of attacks — an attacker who can delete your backups has no reason to fear your backups. The current rule:
    
    - **3** copies of the data (production + 2 backups). Why: one extra copy still dies with a correlated failure.
    - **2** different media / storage types. Why: a single storage class has a single failure mode.
    - **1** offsite. Why: fire, flood, region outage, or a billing-locked cloud account kills everything in one place. Offsite ≠ another folder on the same host.
    - **1** immutable or offline. Why: a mutable copy an attacker (or a buggy script) can delete is not a safety net. See §5.
    - **0** verification errors. Why: an unverified backup is Schrödinger's backup — both good and corrupt until you restore it. See §6–7.
    
    Map it concretely: production Postgres (copy 1) → pgBackRest repo on local disk, different media (copy 2) → same repo replicated to an S3 bucket in another region with **Object Lock** (offsite + immutable) → `pgbackrest verify` + a scheduled test restore (the 0).
    
    ## 3. Pick the mechanism per data store
    
    PITR availability is the column that decides whether you can hit a sub-hour RPO.
    
    | Data store | Tool | PITR? | The one gotcha |
    |---|---|---|---|
    | Managed DB (RDS, Cloud SQL, Supabase, Neon) | Built-in automated backups | Yes | Set retention to your real window; know the ceiling (RDS caps at 35 days). Add a cross-region/cross-account copy you control. |
    | Self-hosted PostgreSQL | pgBackRest or WAL-G + WAL archiving | Yes | Single-threaded `archive_command` falling behind = WAL storm (see §4). Use `archive-async=y`. |
    | MySQL / MariaDB | Percona XtraBackup + binlog | Yes (binlog) | `mysqldump` alone has no PITR — you also need `--single-transaction` and binlog shipping. |
    | Redis | RDB snapshot + AOF | Partial | Cache-first caveat: if Redis is just a cache, a backup may be pointless; if it's a system of record, enable AOF `everysec` and treat it like a DB. |
    | Files / app data + object storage | restic or BorgBackup → Object-Lock bucket | n/a (versions) | Both dedup + AES-256 encrypt. restic = concurrent shared repos + faster restore; Borg = smaller repos but one exclusive lock per repo. |
    | App config / secrets | Versioned + encrypted store | n/a | Back these up too, or your restored DB has nothing to connect to. Key must live somewhere the disaster doesn't. |
    
    Concrete copy-paste config (pgBackRest stanza, RDS CLI, XtraBackup, restic/Borg) lives in `references/engine-recipes.md`.
    
    ## 4. PITR, generically
    
    The model is the same for every engine that supports it: **a base/full backup + a continuous change log replayed forward to a target time.**
    
    ```text
    base backup (T0) ──► change log: WAL / binlog / txn-log ──► replay to "2026-06-02 13:59:00"
    ```
    
    You restore the base, then replay the log up to one second before the bad event. That second is why PITR beats snapshots for logical corruption.
    
    - **Amazon RDS**: a daily automated snapshot plus transaction logs shipped to S3 **every 5 minutes**, restorable to any second within a retention window of up to **35 days**. Snapshots after the first are incremental. Trap: **AWS Backup does not support a PITR restore *into another region*** — you can copy the backup cross-region, but the restore-to-a-point-in-time happens in the original region. Plan failover accordingly.
    - **Self-hosted Postgres traps** (both kill PITR silently):
      - *WAL storm*: a single-threaded `archive_command` can't keep up under write load, `pg_wal/` fills the disk, and Postgres **stops accepting writes**. Use async/parallel archiving (`archive-async=y`, `--process-max` tuned to disk count — it's I/O-bound, not CPU-bound).
      - *Retention trap*: expiring WAL too aggressively means the chain to your target time is gone and PITR fails. Retain WAL at least as long as your oldest restorable full backup.
    
    Engine commands are in `references/engine-recipes.md`. For the Postgres-internals side of this (writing the `archive_command`, tuning), use `../postgresdb/SKILL.md`.
    
    ## 5. Offsite + immutability
    
    - **Immutable = Object Lock / WORM** on the backup bucket: copies cannot be edited or deleted until the retention window expires, even by a root key. Enable it on a dedicated backup bucket.
    - **Retention ≥ 90 days.** A 30-day default is often too short: malware commonly *dwells* for weeks before triggering, so a 30-day lock can expire on the very copies that predate the infection. 90+ days outlives typical dwell time.
    - **Offsite means a different blast radius**, not a different folder. Different region, and ideally a different account/project with **separate credentials** — so a compromised application key cannot reach in and delete the backups. The app that writes backups should not hold the key that can delete them.
    
    ## 6. The tested restore — this is the part everyone skips
    
    The number-one reason PITR fails in a real disaster is that it was **never tested**. Schedule restores like you schedule backups.
    
    | Test | Cadence | Scope |
    |---|---|---|
    | File / single-object restore | Monthly | Pull one file/table back, in an isolated env, confirm integrity |
    | Application recovery | Quarterly | Stand up the app against the restored data, run smoke queries |
    | Full-environment failover | Annually | Rebuild the whole stack from backups in an isolated account/region |
    
    Rules:
    - **Always restore into an isolated environment** — never over production, never sharing its credentials.
    - **Record the *actual* recovery time every run** and update your RTO to that reality. An RTO of "4 hours" that you've never measured is fiction.
    - A restore test that "looks fine" isn't done until you've run a validation query / row-count / checksum that proves the data is *correct*, not just present.
    
    The fill-in-the-blank restore runbook, the scheduled-test checklist, and the actual-RTO log format are in `references/restore-runbook.md`.
    
    ## 7. Verify & monitor
    
    - **Integrity-verify the backups themselves**: `restic check` / `borg check` / `pgbackrest verify` re-read and checksum chunks. A backup that won't pass `check` won't restore.
    - **Alert on three conditions**: a backup job *failed*, the newest good backup is *older than your RPO* (backup-age check — silence is the dangerous failure mode), and a *restore test is overdue*.
    - Wiring those alerts into a metrics/paging stack is a monitoring concern — own the *what to alert on* here, hand the *how* to `../monitoring/SKILL.md`.
    
    ## Anti-patterns
    
    | Anti-pattern | Why it bites | Do instead |
    |---|---|---|
    | Backups on the same host/disk as the source | One disk/host failure takes the source and the backup together | Offsite copy, different blast radius (§2) |
    | Never test-restored | The restore fails for the first time *during* the disaster | Monthly/quarterly/annual scheduled restores (§6) |
    | "The replica is our backup" | It replicates your `DROP TABLE` faithfully | Replica for RTO, PITR for the bad write (§1) |
    | Mutable bucket the app key can delete | Ransomware/compromised key wipes backups too | Object Lock + separate credentials (§5) |
    | 30-day retention only | Malware dwell time outlasts the lock | Retention ≥ 90 days (§5) |
    | `pg_dump`/dump to `/tmp` | Reboot or full disk silently loses it; no PITR | Dedicated repo + WAL/binlog shipping (§3–4) |
    | RPO promised, cadence can't meet it | "RPO 5 min" with a nightly job = up to 24 h loss | Derive cadence from RPO *first* (§1) |
    | Single-threaded `archive_command` under load | WAL storm fills `pg_wal`, writes stop | Async/parallel archiving (§4) |
    | Encrypted backup, key only in the vault that's gone | Backup is recoverable but unreadable | Store the key in a separate blast radius (§3) |
    | Trusting managed snapshots without knowing the ceiling | RDS caps at 35 days; longer needs an exported copy | Know the retention ceiling, export beyond it (§3) |
    | Counting "job succeeded" as success | The job wrote a corrupt/empty archive | `check`/`verify` + a real test restore (§6–7) |
    
    `scripts/verify.sh <path-to-policy-or-runbook>` lints a produced backup-policy/runbook artifact for the five pillars (RPO, RTO, offsite, immutability, scheduled-restore + verification). It is a completeness lint, not a backup executor.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related