coolify
Use when self-hosting apps and databases with Coolify on a VPS you own — install, first-admin lockdown, Git-to-deploy (Nixpacks/Dockerfile/compose), managed Postgres/Redis, scheduled S3 backups, domains + auto-SSL. NOT a PaaS someone else runs (that is `railway`), NOT sizing/hard
#devops #deployment
Install
npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/coolify
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
git clone https://github.com/ericrisco/rsc-harness.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Coolify — own-the-box PaaS on your VPS
Coolify is an open-source, self-hostable control plane that turns a plain Linux VPS into a Vercel/Heroku replacement: Git-to-deploy, managed databases, automatic SSL, scheduled backups — for a flat VPS bill instead of per-request metering. This skill is operational. It produces the exact install one-liner, the port matrix, env wiring, a lint-clean compose artifact, and a backup-to-S3 cron with a tested restore. Coolify 4.x (stable in 2026); v4.1 added Railpack, structured audit logging, and a read-only MCP server.
Mental model — read this first
- Coolify is the control plane, not the box. It runs on a VPS you provisioned elsewhere. Hardware sizing, the cloud firewall, SSH hardening, OS patching → that is the VPS layer (hetzner / digitalocean), not this skill. Coolify orchestrates what runs on the box.
- You own the data and the uptime now. No vendor takes nightly snapshots for you. So backups are non-optional and a backup you have never restored is not a backup (see the backup section).
- Everything underneath is Docker. Coolify generates compose/Dockerfile runs and a Traefik proxy. So
every Docker rule still applies — named volumes, healthchecks, pinned images, env-injected secrets.
For image-authoring depth (multi-stage, layer caching) go to the
dockerskill; here we wire it. - First account to register owns the instance — forever. There is no "admin invite" recovery if a stranger registers first. Claim it within seconds of install (see step 4).
Decision — where should this app actually run?
| Option | Who owns the box | Cost shape | Ops burden | Choose when |
|---|---|---|---|---|
| Coolify self-hosted (this skill) | You (your VPS) | Flat VPS/month | You patch + back up | Predictable bill, data sovereignty, many apps on one box |
| Coolify Cloud | You bring the server; they host the control plane | VPS + small SaaS fee | They run the control plane | Want Coolify UX without babysitting the dashboard's own uptime |
| Managed PaaS (railway / render / fly-io / vercel / netlify) | The provider | Per-usage, scales up fast | Near-zero | Spiky scale-to-zero, no box to own — the inverse of Coolify |
Choose Coolify when a few always-on apps on a known-cost box beats per-request metering. If the user wants
git push and never touches a server, that is railway — route there.
Install — the 6-minute path
Provision the box first (this is the hetzner / digitalocean step, not this skill). Floor is 2 CPU / 2 GB RAM / 30 GB disk; Coolify itself idles around 1 GB RAM. For comfortable multi-app production aim for 4 vCPU / 8 GB / 100 GB NVMe. Ubuntu 22.04/24.04 LTS recommended (Debian, Fedora/Alma/Rocky, Alpine, Arch, Raspberry Pi OS 64-bit also supported; non-LTS needs manual steps — see references).
# 1. SSH in AS ROOT. Non-root is not fully supported by the installer — log in as root.
ssh root@<server-ip>
# 2. One-liner install (root). Takes ~2–5 min: installs Docker, sets up Coolify's own stack.
curl -fsSL https://cdn.coollabs.io/coolify/install.sh | bash
# 3. The installer prints the dashboard URL: http://<server-ip>:8000
# 4. CRITICAL — within seconds, open http://<server-ip>:8000 and REGISTER.
# That first account becomes the permanent root admin of the instance. First-come-owns-the-box;
# there is no later "claim ownership" flow. Do not walk away between step 2 and this.
# 5. In the dashboard: set the instance FQDN (coolify.example.com) and force HTTPS, so you stop
# hitting the raw :8000 IP and the dashboard itself gets a real certificate.
Why root: the installer wires Docker and system services; the docs state non-root is not fully supported. Why claim immediately: the registration page is open until someone takes it.
Full install transcript, non-LTS/other-OS steps, dashboard-domain setup, Traefik (default) vs Caddy,
wildcard domains + the DNS-01 challenge, and common SSL failures with their fixes →
references/install-and-proxy.md.
Port & firewall matrix
Set these at the cloud firewall (hetzner/digitalocean layer) AND keep them in mind on the box.
| Port | Purpose | Exposure |
|---|---|---|
| 22 | SSH | Restrict to your IP / VPN; key-only |
| 8000 | Coolify dashboard | Restrict to your IP after setup, or front with the proxy on a FQDN; never leave open to the world |
| 80 | Proxy HTTP (Traefik) → redirects to 443 | Public |
| 443 | Proxy HTTPS, app traffic + Let's Encrypt | Public |
| 6001 | Realtime / websocket | Open as the dashboard needs (same audience as 8000) |
| 6002 | Realtime / terminal websocket | Same as 6001 |
Rule: 80/443 are the only ports the public should reach. 8000/6001/6002 are operator surface — lock them to your IP or put the dashboard behind its FQDN. Why: an open :8000 plus an unclaimed instance is a takeover; an open :8000 on a claimed instance is still your full control plane exposed to brute force.
Deploy an app
Pick the build pack first, then wire source → env → storage → domain.
| Build pack | Use when | Note |
|---|---|---|
| Nixpacks (default) | Standard app, you want zero config | Auto-detects the language/runtime; start here |
| Railpack (v4.1, beta) | Need build-time env vars, config merge, or multi-stage control | Newer; reach for it when Nixpacks can't express the build |
| Dockerfile | You already have a Dockerfile / want full build control | You own the build; pin a base image |
| docker-compose | Multi-service app deployed as a unit | Coolify runs your compose; this is the verify.sh-linted artifact shape |
Then:
- Connect the Git source — GitHub/GitLab app or a deploy key; pick branch; enable auto-deploy on push
(or trigger from CI — that wiring is the
github-actionsskill). - Set env vars as Coolify secrets — inject at runtime via the UI/API; never bake secrets into the image or commit them to the compose. Why: a baked secret ships in every image layer and leaks on pull.
- Add persistent storage — any path that must survive a redeploy (uploads, sqlite) gets a named volume. A container without one loses its writes on every recreate.
- Attach the domain + SSL — add
app.example.com, point its DNS A record at the box (DNS itself is thedomains-dnsskill), and the Traefik proxy provisions Let's Encrypt automatically on ports 80/443.
Build-pack decision deep-dive plus a worked, lint-clean docker-compose.yml (env-ref secrets, named
volumes, healthcheck, pinned image) — the verify.sh target → references/deploy-recipes.md.
Managed databases
Coolify provisions Postgres, MySQL, MariaDB, MongoDB, Redis (and more) in-product — one-click, with generated credentials.
- Connect apps over the internal Docker network hostname, not a public IP. Coolify gives each DB an internal service name; the app reaches it on the private network. Why: zero public attack surface.
- Do NOT publicly expose the DB port unless you genuinely need an external client. If you must, bind it deliberately and firewall it to known IPs — an open 5432/3306 is scanned within minutes.
- Every DB gets a named volume. Without it, recreating the service wipes the data. This is the single most common Coolify data-loss footgun.
This skill provisions and connects the DB; it does not teach SQL, indexing, or query tuning — that is the
postgresdb skill. Per-engine details and connection strings → references/databases-and-backups.md.
Backups to S3 — non-negotiable
You own the data now. Coolify runs scheduled dumps (pg_dump / mysqldump / mongodump) on a cron
expression, stores them locally, and optionally pushes to S3-compatible storage: AWS S3, Cloudflare R2,
Backblaze B2, MinIO, Wasabi.
- Add an S3 destination — endpoint, region, bucket, access key/secret (as Coolify secrets, never in a committed file). Off-box storage is the point: a backup on the same disk dies with the disk.
- Set the cron schedule — e.g.
0 3 * * *for nightly 03:00. Match frequency to how much data you can afford to lose (RPO doctrine across systems is thebackupsskill; here we wire the in-product job). - Set retention — keep N days/copies so the bucket doesn't grow unbounded.
- Run the restore drill — once, before you need it. Download a dump, decompress, replay it into a throwaway DB, confirm the row counts. An untested backup is a guess. Schedule a recurring drill.
S3 destination setup for R2/B2/MinIO/Wasabi/AWS, per-engine dump/restore commands, retention, the
step-by-step restore runbook and its consistency caveats → references/databases-and-backups.md.
Operate the box
- Resource planning. Coolify idles ~1 GB RAM; budget the rest for your apps + databases. Multi-app sweet spot is 4 vCPU / 8 GB / 100 GB NVMe. Watch RAM headroom before adding the next app.
- Updates. Update Coolify from the dashboard's settings; pin/snapshot before a major bump.
- Logs & audit. Per-resource logs live in the UI. v4.1 adds a structured audit log — use it to see who changed what.
- MCP server (v4.1, read-only). Coolify exposes an instance-level MCP server with read-only tools for AI-agent integration. Useful for letting an agent inspect status; it is read-only by design — do not treat it as a deploy channel, and still lock down the network in front of it.
Anti-patterns → STOP
| Rationalization | Reality → STOP |
|---|---|
| "I'll install with my sudo user, root feels risky" | The installer wires Docker + system services and the docs say non-root is not fully supported. SSH in as root for install. |
| "I'll register the admin account later" | First account to hit :8000 owns the instance permanently. A stranger registering first = takeover. Claim it within seconds. |
| "Leave :8000 open, it's password-protected" | That is your full control plane exposed to brute force. Restrict 8000/6001/6002 to your IP; only 80/443 are public. |
"Pin the app image to :latest, it's simpler" |
:latest floats — a silent base change breaks a redeploy you can't reproduce. Pin a tag or digest. |
| "The database doesn't need a named volume yet" | Recreating the service wipes an anonymous volume. Data gone. Named volume from day one. |
| "Put the DB password in the compose so deploys are reproducible" | Secrets in a committed file leak in git history and image layers. Inject via Coolify env/secrets, env-ref only. |
| "Expose 5432 so I can connect from my laptop" | An open DB port is scanned in minutes. Use the internal hostname; if you truly need external access, firewall it to known IPs. |
| "Backups are configured, we're covered" | A backup you've never restored is a guess. Run the restore drill before you need it. |
| "Coolify will harden the server for me" | Coolify is the control plane, not the OS-hardening layer. Firewall/SSH/patching is the VPS skill (hetzner/digitalocean). |
| "Use Coolify because I just want to git push and forget the server" | That's the opposite of own-the-box. Use a managed PaaS — railway. |
Verify
Run scripts/verify.sh against your project (or this skill's references/). It statically lints the
example/your docker-compose.y*ml: fails on a hardcoded secret literal (must be env-ref), a DB service
without a named volume, a missing healthcheck:, or a floating :latest tag on a build-context service;
and confirms the canonical port matrix (8000/80/443/6001/6002) is documented. Read-only, no network, no
live deploy. Exits 0 on a clean/empty target.
Project grounding (02-DOCS + CLAUDE.md)
When this skill runs in a project with a 02-DOCS/ layer (the harness Karpathy wiki), record this
instance's deploy topology there and index it in 02-DOCS/wiki/index.md, so the next agent inherits it
instead of re-deriving it.
- Find the article
02-DOCS/wiki/stack/coolify.md, indexed in02-DOCS/wiki/index.md(the Knowledge map index; rootCLAUDE.mdpoints to it). - If missing or stale, create/update it with the real choices — the box (provider/specs), instance
FQDN, which apps/databases run on it, build packs in use, the backup destination + cron + retention,
and the port/firewall decisions — then index it in
02-DOCS/wiki/index.md(the Knowledge map; rootCLAUDE.mdkeeps only a short pointer to it). - Read it first on every use and stay consistent; when the topology changes, update the article (bump
its
Updateddate) in the same change. Never commit credentials here — record where secrets live, not their values.
No 02-DOCS/ layer? Skip silently. Topology is recorded, not gated — never block the task on this.
Files (rsc-harness)
-
evals
-
cases.yaml 4.1 KB
skill: coolify should_trigger: - prompt: "Install Coolify on my fresh Ubuntu 24.04 VPS and deploy my app from a GitHub repo." why: "Core install + Git deploy path: install one-liner as root, first-admin lockdown, build-pack + source wiring." - prompt: "My Vercel bill exploded this month. I want to move everything onto a box I own and pay a flat rate instead. How do I set that up?" why: "Non-obvious cost/sovereignty trigger — user never says Coolify, but 'box I own + flat rate' instead of a metered PaaS is exactly the own-the-box value prop." - prompt: "Set up nightly Postgres backups to Cloudflare R2 for the database I run in Coolify." why: "In-product scheduled backup to S3-compatible storage (R2) with cron — a named core capability, including the restore-drill rule." - prompt: "Montar mi propio PaaS en un VPS barato con Coolify y desplegar varias apps con SSL automático." why: "Spanish trigger; self-hosting your own PaaS on a cheap VPS with multiple apps and automatic SSL is squarely this skill." - prompt: "Vull autoallotjar les meves apps amb Coolify i una base de dades gestionada al meu servidor." why: "Catalan trigger; self-hosting apps + a managed database on your own server is the provision-and-connect core." - prompt: "I already have a server running. I just want a self-hosted control plane to manage my deploys, env vars and SSL certs from one dashboard." why: "Non-obvious 'control plane on my own server' phrasing — the box exists, the user wants the orchestration layer, which is precisely what Coolify is." should_not_trigger: - prompt: "Which Hetzner plan should I pick and how do I harden SSH and the firewall on the box itself?" route_to: "hetzner" why: "VPS hardware sizing + OS/SSH/firewall hardening is the box layer, not the control plane Coolify runs on." - prompt: "Help me write a multi-stage Dockerfile and optimize my image layer caching." route_to: "docker" why: "Generic Docker image authoring — Coolify runs images but doesn't teach how to author/optimize them." - prompt: "I just want a managed PaaS where I git push and they run it — I don't want to own or maintain any server." route_to: "railway" why: "Push-and-forget managed PaaS is the exact inverse of Coolify's own-the-box premise." - prompt: "My Postgres query is doing a seq scan and is slow — which index should I add?" route_to: "postgresdb" why: "SQL/index/query tuning is engine-level work; Coolify only provisions and backs up the DB, it doesn't tune it." - prompt: "Set up the A record and nameservers for my domain at the registrar." route_to: "domains-dns" why: "DNS records / registrar / propagation is the domains-dns layer; Coolify consumes the resolved hostname, it doesn't manage DNS." capability: - scenario: "Stand up Coolify on a fresh Ubuntu 24.04 VPS, deploy a Next.js app from a GitHub repo with a managed Postgres, and configure nightly backups to S3-compatible storage with a custom domain + automatic SSL." must_include: - "Runs the install one-liner (curl -fsSL https://cdn.coollabs.io/coolify/install.sh | bash) AS ROOT, noting non-root is not fully supported" - "Calls out first-admin lockdown: register at http://<ip>:8000 immediately because the first account permanently owns the instance" - "Gives the port/firewall matrix: 8000 (and 6001/6002) restricted to the operator, 80/443 public" - "Picks a build pack (Nixpacks default) and connects the GitHub source with a branch" - "Sets app secrets as Coolify env vars / secrets injected at runtime, never baked into the image or committed" - "Provisions managed Postgres and connects the app over the INTERNAL Docker hostname, with the DB port NOT publicly exposed" - "Ensures the database has a named volume so data survives recreate, and the service has a healthcheck" - "Adds the custom domain and gets automatic Let's Encrypt SSL via the Traefik proxy on 80/443" - "Configures the S3 backup destination + cron schedule + retention AND an explicit restore-test/drill step" - "Acknowledges the boundary that VPS sizing/firewall/SSH hardening is hetzner/digitalocean, not this skill" -
README.md 1.5 KB
# Eval harness — `coolify` This harness checks two things: **triggering** (does the skill fire on own-the-box self-hosting prompts and stay quiet on near-misses that belong to a sibling?) and **capability** (does loading the skill measurably improve the deploy/backup answer?). Cases live in `cases.yaml`. These run through an **agent harness**, not a pure script — a human or driver agent feeds each prompt to a fresh agent session and grades the output. **Triggering.** For each `should_trigger` / `should_not_trigger` prompt, start a clean session with only `coolify` discoverable, paste the prompt verbatim, run 3–5 trials (the decision is stochastic), and score: `should_trigger` passes if the skill fires in the majority of trials; `should_not_trigger` passes if it stays quiet and would plausibly route to the named sibling (`hetzner`, `docker`, `railway`, `postgresdb`, `domains-dns`). Pass bar: ≥ 90% accuracy across all trigger cases. **Capability.** Run the scenario twice — WITHOUT the skill (clean agent) and WITH it — and grade each transcript against the `must_include` checklist, one point per bullet that is specifically and correctly covered (not just name-dropped). Pass bar: WITH the skill covers ≥ 80% of bullets and clearly beats WITHOUT (target ≥ +30 points, or fail→pass). If the baseline already nails it, the rubric is too easy. This is judgment-based grading: record trial counts and per-bullet verdicts so a reviewer can audit — never report a bare pass/fail.
-
-
references
-
databases-and-backups.md 4.5 KB
# Managed databases & backups — deep dive Source facts: coolify.io/docs/databases/backups and database docs, accessed 2026-06-02. Coolify provisions Postgres, MySQL, MariaDB, MongoDB, Redis (and more) in-product with generated credentials. This skill provisions and backs up; SQL/schema/tuning is the `postgresdb` skill; backup strategy/RPO/RTO doctrine across systems is the `backups` skill. Here we wire the concrete in-product job. ## Connection wiring Each managed database gets an **internal Docker network hostname** (the service name Coolify assigns). Connect apps to it over that private hostname — never a public IP. ```text # GOOD — app env var points at the internal hostname on the private network: DATABASE_URL=postgres://<user>:<password>@<internal-db-hostname>:5432/<db> # Inject <password> as a Coolify secret, not a committed literal. ``` Do **not** enable the database's "public port" toggle unless an external client genuinely needs it. An open 5432/3306/27017/6379 is scanned within minutes. If you must expose it, bind deliberately and firewall to known source IPs at the cloud layer. **Every database service needs a named volume.** Recreating a service with only an anonymous volume wipes the data — the most common Coolify data-loss incident. Coolify's managed DBs set this up; verify it before trusting the instance with real data. ## Scheduled backups Coolify runs the engine's native dump on a cron expression, stores it locally, and optionally pushes to S3-compatible storage. | Engine | Dump command Coolify runs | Restore = | | --- | --- | --- | | PostgreSQL | `pg_dump` | download → decompress → `psql`/`pg_restore` replay | | MySQL / MariaDB | `mysqldump` | download → decompress → `mysql <` replay | | MongoDB | `mongodump` | download → `mongorestore` | Configure per database: 1. **Enable scheduled backup** on the resource. 2. **Cron schedule** — e.g. `0 3 * * *` nightly 03:00; `0 */6 * * *` every 6h. Match to your tolerable data loss (RPO). 3. **Retention** — keep N copies/days so storage doesn't grow unbounded. 4. **Destination** — local only, or local + an S3-compatible target (below). ## S3-compatible destinations Add these as Coolify storage destinations. Provide endpoint + region + bucket + access key/secret — always as Coolify secrets, never in a committed file. | Provider | Endpoint shape | Notes | | --- | --- | --- | | AWS S3 | `s3.<region>.amazonaws.com` | Standard region + bucket | | Cloudflare R2 | `https://<accountid>.r2.cloudflarestorage.com` | Region usually `auto`; zero egress fees | | Backblaze B2 | `https://s3.<region>.backblazeb2.com` | S3-compatible endpoint, cheap storage | | MinIO (self-hosted) | `https://minio.example.com` | Your own object store; force path-style if needed | | Wasabi | `https://s3.<region>.wasabisys.com` | Flat-rate storage | Off-box storage is the entire point: a dump sitting on the same disk as the database dies with the disk. ## The restore runbook — run it before you need it A backup you have never restored is a guess. Do this once now, and on a recurring drill. ```bash # 1. Pull the latest dump from the destination (UI download, or via the S3 client). # e.g. from R2/B2/MinIO/S3 with the AWS CLI pointed at the endpoint: aws --endpoint-url "$S3_ENDPOINT" s3 cp "s3://$BUCKET/<path>/dump.sql.gz" ./dump.sql.gz # 2. Decompress. gunzip dump.sql.gz # 3. Replay into a THROWAWAY database (never the live one during a drill). # Postgres: psql "postgres://user:pass@host:5432/restore_test" < dump.sql # MySQL/MariaDB: mysql -h host -u user -p restore_test < dump.sql # MongoDB (from a mongodump archive/dir): mongorestore --uri "mongodb://user:pass@host:27017/restore_test" ./dump # 4. VERIFY: compare row counts / key tables against production expectations. # e.g. Postgres: psql "postgres://.../restore_test" -c "SELECT count(*) FROM orders;" # 5. Drop the throwaway DB. Record the drill date in 02-DOCS/wiki/stack/coolify.md. ``` ## Consistency caveats - These dumps are **logical, point-in-time snapshots at dump start**, not continuous PITR. Between dumps, writes are unprotected — size the cron to your RPO. - For true point-in-time recovery (WAL archiving / continuous replication) you need engine-level setup beyond Coolify's scheduled dump — that is `postgresdb` / `backups` territory. - For a large/active database, a logical dump can take time and add load; schedule it off-peak and watch that the dump window doesn't overlap the next one. - Test restores against the **same major engine version**; cross-version replays can fail or silently change behavior. -
deploy-recipes.md 4.7 KB
# Deploy recipes — build packs & a lint-clean compose Source facts: Coolify GitHub + v4.1 release notes, accessed 2026-06-02. ## Build-pack decision | Build pack | Reach for it when | Trade-off | | --- | --- | --- | | **Nixpacks** (default) | Standard app; want zero config | Auto-detects runtime; least control over the build | | **Railpack** (v4.1, beta) | Need build-time env vars, config merge, or multi-stage control Nixpacks can't express | Newer/beta; expect rougher edges than Nixpacks | | **Dockerfile** | You already have (or want to own) the build | You maintain the Dockerfile; pin the base image | | **docker-compose** | Multi-service app deployed as one unit | You own the compose — keep it lint-clean (below) | Start with Nixpacks. Escalate to Railpack only when a build-time env var or multi-stage step is required. Drop to a Dockerfile when you need full control of the build. Use compose when one deploy is several services. For image-authoring depth (multi-stage layering, cache mounts) that is the `docker` skill. ## Git source & env 1. Connect the Git provider (GitHub/GitLab app, or a deploy key) and pick the branch. 2. Enable auto-deploy on push, or trigger the deploy from CI (that's the `github-actions` skill). 3. Set every secret as a **Coolify environment variable / secret**, injected at runtime. Never bake a secret into the image and never commit it to the compose — a baked secret ships in every layer. 4. Add a **named volume** for any path that must survive a redeploy (uploads, sqlite, caches you care about). Anonymous volumes are wiped on recreate. 5. Add the domain (`app.example.com`), point its A record at the box (`domains-dns`), and Traefik issues the Let's Encrypt cert automatically on 80/443. ## A worked, lint-clean compose (the verify.sh target) This is the shape `scripts/verify.sh` asserts: secrets are env-refs (no literals), every DB has a named volume, services have a `healthcheck:`, and the build-context app does not float on `:latest`. Coolify injects the `${...}` values from its secrets store. ```yaml # docker-compose.yml — GOOD: env-ref secrets, named volume, healthchecks, pinned db image. services: web: build: context: . # build-context service: pinned via the Dockerfile's FROM, not :latest here restart: unless-stopped environment: # Secrets are REFERENCES, injected by Coolify — never literals committed to git. DATABASE_URL: ${DATABASE_URL} APP_SECRET: ${APP_SECRET} depends_on: db: condition: service_healthy healthcheck: test: ["CMD", "wget", "-qO-", "http://localhost:3000/health"] interval: 30s timeout: 5s retries: 3 labels: # Coolify/Traefik routing + auto-SSL is wired from the UI domain field; labels shown for reference. - "traefik.enable=true" db: image: postgres:16 # pinned major tag, not :latest restart: unless-stopped environment: POSTGRES_DB: ${POSTGRES_DB} POSTGRES_USER: ${POSTGRES_USER} POSTGRES_PASSWORD: ${POSTGRES_PASSWORD} # env-ref, injected by Coolify volumes: - pgdata:/var/lib/postgresql/data # NAMED volume — survives recreate healthcheck: test: ["CMD-SHELL", "pg_isready -U $${POSTGRES_USER}"] interval: 10s timeout: 5s retries: 5 # NOTE: no `ports:` mapping — db is reached over the internal network, not exposed publicly. volumes: pgdata: # declared named volume ``` ```yaml # BAD — every line here is a verify.sh failure or a production incident: services: web: image: myapp:latest # floating tag — irreproducible redeploys environment: DATABASE_URL: postgres://app:hunter2@db:5432/app # hardcoded secret literal in git # no healthcheck db: image: postgres:latest # floating tag environment: POSTGRES_PASSWORD: hunter2 # hardcoded secret literal ports: - "5432:5432" # DB port exposed to the public internet # no volume — data is wiped on every recreate ``` ## Why each rule earns its place - **Pinned image** — `:latest` silently changes the base; a redeploy you can't reproduce is a redeploy you can't roll back. - **Named volume on the DB** — recreate without it and the data is gone. This is the top Coolify footgun. - **Healthcheck** — Coolify/Traefik use it to know when the container is actually ready; without it, traffic hits a half-started app and `depends_on: service_healthy` can't gate startup. - **Env-ref secrets** — a literal in the compose leaks in git history and image layers. Inject from Coolify's secret store. - **No public DB port** — reach the database over the internal hostname; an exposed port is scanned in minutes. -
install-and-proxy.md 4.4 KB
# Install & proxy — deep dive Source facts: coolify.io/docs/get-started/installation and proxy docs, accessed 2026-06-02. Coolify 4.x. ## Full install transcript (Ubuntu 22.04/24.04 LTS) ```bash # Provision the box first (hetzner / digitalocean). Floor: 2 CPU / 2 GB / 30 GB. # Comfortable multi-app production: 4 vCPU / 8 GB / 100 GB NVMe. Coolify idles ~1 GB RAM. ssh root@<server-ip> # root required; non-root is not fully supported # Optional sanity: confirm OS + free disk before install cat /etc/os-release df -h / # Install (downloads + runs the official script; installs Docker + Coolify's own stack) curl -fsSL https://cdn.coollabs.io/coolify/install.sh | bash # Some docs show `| sudo bash` when piping; on a root shell `| bash` is correct. ``` The script takes ~2–5 minutes. When it finishes it prints the dashboard URL: `http://<server-ip>:8000`. ### Claim the admin account (do this immediately) Open `http://<server-ip>:8000` and register. The **first** account created becomes the permanent root admin of the instance — there is no later ownership-claim or admin-invite recovery flow. Until you register, anyone who can reach `:8000` can take the instance. Do not pause between the install finishing and this step. After registering, in the dashboard: 1. Settings → Instance: set the **FQDN** (e.g. `coolify.example.com`) and enable **force HTTPS** so the dashboard itself gets a Let's Encrypt cert and you stop using the raw `:8000` IP. 2. Point the dashboard FQDN's DNS A record at the box (DNS records are the `domains-dns` skill). ## Other OS / non-LTS notes - **Ubuntu non-LTS** (e.g. interim releases): supported but the installer may need manual Docker repo steps — install Docker Engine first, then re-run the Coolify script. - **Debian / Fedora / AlmaLinux / Rocky / CentOS / SUSE / Arch / Alpine**: supported. On distros where the script can't add the Docker repo automatically, install Docker Engine via the distro's official method, then run the Coolify script. - **Raspberry Pi OS**: 64-bit only. - General rule: if the one-liner fails partway, the usual cause is Docker not installing cleanly. Install Docker by hand, verify `docker run --rm hello-world`, then re-run the Coolify script (it is idempotent). ## Proxy: Traefik (default) vs Caddy Coolify ships **Traefik** as the default reverse proxy and can use **Caddy** instead. The proxy owns ports 80/443, routes hostnames to containers, and provisions Let's Encrypt certificates automatically. - **Traefik (default)**: keep it unless you have a specific reason. Most templates and docs assume it. - **Caddy**: switch in Settings → Proxy if you prefer Caddy's config model. Don't switch mid-flight on a busy instance without a maintenance window — the proxy restarts. You rarely hand-edit proxy config; Coolify generates labels/routes from each resource's domain settings. For a per-resource custom rule, use the resource's "Custom Traefik/Caddy labels" field rather than editing the proxy container directly (your edit would be overwritten on regeneration). ## Wildcard domain + DNS-01 challenge For `*.apps.example.com` (one cert for many subdomains) Let's Encrypt requires the **DNS-01** challenge, which needs API access to your DNS provider so the proxy can create the `_acme-challenge` TXT record. 1. Create a scoped DNS API token at your provider (Cloudflare, etc.). 2. Configure the DNS provider + token in the proxy/ACME settings. 3. Request the wildcard cert. The HTTP-01 challenge (the default for single hosts) cannot issue wildcards. Single-host certs use HTTP-01 and need only port 80 reachable + the A record pointing at the box. ## Common SSL failures and fixes | Symptom | Likely cause | Fix | | --- | --- | --- | | Cert never issues, "connection refused" in ACME logs | Port 80 not reachable from the internet | Open 80 at the cloud firewall; HTTP-01 needs it even for HTTPS-only sites | | "DNS problem: NXDOMAIN" | A record missing or not propagated | Confirm the A record points at the box; wait for propagation (domains-dns) | | Wildcard cert fails on HTTP-01 | Wildcards require DNS-01 | Configure the DNS provider token and use DNS-01 | | Rate-limit / "too many certificates" | Hit Let's Encrypt issuance limit by retrying | Wait out the limit; use staging while debugging, then switch to production | | Dashboard at `:8000` has no cert | FQDN + force HTTPS not set | Set the instance FQDN and enable force HTTPS in Settings |
-
-
scripts
-
verify.sh 5.2 KB
#!/usr/bin/env bash set -euo pipefail # verify.sh — coolify skill gate. Run from your PROJECT root (or this skill's dir). # # Static, read-only lint of any committed docker-compose artifact (docker-compose.yml/.yaml, # compose.yml/.yaml) plus a check that this skill's SKILL.md documents the canonical port matrix. # It NEVER writes, NEVER connects to the network, and NEVER deploys. # # Per compose file it asserts: # (a) no hardcoded secret literal — values for *PASSWORD*/*SECRET*/*TOKEN*/*_KEY*/DATABASE_URL # must be env-refs (${...}) , not inline literals. # (b) any database service (postgres/mysql/mariadb/mongo) is backed by a NAMED volume (data-loss guard). # (c) at least one healthcheck: present. # (d) no floating ":latest" image tag, and no image without an explicit tag (also "floating"). # And once: SKILL.md mentions every canonical port (8000 80 443 6001 6002). # # Exit code: non-zero ONLY on a hard failure (a/b/c/d or a missing port). An EMPTY/CLEAN target # (no compose files found) exits 0 with a skip — never a false failure. Missing optional context # (no SKILL.md to check) is advisory, not fatal. # # Portability: stock macOS bash 3.2 (no mapfile, no associative arrays). Arrays are initialised so # they expand safely under `set -u`. YELLOW=$'\033[33m'; GREEN=$'\033[32m'; RED=$'\033[31m'; NC=$'\033[0m' EXIT=0 skip() { printf '%s[skip]%s %s\n' "$YELLOW" "$NC" "$*"; } warn() { printf '%s[warn]%s %s\n' "$YELLOW" "$NC" "$*"; } ok() { printf '%s[ok]%s %s\n' "$GREEN" "$NC" "$*"; } err() { printf '%s[fail]%s %s\n' "$RED" "$NC" "$*"; EXIT=1; } ROOT="$(pwd)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" # --- discover compose files (skip vendor dirs) --- COMPOSE_FILES=() while IFS= read -r -d '' f; do COMPOSE_FILES+=("$f") done < <( find "$ROOT" \ \( -path '*/node_modules/*' -o -path '*/.git/*' -o -path '*/vendor/*' -o -path '*/.venv/*' \) -prune -o \ -type f \( -name 'docker-compose.yml' -o -name 'docker-compose.yaml' \ -o -name 'compose.yml' -o -name 'compose.yaml' \) -print0 2>/dev/null ) lint_compose() { file="$1" base="$(basename "$file")" # (a) hardcoded secrets: a *PASSWORD/SECRET/TOKEN/_KEY/DATABASE_URL key whose value is NOT a ${...} ref # and is not empty. We look at "key: value" pairs (env blocks are key: value in compose). if grep -Eiq '^[[:space:]]*[A-Za-z_]*(password|secret|token|_key|database_url)[A-Za-z_]*[[:space:]]*:[[:space:]]*[^$[:space:]"'"'"'].*' "$file"; then # exclude the case where the value is an env-ref like ${FOO} if grep -Ei '^[[:space:]]*[A-Za-z_]*(password|secret|token|_key|database_url)[A-Za-z_]*[[:space:]]*:[[:space:]]*' "$file" \ | grep -Eiv ':[[:space:]]*("?\$\{|\$[A-Za-z_])' \ | grep -Eq ':[[:space:]]*[^[:space:]$]'; then err "$base: hardcoded secret literal — use an env-ref (\${VAR}) injected by Coolify, not an inline value" fi fi # (b) named volume for any database service. If a db image is present, require a top-level `volumes:` map. if grep -Eiq 'image:[[:space:]]*("?)(postgres|mysql|mariadb|mongo)' "$file"; then if grep -Eq '^volumes:' "$file"; then ok "$base: database service has a named volume declaration" else err "$base: database service without a named volume — data is wiped on recreate" fi fi # (c) at least one healthcheck if grep -Eq '^[[:space:]]*healthcheck:' "$file"; then ok "$base: healthcheck present" else err "$base: no healthcheck: — Coolify/Traefik can't tell when the container is ready" fi # (d) floating tags: image: foo:latest, or image: foo with no tag at all while IFS= read -r line; do img="$(printf '%s' "$line" | sed -E 's/^[[:space:]]*image:[[:space:]]*"?//; s/"?[[:space:]]*$//')" case "$img" in \$*) : ;; # env-ref image, skip *:latest) err "$base: floating ':latest' tag on image '$img' — pin a tag or digest" ;; *@sha256:*) : ;; # digest-pinned, fine *:*) : ;; # has an explicit tag, fine *) err "$base: image '$img' has no tag (floating) — pin a tag or digest" ;; esac done < <(grep -E '^[[:space:]]*image:[[:space:]]*' "$file" || true) } if [ "${#COMPOSE_FILES[@]}" -eq 0 ]; then skip "no docker-compose files found — nothing to lint (clean)" else for f in "${COMPOSE_FILES[@]}"; do lint_compose "$f" done fi # --- port matrix documented in SKILL.md (look next to this script, then under ROOT) --- SKILL_MD="" if [ -f "$SCRIPT_DIR/../SKILL.md" ]; then SKILL_MD="$SCRIPT_DIR/../SKILL.md" else cand="$(find "$ROOT" -type f -name 'SKILL.md' -path '*coolify*' -print 2>/dev/null | head -n1 || true)" [ -n "$cand" ] && SKILL_MD="$cand" fi if [ -n "$SKILL_MD" ] && [ -f "$SKILL_MD" ]; then missing="" for p in 8000 80 443 6001 6002; do grep -Eq "(^|[^0-9])$p([^0-9]|$)" "$SKILL_MD" || missing="$missing $p" done if [ -n "$missing" ]; then err "SKILL.md missing canonical port(s):$missing" else ok "SKILL.md documents the canonical port matrix (8000/80/443/6001/6002)" fi else skip "no coolify SKILL.md found to check the port matrix" fi printf '\n' if [ "$EXIT" -eq 0 ]; then ok "verify.sh passed"; else err "verify.sh found failures"; fi exit "$EXIT"
-
-
SKILL.md 13.1 KB
--- name: coolify description: "Use when self-hosting apps and databases with Coolify on a VPS you own — install, first-admin lockdown, Git-to-deploy (Nixpacks/Dockerfile/compose), managed Postgres/Redis, scheduled S3 backups, domains + auto-SSL. NOT a PaaS someone else runs (that is `railway`), NOT sizing/hardening the box (that is `hetzner`), NOT authoring Dockerfiles (that is `docker`)." tags: [coolify, self-hosting, paas, vps, docker, deployment, devops] recommends: [hetzner, digitalocean, docker, postgresdb, backups, domains-dns, github-actions] origin: risco --- # Coolify — own-the-box PaaS on your VPS Coolify is an open-source, self-hostable control plane that turns a plain Linux VPS into a Vercel/Heroku replacement: Git-to-deploy, managed databases, automatic SSL, scheduled backups — for a flat VPS bill instead of per-request metering. This skill is operational. It produces the exact install one-liner, the port matrix, env wiring, a lint-clean compose artifact, and a backup-to-S3 cron with a tested restore. Coolify 4.x (stable in 2026); v4.1 added Railpack, structured audit logging, and a read-only MCP server. ## Mental model — read this first 1. **Coolify is the control plane, not the box.** It runs *on* a VPS you provisioned elsewhere. Hardware sizing, the cloud firewall, SSH hardening, OS patching → that is the VPS layer (hetzner / digitalocean), not this skill. Coolify orchestrates what runs on the box. 2. **You own the data and the uptime now.** No vendor takes nightly snapshots for you. So backups are non-optional and a backup you have never restored is not a backup (see the backup section). 3. **Everything underneath is Docker.** Coolify generates compose/Dockerfile runs and a Traefik proxy. So every Docker rule still applies — named volumes, healthchecks, pinned images, env-injected secrets. For image-authoring depth (multi-stage, layer caching) go to the `docker` skill; here we wire it. 4. **First account to register owns the instance — forever.** There is no "admin invite" recovery if a stranger registers first. Claim it within seconds of install (see step 4). ## Decision — where should this app actually run? | Option | Who owns the box | Cost shape | Ops burden | Choose when | | --- | --- | --- | --- | --- | | **Coolify self-hosted** (this skill) | You (your VPS) | Flat VPS/month | You patch + back up | Predictable bill, data sovereignty, many apps on one box | | **Coolify Cloud** | You bring the server; they host the control plane | VPS + small SaaS fee | They run the control plane | Want Coolify UX without babysitting the dashboard's own uptime | | **Managed PaaS** (railway / render / fly-io / vercel / netlify) | The provider | Per-usage, scales up fast | Near-zero | Spiky scale-to-zero, no box to own — the inverse of Coolify | Choose Coolify when a few always-on apps on a known-cost box beats per-request metering. If the user wants `git push` and never touches a server, that is railway — route there. ## Install — the 6-minute path Provision the box first (this is the hetzner / digitalocean step, not this skill). Floor is 2 CPU / 2 GB RAM / 30 GB disk; Coolify itself idles around 1 GB RAM. For comfortable multi-app production aim for 4 vCPU / 8 GB / 100 GB NVMe. Ubuntu 22.04/24.04 LTS recommended (Debian, Fedora/Alma/Rocky, Alpine, Arch, Raspberry Pi OS 64-bit also supported; non-LTS needs manual steps — see references). ```bash # 1. SSH in AS ROOT. Non-root is not fully supported by the installer — log in as root. ssh root@<server-ip> # 2. One-liner install (root). Takes ~2–5 min: installs Docker, sets up Coolify's own stack. curl -fsSL https://cdn.coollabs.io/coolify/install.sh | bash ``` ```text # 3. The installer prints the dashboard URL: http://<server-ip>:8000 # 4. CRITICAL — within seconds, open http://<server-ip>:8000 and REGISTER. # That first account becomes the permanent root admin of the instance. First-come-owns-the-box; # there is no later "claim ownership" flow. Do not walk away between step 2 and this. # 5. In the dashboard: set the instance FQDN (coolify.example.com) and force HTTPS, so you stop # hitting the raw :8000 IP and the dashboard itself gets a real certificate. ``` Why root: the installer wires Docker and system services; the docs state non-root is not fully supported. Why claim immediately: the registration page is open until someone takes it. Full install transcript, non-LTS/other-OS steps, dashboard-domain setup, Traefik (default) vs Caddy, wildcard domains + the DNS-01 challenge, and common SSL failures with their fixes → `references/install-and-proxy.md`. ## Port & firewall matrix Set these at the cloud firewall (hetzner/digitalocean layer) AND keep them in mind on the box. | Port | Purpose | Exposure | | --- | --- | --- | | 22 | SSH | Restrict to your IP / VPN; key-only | | 8000 | Coolify dashboard | **Restrict** to your IP after setup, or front with the proxy on a FQDN; never leave open to the world | | 80 | Proxy HTTP (Traefik) → redirects to 443 | Public | | 443 | Proxy HTTPS, app traffic + Let's Encrypt | Public | | 6001 | Realtime / websocket | Open as the dashboard needs (same audience as 8000) | | 6002 | Realtime / terminal websocket | Same as 6001 | Rule: 80/443 are the only ports the public should reach. 8000/6001/6002 are operator surface — lock them to your IP or put the dashboard behind its FQDN. Why: an open :8000 plus an unclaimed instance is a takeover; an open :8000 on a claimed instance is still your full control plane exposed to brute force. ## Deploy an app Pick the build pack first, then wire source → env → storage → domain. | Build pack | Use when | Note | | --- | --- | --- | | **Nixpacks** (default) | Standard app, you want zero config | Auto-detects the language/runtime; start here | | **Railpack** (v4.1, beta) | Need build-time env vars, config merge, or multi-stage control | Newer; reach for it when Nixpacks can't express the build | | **Dockerfile** | You already have a Dockerfile / want full build control | You own the build; pin a base image | | **docker-compose** | Multi-service app deployed as a unit | Coolify runs your compose; this is the verify.sh-linted artifact shape | Then: 1. **Connect the Git source** — GitHub/GitLab app or a deploy key; pick branch; enable auto-deploy on push (or trigger from CI — that wiring is the `github-actions` skill). 2. **Set env vars as Coolify secrets** — inject at runtime via the UI/API; never bake secrets into the image or commit them to the compose. Why: a baked secret ships in every image layer and leaks on pull. 3. **Add persistent storage** — any path that must survive a redeploy (uploads, sqlite) gets a named volume. A container without one loses its writes on every recreate. 4. **Attach the domain + SSL** — add `app.example.com`, point its DNS A record at the box (DNS itself is the `domains-dns` skill), and the Traefik proxy provisions Let's Encrypt automatically on ports 80/443. Build-pack decision deep-dive plus a worked, lint-clean `docker-compose.yml` (env-ref secrets, named volumes, healthcheck, pinned image) — the verify.sh target → `references/deploy-recipes.md`. ## Managed databases Coolify provisions Postgres, MySQL, MariaDB, MongoDB, Redis (and more) in-product — one-click, with generated credentials. - **Connect apps over the internal Docker network hostname**, not a public IP. Coolify gives each DB an internal service name; the app reaches it on the private network. Why: zero public attack surface. - **Do NOT publicly expose the DB port** unless you genuinely need an external client. If you must, bind it deliberately and firewall it to known IPs — an open 5432/3306 is scanned within minutes. - **Every DB gets a named volume.** Without it, recreating the service wipes the data. This is the single most common Coolify data-loss footgun. This skill provisions and connects the DB; it does not teach SQL, indexing, or query tuning — that is the `postgresdb` skill. Per-engine details and connection strings → `references/databases-and-backups.md`. ## Backups to S3 — non-negotiable You own the data now. Coolify runs scheduled dumps (`pg_dump` / `mysqldump` / `mongodump`) on a cron expression, stores them locally, and optionally pushes to S3-compatible storage: AWS S3, Cloudflare R2, Backblaze B2, MinIO, Wasabi. 1. **Add an S3 destination** — endpoint, region, bucket, access key/secret (as Coolify secrets, never in a committed file). Off-box storage is the point: a backup on the same disk dies with the disk. 2. **Set the cron schedule** — e.g. `0 3 * * *` for nightly 03:00. Match frequency to how much data you can afford to lose (RPO doctrine across systems is the `backups` skill; here we wire the in-product job). 3. **Set retention** — keep N days/copies so the bucket doesn't grow unbounded. 4. **Run the restore drill — once, before you need it.** Download a dump, decompress, replay it into a throwaway DB, confirm the row counts. An untested backup is a guess. Schedule a recurring drill. S3 destination setup for R2/B2/MinIO/Wasabi/AWS, per-engine dump/restore commands, retention, the step-by-step restore runbook and its consistency caveats → `references/databases-and-backups.md`. ## Operate the box - **Resource planning.** Coolify idles ~1 GB RAM; budget the rest for your apps + databases. Multi-app sweet spot is 4 vCPU / 8 GB / 100 GB NVMe. Watch RAM headroom before adding the next app. - **Updates.** Update Coolify from the dashboard's settings; pin/snapshot before a major bump. - **Logs & audit.** Per-resource logs live in the UI. v4.1 adds a structured audit log — use it to see who changed what. - **MCP server (v4.1, read-only).** Coolify exposes an instance-level MCP server with read-only tools for AI-agent integration. Useful for letting an agent inspect status; it is read-only by design — do not treat it as a deploy channel, and still lock down the network in front of it. ## Anti-patterns → STOP | Rationalization | Reality → STOP | | --- | --- | | "I'll install with my sudo user, root feels risky" | The installer wires Docker + system services and the docs say non-root is not fully supported. SSH in as root for install. | | "I'll register the admin account later" | First account to hit `:8000` owns the instance permanently. A stranger registering first = takeover. Claim it within seconds. | | "Leave :8000 open, it's password-protected" | That is your full control plane exposed to brute force. Restrict 8000/6001/6002 to your IP; only 80/443 are public. | | "Pin the app image to `:latest`, it's simpler" | `:latest` floats — a silent base change breaks a redeploy you can't reproduce. Pin a tag or digest. | | "The database doesn't need a named volume yet" | Recreating the service wipes an anonymous volume. Data gone. Named volume from day one. | | "Put the DB password in the compose so deploys are reproducible" | Secrets in a committed file leak in git history and image layers. Inject via Coolify env/secrets, env-ref only. | | "Expose 5432 so I can connect from my laptop" | An open DB port is scanned in minutes. Use the internal hostname; if you truly need external access, firewall it to known IPs. | | "Backups are configured, we're covered" | A backup you've never restored is a guess. Run the restore drill before you need it. | | "Coolify will harden the server for me" | Coolify is the control plane, not the OS-hardening layer. Firewall/SSH/patching is the VPS skill (hetzner/digitalocean). | | "Use Coolify because I just want to git push and forget the server" | That's the opposite of own-the-box. Use a managed PaaS — railway. | ## Verify Run `scripts/verify.sh` against your project (or this skill's `references/`). It statically lints the example/your `docker-compose.y*ml`: fails on a hardcoded secret literal (must be env-ref), a DB service without a named volume, a missing `healthcheck:`, or a floating `:latest` tag on a build-context service; and confirms the canonical port matrix (8000/80/443/6001/6002) is documented. Read-only, no network, no live deploy. Exits 0 on a clean/empty target. ## Project grounding (02-DOCS + CLAUDE.md) When this skill runs in a project with a `02-DOCS/` layer (the harness Karpathy wiki), record this instance's deploy topology there and index it in `02-DOCS/wiki/index.md`, so the next agent inherits it instead of re-deriving it. 1. **Find the article** `02-DOCS/wiki/stack/coolify.md`, indexed in `02-DOCS/wiki/index.md` (the Knowledge map index; root `CLAUDE.md` points to it). 2. **If missing or stale**, create/update it with the real choices — the box (provider/specs), instance FQDN, which apps/databases run on it, build packs in use, the backup destination + cron + retention, and the port/firewall decisions — then index it in `02-DOCS/wiki/index.md` (the Knowledge map; root `CLAUDE.md` keeps only a short pointer to it). 3. **Read it first on every use** and stay consistent; when the topology changes, update the article (bump its `Updated` date) in the same change. Never commit credentials here — record *where* secrets live, not their values. No `02-DOCS/` layer? Skip silently. Topology is recorded, not gated — never block the task on this.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.