automation-strategy
Use when deciding whether a process is worth automating, sizing ROI and build-vs-buy, choosing an automation platform, or diagnosing why a fleet of automations keeps breaking — the decision layer before anyone builds. NOT building the flow (that is `automation-flows`, or `n8n` /
#automation
Install
npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/automation-strategy
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
git clone https://github.com/ericrisco/rsc-harness.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Automation strategy — decide whether, where, and how before anyone builds
This is the discipline layer of the automation suite. It answers three questions that come before the first node is wired: should this be automated at all?, on what platform, and is a platform even the right tool vs code?, and what orchestration guarantees does every automation have to meet so the fleet doesn't rot? No platform API, no importable JSON, no code lives here — the moment the decision is made, you route to the skill that builds it.
Building the flow (design + importable artifact) → ../automation-flows/SKILL.md. Driving a specific platform's live API/MCP → ../n8n/SKILL.md, ../make/SKILL.md, ../zapier/SKILL.md, ../power-automate/SKILL.md. A typed API client in code → ../api-connector-builder/SKILL.md. Receiving webhooks in your own app → ../webhooks/SKILL.md.
Every fact about pricing and platform capability below moves fast. Treat the ≈ figures as hedges and re-check the vendor page at author time; the honest limits (Zapier can't create Zaps via public API; Power Automate can't create end-user "My flows" cleanly via API) are called out where they matter and expanded in the references.
1. Should you automate at all?
Most "we should automate this" instincts are wrong — not because automation is bad, but because the process underneath isn't ready, or the payoff isn't there. Gate on four factors, all of them, before writing a line:
- Frequency — how many times per week/month does this run? A once-a-quarter task rarely earns its maintenance cost.
- Time saved per run — minutes of human toil removed each time. Two minutes × 1,000 runs beats an hour × 3.
- Error rate of doing it by hand — high manual error rate (typos, missed steps, missed SLAs) is often a bigger win than the time saved. Automation's real product is consistency, not just speed.
- Stability of the process — how often do the steps, the systems, or the rules change? An unstable process is the one thing that turns automation into a liability.
Rough gate: automate when it is frequent AND stable AND either saves meaningful time-per-run or removes a costly manual error rate. If it's frequent but unstable, or high-value but requires judgment on every run, do not automate yet — see below. Full scoring rubric with worked numbers: references/automate-decision-and-roi.md.
The trap: automating a broken or unstable process
Automating a bad process just makes you produce bad output faster, and now nobody can see the badness because a machine hides it. Fix and standardize the process first, then automate the standardized version. If the steps aren't written down the same way twice, if it "depends", if three people do it three ways — you are not ready. Standardize on paper, run it manually a few times against the written steps, then encode it. Automating chaos gives you automated chaos with a maintenance bill.
Human-in-the-loop: what you deliberately do NOT fully automate
Some steps should stay manual by design, wrapped in an approval gate rather than removed:
- Judgment calls — pricing exceptions, hiring decisions, content that carries brand/legal risk, anything where "usually right" isn't good enough.
- Irreversible or high-blast-radius actions — sending money, deleting data, emailing your whole list, publishing publicly, terminating accounts. Automate the preparation, require a human click for the commit.
- Low-frequency, high-stakes — the payoff is small and the cost of a silent wrong run is huge.
The pattern is assisted, not autonomous: the automation gathers, drafts, and stages; a human approves the commit. Cheap insurance against the day the automation is confidently wrong at scale.
2. ROI and build-vs-buy
Payback math
build cost = hours to build × loaded hourly rate
monthly saving = (runs/month × time saved per run × loaded rate)
− monthly platform/run cost
− monthly maintenance cost ← the line people forget
payback (months)= build cost / monthly saving
Two rules of thumb: if payback > ~12 months, the process is probably too rare or too cheap to bother; if maintenance eats > ~30–50% of the gross saving, you've built a pet, not a tool. Maintenance is not optional — APIs change, auth expires, edge cases surface. Budget it explicitly or the ROI is fiction. Worked examples: references/automate-decision-and-roi.md.
No/low-code platform vs custom code
| Choose a no/low-code platform when… | Choose custom code when… |
|---|---|
| Gluing 2–8 SaaS apps with pre-built connectors | Logic is genuinely complex, algorithmic, or stateful |
| Business owner needs to read/edit the flow | High volume where per-run platform billing dominates TCO |
| Speed to first version matters more than control | You need version control, tests, code review, CI on the logic |
| Volume is modest and connectors exist | The vendor has no connector and you need deep API control |
The honest middle: platforms win on time-to-first-version and readability; code wins on complex logic, testability, and cost at scale. Many mature setups are hybrid — a platform orchestrates and a small custom service (or api-connector-builder client) does the hard part.
Total cost of ownership
TCO is not the sticker price. Count: the platform bill under your real volume and its billing unit (this is where surprises live — see §3), build time, ongoing maintenance, the cost of an outage when it fails silently, and the switching cost if you outgrow the platform. A "free" automation that a senior engineer babysits monthly is not free.
3. Platform selection
Pick on billing unit first — it's the axis that dictates cost at scale, and choosing on familiarity instead is the classic 10× mistake. Then filter on data residency, ecosystem, programmatic maturity, and portability. This is a routing decision: once chosen, hand off to that platform's skill.
| Axis | n8n | Make | Zapier | Power Automate |
|---|---|---|---|---|
| Billing unit | per execution (whole run = 1, any # of steps); self-host = $0 | per credit (each module action = 1; was "operations", renamed 2025-08-27) | per task (every action counts) | per user or per flow (premium); bundled with many M365 plans |
| Best cost shape | few long/complex flows, high volume | moderate multi-step flows | many short simple flows | teams already paying for M365 |
| Self-host / data residency | yes — your infra, your data | no (EU/US regions only) | no | Microsoft cloud / Dataverse regions |
| Ecosystem fit | technical teams, raw HTTP, code nodes | mid-market, deep per-app modules | widest app catalog (~9,000+) | native to Microsoft 365 — Teams, SharePoint, Outlook, Dataverse |
| Programmatic + MCP maturity | public REST API to CRUD workflows; community MCP; strongest for code-driven ops | management API + MCP support maturing | rich API for actions + Zapier MCP to expose actions to agents — but no public API to create Zaps | Flow management APIs + solutions/ARM — but no clean public API to create end-user "My flows" |
| Portability | high — export/import JSON + self-hostable runtime | medium — blueprint export/import, but runs only on Make | low — no portable export | low — tied to MS solutions |
Routing shortcuts: already living in Microsoft 365 → Power Automate (../power-automate/SKILL.md). Data must stay on your own infra, or high volume where execution-billing wins → n8n (../n8n/SKILL.md). Obscure app, non-technical owner → Zapier (../zapier/SKILL.md). Deep per-app modules, mid volume → Make (../make/SKILL.md).
Two capability limits to state honestly, not paper over: you can drive Zapier's actions programmatically and via MCP, but you cannot create a Zap through a public API — Zap authoring is UI-only. Power Automate exposes flow management and solution deployment, but there is no clean public API to author an end-user's personal "My flows". If a plan depends on API-creating either, the plan is wrong. Full matrix, MCP notes, and caveats: references/platform-selection-matrix.md.
4. Orchestration strategy (engine-agnostic)
These are the guarantees every automation must meet regardless of platform. This is the policy; the per-platform recipe lives in ../automation-flows/SKILL.md and the platform skills. Deep dive: references/orchestration-and-antipatterns.md.
- Idempotency & dedup. Delivery is at-least-once and retries replay; assume every event can arrive twice. Require a stable dedup key (event id / external id) checked before any non-idempotent write (charge, email, insert). No dedup on a non-idempotent write = duplicate charges waiting to happen.
- Error paths, retries & backoff. Every automation has a defined destination for failure. Retry transient errors (timeouts, 429/5xx) with backoff and a cap; do not retry permanent errors (400/validation) — route them out. A flow with no error path isn't finished.
- Dead-letter. After retries are exhausted, the failed item lands in a durable place a human can inspect and replay — a queue, a table, an "incomplete executions" store. Dropped-on-the-floor failures are invisible data loss.
- Observability & alerting. Emit success/failure counts; put a heartbeat on scheduled jobs and alert on silence (a cron that stops firing is the failure you never hear about). Route alerts to a named owner, not a dead inbox.
- Secrets handling. Credentials in the platform's credential store / a vault, referenced by name — never pasted inline where an export leaks them. Least privilege, and a rotation plan. Engineering detail →
../secure-coding/SKILL.md. - Versioning & environments. Separate dev/staging from prod; export before you edit; keep a rollback. Changes to a live automation are production changes and deserve the same care.
- Avoid fan-out storms. Cap concurrency, batch, and rate-limit. Watch for echo loops (flow A writes a record that triggers flow B that triggers flow A) and for a single event fanning out into thousands of downstream calls that trip rate limits or blow the bill.
5. Anti-patterns
| Anti-pattern | Why it bites | Do instead |
|---|---|---|
| No error path | First API hiccup, the run dies; you learn from an angry customer | Define a failure destination + alert before shipping |
| Silent failure | Failures that don't alert erode trust worse than no automation — everyone assumes it worked | Alert on failure and on silence (heartbeat) |
| No kill-switch | When it misfires at scale, there's no fast way to stop the bleeding | Every outward-writing automation gets a documented pause/disable toggle |
| Over-automation | Automating rare / low-value / judgment work costs more to maintain than doing it by hand | Gate on §1; leave judgment steps human-in-the-loop |
| Automating a broken process | You produce bad output faster and hide the badness behind a machine | Standardize the process first, then encode it |
| Hidden single point of failure | One personal token / one undocumented flow everyone depends on | Service accounts, documented dependencies, no personal-account glue |
| No owner | Orphaned automations rot; nobody notices the break or knows how it works | Name an owner + a one-page runbook per automation |
| Choosing platform by familiarity | Bill 10× higher than the right billing unit; "our automation bill exploded" | Pick by billing unit for your run shape (§3) |
Checklist
- Ran the §1 gate: frequency × time-saved × error-rate × stability — and confirmed the process is standardized, not chaos being paved.
- Identified any judgment / irreversible / high-blast-radius step and kept it human-in-the-loop with an approval gate.
- Computed payback including maintenance, and confirmed it's under the rule-of-thumb ceiling.
- Decided platform vs custom code on the build-vs-buy criteria, not on what's familiar.
- Chose the platform by billing unit for the real run shape, then data residency / ecosystem / portability — and did NOT assume a capability the platform lacks (Zap creation, PA "My flows").
- Specified the orchestration guarantees §4 requires (idempotency key, error path, dead-letter, heartbeat/alert, secrets store, dev/prod split, fan-out caps) as requirements handed to the build skill.
- Every automation has a named owner, a runbook, and a kill-switch.
- Routed the actual build to
automation-flows/ the chosen platform skill — no flow was built inside this skill.
Files (rsc-harness)
-
evals
-
cases.yaml 4.4 KB
skill: automation-strategy should_trigger: - prompt: "Should we automate our weekly invoice-reconciliation process? It changes a bit every month." why: The core whether-to-automate decision — and the "changes every month" signals the stability trap this skill exists to catch. - prompt: "n8n vs Make vs Zapier vs Power Automate — which one should we standardize on as a company?" why: Platform selection across all four engines by billing model and ecosystem fit; a strategy/routing decision, not a build. - prompt: "What's the ROI on automating this task, and should we build a custom script or use a no-code platform?" why: Payback math plus build-vs-buy — squarely §2, before anyone builds. - prompt: "Our automations keep breaking and nobody knows why. Help us design a proper automation strategy." why: Fleet-health / orchestration-discipline diagnosis — missing error paths, ownership, observability across a fleet. - prompt: "We're a Microsoft 365 shop. Does it even make sense to pay for Zapier, or use what we have?" why: Ecosystem-fit selection (MS 365 → Power Automate) and TCO reasoning; a platform-choice decision. - prompt: "¿Merece la pena automatizar esto o me sale más caro de mantener que hacerlo a mano?" why: Spanish; ROI / maintenance-cost gate — the "pet not a tool" payback question. - prompt: "Automating this last quarter cost us more than it saved. What did we get wrong?" why: TCO / over-automation post-mortem — the maintenance-eats-the-ROI and over-automation anti-patterns. should_not_trigger: - prompt: "Give me the importable n8n JSON for a webhook → Notion → Slack flow with error handling." route_to: automation-flows why: A concrete build/artifact request; automation-flows owns design and the importable workflow JSON. - prompt: "Use the n8n API to create and activate this workflow on my instance." route_to: n8n why: Driving a specific platform's live API to CRUD/activate — the n8n platform skill, not the strategy layer. - prompt: "Write a TypeScript client for the Acme REST API with OAuth, pagination, and exponential backoff." route_to: api-connector-builder why: Generic typed API client in code; that is api-connector-builder, not automation strategy. - prompt: "Build the Express endpoint that receives and verifies Stripe webhook signatures." route_to: webhooks why: Building the inbound receiver in your own app; webhooks owns receipt, this skill treats webhooks only as triggers. - prompt: "Set up the Make scenario: on new HubSpot deal, create a Trello card and email the owner." route_to: automation-flows why: A specific flow to build on a named platform — design goes to automation-flows, operation to the make skill; not a strategy question. - prompt: "How do I rotate and store the API keys my flow uses securely in code?" route_to: secure-coding why: Secrets-engineering implementation detail; §4 sets the policy but secure-coding owns the how. capability: - scenario: "A 30-person Microsoft 365 company asks: 'We do a manual monthly vendor-payment run — pull approved invoices, pay them, notify finance. Should we automate it, on what platform, and how do we not screw it up?' Produce the strategy recommendation." must_include: - "Runs the §1 gate on frequency × time-saved × error-rate × stability before recommending automation." - "Flags the payment/commit step as irreversible/high-blast-radius and keeps it human-in-the-loop behind an approval gate (automate prep, human clicks to pay)." - "Payback math that explicitly includes ongoing maintenance cost, not just time saved." - "Recommends a platform by billing unit and ecosystem fit — notes Microsoft 365 fit points to Power Automate — and routes the build to that platform skill / automation-flows rather than building here." - "States the honest limit that Power Automate has no clean public API to author end-user 'My flows' (and Zapier can't create Zaps via API), so a plan can't depend on API-creating them." - "Specifies orchestration guarantees as requirements: idempotency/dedup key, error path + retry/backoff, dead-letter, heartbeat/alert to a named owner, secrets in a vault, dev/prod split, and a kill-switch." - "Does NOT emit importable workflow JSON or platform API code — defers the build to automation-flows / the platform skill." -
README.md 1.2 KB
# Evals — automation-strategy `cases.yaml` holds three groups. `should_trigger` covers the decisions this skill owns: whether to automate (with the stability trap), four-way platform selection, ROI / build-vs-buy, fleet-health "our automations keep breaking", ecosystem-fit selection, and TCO/over-automation post-mortems, with Spanish phrasing. `should_not_trigger` lists adjacent prompts routed to the sibling that owns them — concrete builds/artifacts → automation-flows, driving a live platform → the n8n/make skills, code clients → api-connector-builder, webhook receipt → webhooks, secrets engineering → secure-coding. `capability` is one end-to-end strategy scenario whose `must_include` rubric checks the §1 gate, human-in-the-loop on the irreversible step, maintenance-inclusive payback, billing-unit + ecosystem platform routing, the honest Zapier/Power-Automate API-creation limits, the orchestration guarantees, and that it defers the build (no JSON/API code emitted here). No automated runner. Score by judgement: feed each prompt to the routing layer and confirm it activates or routes to the listed sibling; for `capability`, have the skill produce the strategy recommendation and check every `must_include` item by hand.
-
-
references
-
automate-decision-and-roi.md 6.1 KB
# Automate decision & ROI — the numbers behind §1–§2 Depth for "should we automate this, and does it pay back?" Engine-agnostic. Once the answer is yes, route the build to `../../automation-flows/SKILL.md` or the chosen platform skill. ## The decision scorecard Score each factor honestly. None of these is a magic threshold — they're a forcing function to stop you automating on vibes. | Factor | Ask | Weak signal (lean NO) | Strong signal (lean YES) | | --- | --- | --- | --- | | **Frequency** | runs per week/month | a handful per quarter | dozens+ per week | | **Time saved / run** | minutes of human toil removed | seconds, or "it's quick anyway" | many minutes, or a context-switch tax | | **Manual error rate** | how often a human gets it wrong | near-zero, low stakes | frequent, or costly when wrong | | **Stability** | how often steps/systems/rules change | changes most months | unchanged for many months | | **Toil / morale** | is it soul-crushing repetitive work | mildly annoying | actively burning out a person | **The gate:** automate when **frequent AND stable AND (meaningful time-per-run OR costly manual error rate)**. Instability vetoes everything else — a high-value process that changes monthly will cost more in rework than it saves. When in doubt, the tie-breaker is stability, then frequency. ### Why stability is the veto Automation encodes *today's* process. Every change to the underlying process is now a change to code/config someone has to make, test, and redeploy. On a stable process that's rare. On a churning process it's a second job. If the process is still being figured out, keep it manual and cheap until it stops moving. ## The unstable / broken-process trap Automating a broken process gives you **automated brokenness at higher throughput**, plus a machine that now obscures the brokenness from the humans who used to feel it. Sequence to do it right: 1. **Write the process down** — the exact steps, inputs, decision rules, outputs. 2. **Standardize** — if three people do it three ways, pick one way. Remove the "it depends" branches or make them explicit rules. 3. **Run it manually against the written steps** a few times. Fix the doc where reality diverges. 4. **Only now automate** the standardized version. Smell tests that you're not ready: the steps can't be written without "usually" or "it depends"; different people produce different outputs from the same input; the systems involved are mid-migration. ## Human-in-the-loop taxonomy Keep these steps assisted, not autonomous. The automation stages the work; a human commits it. | Step type | Example | Pattern | | --- | --- | --- | | **Judgment** | pricing exception, hiring, brand/legal-sensitive content | draft + route for approval | | **Irreversible** | send money, delete data, terminate an account | prepare automatically, require a human click to commit | | **High blast radius** | email the whole list, publish publicly, mass update | stage + preview + explicit go | | **Low-freq / high-stakes** | quarterly filing, contract renewal | checklist-assist, human executes | Rule of thumb: if a wrong autonomous run would be expensive *and* hard to undo, insert an approval gate. It costs one click and buys you the day the automation is confidently wrong. ## Payback math (worked) ```text build cost = build hours × loaded hourly rate gross saving/mo = runs/month × time saved per run (hrs) × loaded hourly rate net saving/mo = gross saving/mo − platform cost/mo − maintenance cost/mo payback (months)= build cost / net saving/mo ``` **Loaded rate** = salary + overhead (benefits, tooling, management), not the bare wage — usually 1.3–2× the wage. Use it for both build cost and time saved so the comparison is fair. ### Example A — clear win - Task runs 800×/month, saves 4 min each → 53.3 hrs/month saved. - Loaded rate $60/hr → gross saving ≈ $3,200/month. - Build = 20 hrs × $60 = $1,200. Platform ≈ $30/month. Maintenance ≈ 1 hr/month = $60. - Net saving ≈ $3,200 − $90 = $3,110/month. **Payback ≈ 0.4 months.** Build it. ### Example B — a pet, not a tool - Task runs 6×/month, saves 15 min each → 1.5 hrs/month. - Gross saving ≈ $90/month. Build = 25 hrs × $60 = $1,500. Maintenance ≈ 1.5 hrs/month = $90. - Net saving ≈ $90 − $90 = **~$0/month.** Payback = never. Don't automate; write a checklist. ### The two rules of thumb - **Payback > ~12 months** → the process is probably too rare or too cheap. Reconsider. - **Maintenance > ~30–50% of gross saving** → you've built something that eats its own ROI. The fragile-integration tax (auth expiry, API changes, edge cases) is real; if it's this high, either simplify the flow or question whether it should exist. Maintenance is the line people omit and the reason ROI projections lie. Budget it explicitly. ## Build vs buy (platform vs custom code) | Dimension | No/low-code platform | Custom code | | --- | --- | --- | | Time to first version | hours | days+ | | Readable by non-engineers | yes | no | | Complex/stateful/algorithmic logic | awkward, hits ceilings | natural | | Testability, review, CI | weak | strong | | Cost at high volume | per-run billing can dominate | marginal cost near zero | | Vendor/connector coverage | huge | you build each integration | | Portability / lock-in | often locked in | yours | **Hybrid is normal and often best:** a platform orchestrates the glue and human-readable flow; a small custom service or an `api-connector-builder` client does the one hard, high-volume, or algorithmic part. Don't force either extreme. ## Total cost of ownership checklist - Platform bill **at your real volume, in the platform's billing unit** (tasks vs credits vs executions vs seats — see the platform matrix; this is where the surprise lives). - Build hours (loaded). - Ongoing maintenance hours/month (loaded) — the honest number, not the hopeful one. - Cost of a silent outage: what does one undetected failure cost, and how likely is it? - Switching cost if you outgrow the platform (low portability = high later pain). - The babysitting tax: a "free" flow a senior engineer nurses monthly is not free. -
orchestration-and-antipatterns.md 6.6 KB
# Orchestration strategy & anti-patterns — depth for §4–§5 Engine-agnostic policy. These are the guarantees you *require of* an automation before it ships and the failure modes you audit for when a fleet is breaking. The per-platform **recipes** (which node, which toggle, exact retry settings) live in `../../automation-flows/SKILL.md` and each platform skill — this file is the discipline, not the wiring. ## The orchestration guarantees ### 1. Idempotency & dedup Event delivery is **at-least-once**: webhooks are re-sent, retries replay, users double-click. Assume every event can arrive two or more times. A non-idempotent write executed twice = double charge, duplicate row, duplicate email. - Require a **stable dedup key** — the provider's event id or a natural external id, never a timestamp or a hash of "roughly the same thing". - Check the key **before** any non-idempotent action; record it **after** success. - Prefer idempotent operations where the platform/API allows (upsert-by-key beats blind insert; idempotency keys on payment APIs). ```text Bad: event → create record (re-delivery → two records) Good: event → seen(key)? → if new: create → mark key seen ``` ### 2. Error paths, retries & backoff Every automation has a defined destination for failure. Classify the error first: - **Transient** (network timeout, 429 rate-limit, 5xx) → retry with **exponential backoff + jitter**, a **max attempts cap**, and a total time budget. Respect `Retry-After` when present. - **Permanent** (400, 401/403 auth, 404, validation) → do **not** retry; retrying just burns runs and delays the alert. Route straight to the dead-letter path. Distinguishing the two is the single highest-leverage error decision — blind retry-everything is a common cause of both runaway bills and masked auth failures. ### 3. Dead-letter After retries are exhausted, the failed item must land somewhere **durable and inspectable** — a queue, a table, the platform's "incomplete executions" store — with enough context (the payload, the error, the timestamp) to diagnose and **replay** it. Failures dropped on the floor are silent data loss; you find out when a customer does. ### 4. Observability & alerting - Emit **success/failure counts** per automation; a run rate that quietly drops is a signal. - **Heartbeat scheduled jobs and alert on silence.** A cron that stops firing produces no error — it produces *nothing*, which is the failure you never hear about. Alert when an expected run doesn't happen within its window (dead-man's-switch). - Route alerts to a **named owner / on-call channel**, not an inbox nobody reads. An alert nobody receives is not observability. ### 5. Secrets handling - Store credentials in the **platform credential store or a vault**, referenced by name — never pasted inline in a node/step where an export or a screenshot leaks them. - **Least privilege** — scope each token to what the automation actually needs. - **Rotation plan** — tokens expire; a rotation nobody owns becomes a 2 a.m. outage. Engineering detail lives in `../../secure-coding/SKILL.md`. ### 6. Versioning & environments - Separate **dev/staging from prod**. Test against non-prod data; don't debug against live customers. - **Export/snapshot before editing** a live automation — editing prod is a production change. - Keep a **rollback**: the previous known-good version, and a way to get back to it fast. ### 7. Avoid fan-out storms - **Cap concurrency**, batch where possible, and **rate-limit** outbound calls to respect downstream quotas. - Watch for **echo loops**: flow A writes a record that triggers flow B that writes a record that triggers flow A. Break the cycle with a guard (a marker field, an event-source check). - Guard against one event **fanning out into thousands** of downstream calls that trip rate limits or detonate the bill. Debounce bursts. ## Anti-patterns — detection & remediation | Anti-pattern | How you notice it | Remediation | | --- | --- | --- | | **No error path** | first API blip and the run just stops; found out from a customer | define a failure destination + retry policy + alert before shipping | | **Silent failure** | dashboards green, but the work isn't happening; false confidence | alert on failure AND on silence (heartbeat / dead-man's-switch) | | **No kill-switch** | a misfiring flow floods a downstream system and there's no fast stop | every outward-writing automation gets a documented pause/disable toggle (and whoever's on-call knows where it is) | | **Over-automation** | maintenance time exceeds the manual time it replaced | re-run the §1 gate; retire it or revert to a checklist | | **Automating a broken process** | the automation "works" but the outputs are still wrong, just faster | stop, standardize the process on paper, then re-encode | | **Hidden single point of failure** | it breaks when one person leaves or one personal token expires | service accounts, documented dependencies, no personal-account glue | | **No owner** | nobody can say how it works or notices when it breaks | assign an owner + a one-page runbook (what it does, deps, how to pause, how to replay) | | **Choosing platform by familiarity** | the bill "exploded"; billing unit mismatched the run shape | re-select by billing unit for the real shape (see platform matrix) | | **Retry-everything** | runaway run counts, masked auth failures | classify transient vs permanent; only retry transient, with backoff + cap | | **Secrets inline** | credentials visible in an export/screenshot; can't rotate cleanly | move to credential store/vault, reference by name, plan rotation | ## Fleet-health audit ("our automations keep breaking") When someone arrives with a *fleet* that keeps breaking rather than one flow, walk this list — the breakage is almost always a missing discipline, not a bad node: 1. **Inventory & ownership** — is there a list of every automation with a named owner? Orphans first. 2. **Error paths** — how many have a real failure destination + alert vs die silently? 3. **Idempotency** — which do non-idempotent writes with no dedup key? (double-processing symptoms) 4. **Observability** — is anything watching run rates and heartbeats, or do you learn from customers? 5. **Secrets/expiry** — any recent breaks that trace to an expired token or a personal account? 6. **Environments** — are people editing prod directly with no snapshot/rollback? 7. **Fan-out / loops** — any echo loops or bursts tripping rate limits? Fix the systemic gap (usually error paths + observability + ownership) rather than patching flows one at a time. Then feed the corrected guarantees into the build skill as requirements. -
platform-selection-matrix.md 5.4 KB
# Platform selection matrix — n8n · Make · Zapier · Power Automate Depth for §3. This is a *routing* reference: decide, then hand off to the platform's own skill (`../../n8n/`, `../../make/`, `../../zapier/`, `../../power-automate/`) for the live API/MCP work. All pricing/version figures move — re-check the vendor page at author time. ## Billing unit is the decision axis Cost at scale is dominated by the unit each platform charges on. The *same* logical flow costs wildly different amounts depending on shape (few-long vs many-short) because of this. | Platform | Billing unit | What counts as 1 | Cost shape it favors | | --- | --- | --- | --- | | **n8n** | execution | one whole workflow run, any number of steps | few, long, complex flows; high volume; self-host = $0 | | **Make** | credit | each module action (renamed from "operations" 2025-08-27, converted 1:1) | moderate multi-step flows | | **Zapier** | task | each action step that runs | many, short, simple flows | | **Power Automate** | seat / flow | per-user or per-flow premium licensing; standard connectors bundled with many M365 plans | teams already paying for Microsoft 365 | **Worked contrast** — a 10-step flow run 10,000×/month: - **n8n:** 10,000 executions regardless of step count; self-hosted = $0 infra-aside. - **Make:** ~100,000 credits (10 module actions × 10k). - **Zapier:** ~100,000 tasks (10 actions × 10k) — pushed into high tiers fast. - **Power Automate:** governed by seat/flow licensing rather than per-action, so the math is about who's licensed and premium-connector usage, not run count. For complex high-volume flows, n8n's per-execution model commonly cuts cost 80–90% vs per-task/credit platforms. For a Microsoft-centric team, Power Automate can be effectively "free at the margin" because it rides existing M365 licenses. ## Full comparison | Axis | n8n | Make | Zapier | Power Automate | | --- | --- | --- | --- | --- | | Free / entry | self-host free; cloud Starter ≈ €20/mo | ≈ 1,000 credits free; Core ≈ $9/mo | 100 tasks/mo free; Pro ≈ $19.99/mo | bundled in many M365 plans; premium add-ons | | App breadth | ~1,000 nodes + generic HTTP + code | ~1,500, often deep per app | ~9,000+ apps — widest catalog | strong on MS + growing 3rd-party connectors | | Self-host / data residency | **yes — your infra** | no (EU/US regions) | no | Microsoft cloud / Dataverse regions | | Who runs/maintains it | technical; you run the box | mid; visual but rich | non-technical-friendly | IT/business in a Microsoft shop | | Portability | high (export/import JSON + self-hostable runtime) | medium (blueprint export/import, runs only on Make) | low (no portable export) | low (tied to MS solutions) | | Complex logic / code | code nodes, sandboxed by default (n8n 2.0) | functions + routers | limited (Code by Zapier) | expressions + Azure Functions escape hatch | ## Programmatic + MCP maturity (honest) The suite's platform skills drive these; here's the strategy-level read on how far each goes. - **n8n** — public REST API can CRUD and activate workflows programmatically; strongest fit for code-driven ops and GitOps-style management. MCP is community-driven and evolving. Best pick when you want automations that are themselves managed as code. → `../../n8n/SKILL.md`. - **Make** — management API plus maturing MCP / AI-agent support; you can manage scenarios via API to a useful degree. → `../../make/SKILL.md`. - **Zapier** — rich API for *running actions*, and **Zapier MCP** exposes thousands of app actions to AI agents. **Hard limit: there is no public API to create/author a Zap** — Zap building is UI-only. You can trigger and act, not construct. → `../../zapier/SKILL.md`. - **Power Automate** — Flow *management* APIs, plus solutions/ARM/ALM for deploying flows as packaged artifacts. **Hard limit: no clean public API to author an end-user's personal "My flows"**; the supported path is solution-based deployment, not arbitrary per-user flow creation. → `../../power-automate/SKILL.md`. **Do not build a plan that depends on API-creating a Zap or a personal Power Automate "My flow."** Both are UI-authored. If automation-of-the-automation is a requirement, that requirement points at **n8n** (and to a lesser extent Make). ## Ecosystem fit — the tie-breaker after billing - **Microsoft 365 shop** (Teams, SharePoint, Outlook, Dataverse, Excel Online) → **Power Automate**. The native integration and bundled licensing usually beat everything else on both cost and friction. Don't bolt a third-party platform onto a Microsoft estate without a reason. - **Data must stay on your own infrastructure** (compliance, residency, air-gapped) → **n8n** self-hosted is the only option here that keeps data on your box. - **Obscure or long-tail app**, non-technical owner → **Zapier** for catalog breadth. - **Deep per-app capability at mid volume** → **Make**. ## Routing summary ```text Microsoft 365 estate? → Power Automate Data must stay on our infra? → n8n (self-host) High volume / complex / code? → n8n Need to manage flows as code? → n8n (Make second) Widest app catalog / non-tech? → Zapier Deep modules, mid volume? → Make ``` Then route to that platform's skill for the build/operate work. Portability note: n8n's exportable JSON makes exit cheap; the other three lock you in more, so weigh that if you expect to outgrow the choice.
-
-
SKILL.md 13.2 KB
--- name: automation-strategy description: "Use when deciding whether a process is worth automating, sizing ROI and build-vs-buy, choosing an automation platform, or diagnosing why a fleet of automations keeps breaking — the decision layer before anyone builds. NOT building the flow (that is `automation-flows`, or `n8n` / `make` / `zapier` / `power-automate` to drive a live platform)." tags: [automation, strategy, roi, build-vs-buy, platform-selection, orchestration, idempotency, n8n, make, zapier, power-automate] recommends: [automation-flows, n8n, make, zapier, power-automate] profiles: [core, full] origin: risco --- # Automation strategy — decide whether, where, and how before anyone builds This is the discipline layer of the automation suite. It answers three questions that come *before* the first node is wired: **should this be automated at all?**, **on what platform, and is a platform even the right tool vs code?**, and **what orchestration guarantees does every automation have to meet so the fleet doesn't rot?** No platform API, no importable JSON, no code lives here — the moment the decision is made, you route to the skill that builds it. Building the flow (design + importable artifact) → `../automation-flows/SKILL.md`. Driving a specific platform's live API/MCP → `../n8n/SKILL.md`, `../make/SKILL.md`, `../zapier/SKILL.md`, `../power-automate/SKILL.md`. A typed API client in code → `../api-connector-builder/SKILL.md`. Receiving webhooks in your own app → `../webhooks/SKILL.md`. Every fact about pricing and platform capability below moves fast. Treat the `≈` figures as hedges and re-check the vendor page at author time; the honest limits (Zapier can't create Zaps via public API; Power Automate can't create end-user "My flows" cleanly via API) are called out where they matter and expanded in the references. ## 1. Should you automate at all? Most "we should automate this" instincts are wrong — not because automation is bad, but because the process underneath isn't ready, or the payoff isn't there. Gate on four factors, all of them, before writing a line: - **Frequency** — how many times per week/month does this run? A once-a-quarter task rarely earns its maintenance cost. - **Time saved per run** — minutes of human toil removed each time. Two minutes × 1,000 runs beats an hour × 3. - **Error rate of doing it by hand** — high manual error rate (typos, missed steps, missed SLAs) is often a bigger win than the time saved. Automation's real product is *consistency*, not just speed. - **Stability of the process** — how often do the steps, the systems, or the rules change? An unstable process is the one thing that turns automation into a liability. Rough gate: automate when it is **frequent AND stable AND** either saves meaningful time-per-run *or* removes a costly manual error rate. If it's frequent but unstable, or high-value but requires judgment on every run, do not automate yet — see below. Full scoring rubric with worked numbers: `references/automate-decision-and-roi.md`. ### The trap: automating a broken or unstable process Automating a bad process just makes you produce bad output faster, and now nobody can see the badness because a machine hides it. **Fix and standardize the process first, then automate the standardized version.** If the steps aren't written down the same way twice, if it "depends", if three people do it three ways — you are not ready. Standardize on paper, run it manually a few times against the written steps, *then* encode it. Automating chaos gives you automated chaos with a maintenance bill. ### Human-in-the-loop: what you deliberately do NOT fully automate Some steps should stay manual by design, wrapped in an approval gate rather than removed: - **Judgment calls** — pricing exceptions, hiring decisions, content that carries brand/legal risk, anything where "usually right" isn't good enough. - **Irreversible or high-blast-radius actions** — sending money, deleting data, emailing your whole list, publishing publicly, terminating accounts. Automate the *preparation*, require a human click for the *commit*. - **Low-frequency, high-stakes** — the payoff is small and the cost of a silent wrong run is huge. The pattern is *assisted*, not autonomous: the automation gathers, drafts, and stages; a human approves the commit. Cheap insurance against the day the automation is confidently wrong at scale. ## 2. ROI and build-vs-buy ### Payback math ```text build cost = hours to build × loaded hourly rate monthly saving = (runs/month × time saved per run × loaded rate) − monthly platform/run cost − monthly maintenance cost ← the line people forget payback (months)= build cost / monthly saving ``` Two rules of thumb: if **payback > ~12 months**, the process is probably too rare or too cheap to bother; if **maintenance eats > ~30–50% of the gross saving**, you've built a pet, not a tool. Maintenance is not optional — APIs change, auth expires, edge cases surface. Budget it explicitly or the ROI is fiction. Worked examples: `references/automate-decision-and-roi.md`. ### No/low-code platform vs custom code | Choose a no/low-code platform when… | Choose custom code when… | | --- | --- | | Gluing 2–8 SaaS apps with pre-built connectors | Logic is genuinely complex, algorithmic, or stateful | | Business owner needs to read/edit the flow | High volume where per-run platform billing dominates TCO | | Speed to first version matters more than control | You need version control, tests, code review, CI on the logic | | Volume is modest and connectors exist | The vendor has no connector and you need deep API control | The honest middle: platforms win on **time-to-first-version and readability**; code wins on **complex logic, testability, and cost at scale**. Many mature setups are hybrid — a platform orchestrates and a small custom service (or `api-connector-builder` client) does the hard part. ### Total cost of ownership TCO is not the sticker price. Count: the **platform bill under your real volume and its billing unit** (this is where surprises live — see §3), build time, ongoing maintenance, the cost of an outage when it fails silently, and the switching cost if you outgrow the platform. A "free" automation that a senior engineer babysits monthly is not free. ## 3. Platform selection Pick on **billing unit first** — it's the axis that dictates cost at scale, and choosing on familiarity instead is the classic 10× mistake. Then filter on data residency, ecosystem, programmatic maturity, and portability. This is a routing decision: once chosen, hand off to that platform's skill. | Axis | n8n | Make | Zapier | Power Automate | | --- | --- | --- | --- | --- | | **Billing unit** | per **execution** (whole run = 1, any # of steps); self-host = $0 | per **credit** (each module action = 1; was "operations", renamed 2025-08-27) | per **task** (every action counts) | per **user** or per **flow** (premium); bundled with many M365 plans | | **Best cost shape** | few long/complex flows, high volume | moderate multi-step flows | many short simple flows | teams already paying for M365 | | **Self-host / data residency** | **yes** — your infra, your data | no (EU/US regions only) | no | Microsoft cloud / Dataverse regions | | **Ecosystem fit** | technical teams, raw HTTP, code nodes | mid-market, deep per-app modules | widest app catalog (~9,000+) | **native to Microsoft 365** — Teams, SharePoint, Outlook, Dataverse | | **Programmatic + MCP maturity** | public REST API to CRUD workflows; community MCP; strongest for code-driven ops | management API + MCP support maturing | rich API for actions + Zapier MCP to expose actions to agents — **but no public API to create Zaps** | Flow management APIs + solutions/ARM — **but no clean public API to create end-user "My flows"** | | **Portability** | high — export/import JSON + self-hostable runtime | medium — blueprint export/import, but runs only on Make | low — no portable export | low — tied to MS solutions | **Routing shortcuts:** already living in Microsoft 365 → **Power Automate** (`../power-automate/SKILL.md`). Data must stay on your own infra, or high volume where execution-billing wins → **n8n** (`../n8n/SKILL.md`). Obscure app, non-technical owner → **Zapier** (`../zapier/SKILL.md`). Deep per-app modules, mid volume → **Make** (`../make/SKILL.md`). **Two capability limits to state honestly, not paper over:** you can drive Zapier's actions programmatically and via MCP, but you **cannot create a Zap through a public API** — Zap authoring is UI-only. Power Automate exposes flow *management* and solution deployment, but there is **no clean public API to author an end-user's personal "My flows"**. If a plan depends on API-creating either, the plan is wrong. Full matrix, MCP notes, and caveats: `references/platform-selection-matrix.md`. ## 4. Orchestration strategy (engine-agnostic) These are the guarantees every automation must meet regardless of platform. This is the *policy*; the per-platform *recipe* lives in `../automation-flows/SKILL.md` and the platform skills. Deep dive: `references/orchestration-and-antipatterns.md`. - **Idempotency & dedup.** Delivery is at-least-once and retries replay; assume every event can arrive twice. Require a **stable dedup key** (event id / external id) checked *before* any non-idempotent write (charge, email, insert). No dedup on a non-idempotent write = duplicate charges waiting to happen. - **Error paths, retries & backoff.** Every automation has a defined destination for failure. Retry **transient** errors (timeouts, 429/5xx) with backoff and a cap; do **not** retry **permanent** errors (400/validation) — route them out. A flow with no error path isn't finished. - **Dead-letter.** After retries are exhausted, the failed item lands in a durable place a human can inspect and replay — a queue, a table, an "incomplete executions" store. Dropped-on-the-floor failures are invisible data loss. - **Observability & alerting.** Emit success/failure counts; put a **heartbeat on scheduled jobs and alert on silence** (a cron that stops firing is the failure you never hear about). Route alerts to a named owner, not a dead inbox. - **Secrets handling.** Credentials in the platform's credential store / a vault, referenced by name — never pasted inline where an export leaks them. Least privilege, and a rotation plan. Engineering detail → `../secure-coding/SKILL.md`. - **Versioning & environments.** Separate **dev/staging from prod**; export before you edit; keep a rollback. Changes to a live automation are production changes and deserve the same care. - **Avoid fan-out storms.** Cap concurrency, batch, and rate-limit. Watch for **echo loops** (flow A writes a record that triggers flow B that triggers flow A) and for a single event fanning out into thousands of downstream calls that trip rate limits or blow the bill. ## 5. Anti-patterns | Anti-pattern | Why it bites | Do instead | | --- | --- | --- | | **No error path** | First API hiccup, the run dies; you learn from an angry customer | Define a failure destination + alert before shipping | | **Silent failure** | Failures that don't alert erode trust worse than no automation — everyone assumes it worked | Alert on failure *and* on silence (heartbeat) | | **No kill-switch** | When it misfires at scale, there's no fast way to stop the bleeding | Every outward-writing automation gets a documented pause/disable toggle | | **Over-automation** | Automating rare / low-value / judgment work costs more to maintain than doing it by hand | Gate on §1; leave judgment steps human-in-the-loop | | **Automating a broken process** | You produce bad output faster and hide the badness behind a machine | Standardize the process first, then encode it | | **Hidden single point of failure** | One personal token / one undocumented flow everyone depends on | Service accounts, documented dependencies, no personal-account glue | | **No owner** | Orphaned automations rot; nobody notices the break or knows how it works | Name an owner + a one-page runbook per automation | | **Choosing platform by familiarity** | Bill 10× higher than the right billing unit; "our automation bill exploded" | Pick by billing unit for your run shape (§3) | ## Checklist - [ ] Ran the §1 gate: frequency × time-saved × error-rate × **stability** — and confirmed the process is standardized, not chaos being paved. - [ ] Identified any judgment / irreversible / high-blast-radius step and kept it human-in-the-loop with an approval gate. - [ ] Computed payback including **maintenance**, and confirmed it's under the rule-of-thumb ceiling. - [ ] Decided platform vs custom code on the build-vs-buy criteria, not on what's familiar. - [ ] Chose the platform by **billing unit for the real run shape**, then data residency / ecosystem / portability — and did NOT assume a capability the platform lacks (Zap creation, PA "My flows"). - [ ] Specified the orchestration guarantees §4 requires (idempotency key, error path, dead-letter, heartbeat/alert, secrets store, dev/prod split, fan-out caps) as requirements handed to the build skill. - [ ] Every automation has a named owner, a runbook, and a kill-switch. - [ ] Routed the actual build to `automation-flows` / the chosen platform skill — no flow was built inside this skill.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.