Claude Skill

os-what-could-go-wrong

ALWAYS invoke this skill before anything hard to undo gets agreed to - a contract, a purchase, a migration, a launch, a price change, a reorganisation - and whenever the user asks "what could go wrong", "what are we missing", "poke holes in this", or for a premortem or a red team

LLM Mart · 0 points · 2 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download kharmanskyi-open-steps-skills_os-what-could-go-wrong-94b1cba.zip · 24 KB
Part of kharmanskyi/open-steps — 7 skills

Install

skills CLI npx skills add https://github.com/kharmanskyi/open-steps/tree/main/skills/os-what-could-go-wrong
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install kharmanskyi-open-steps@llmmart
Git git clone https://github.com/kharmanskyi/open-steps.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole kharmanskyi/open-steps collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

os-what-could-go-wrong

Assume it already failed, then work backwards to find out why - while there is still time to change it. The attack runs in a fresh agent that had no part in the decision, because an agent that helped shape one reviews it far too gently: it defends its own reasoning, and it misses the thing that kills the plan out of politeness.

os-ask-simple screens a choice before the user picks one. This one runs after the choice is made and before it can no longer be taken back.

Language

Write in the language the user speaks in this session, detected from the conversation. Names, figures and identifiers stay as they are. The agent you dispatch cannot see this conversation and does not inherit the writing style, so the language has to travel with the handover - see step 2.

Step 1 - write down what is actually being decided

Find the decision first. Given as text or a document, that is it. Asked at the end of a discussion, it is the decision the discussion arrived at, and say which one you took it to be. If neither, ask which decision to attack.

Then fill every line. This is the only thing the fresh agent will ever see.

DECISION BRIEF
What will be done: three to seven sentences
What it is for: the problem it solves, and what success looks like, measured
  wherever it can be
The main moves: the money, the people, the systems, the dates
What is fixed: the constraints, plus the surrounding facts that matter
Who it lands on: who and what is affected if this goes wrong
What cannot be undone: which parts are one-way
When we would know: the date success or failure actually gets judged
What we know: facts from documents, data, the repository, past incidents,
  each with where it came from
What we are assuming: every gap nobody could close, written as an assumption

Close the gaps in this order, and stop as soon as a gap is closed.

  1. Look it up yourself. Documents, data, the repository, the last report in ~/.claude/open-steps/reports/, what went wrong last time.
  2. Ask, but earn the ask. Only a gap where guessing wrong would change the verdict, one question at a time through os-ask-simple, three at most. Nobody there to answer, or an answer that would not move the verdict -> 3.
  3. Write the guess down as a guess. Put the assumed value in its line, mark it (assumed), and repeat it under "What we are assuming".

The brief states facts and open questions. It never makes the case for the decision: a brief that argues gets a report that agrees.

Step 2 - hand it to an agent that had no part in it

Pick the depth, say which in one line, and carry on; the user can change it.

Depth When
Full Hard to undo, or being wrong costs money, trust or data beyond one team
Quick Reversible, and cheap to be wrong about, whoever it touches

When in doubt, Full. The cost of a full look is a few minutes; the cost of a quick look at a one-way door is the door.

Dispatch one general-purpose agent. Send it four things and nothing else: the analysis prompt copied exactly from the bottom of this skill, MODE: Full or MODE: Quick, LANGUAGE: <the language above>, and the brief. Never send an instruction to run this skill: the fresh session would load it and start over.

One agent, not several. Where no fresh agent can be started, say so in the first line of the report, name the tool, and never use the word independent.

Step 3 - give it to the user straight

The user never sees what the agent returned: a tool result is visible only to you. So your final message is that report, copied whole, first line to last. One line goes before it: the depth, and whether a fresh agent ran. Anything of your own comes after it, never instead of it: no summary in its place, no "details above", no reassurance the analysis did not earn, no dropped card because the user seemed committed. Bad news that arrives late is worth nothing.

Then offer to turn the "Fix before you commit" list into real things: edits to the plan, tickets, an owner and a date per item, a reminder for each early warning. "Go ahead" is delivered just as plainly.

Hard rules

  1. The agent that helped decide never attacks the decision. Dispatch a fresh one every time, even when you already hold the whole thing in context. Skipping this does not save a step, it changes the answer.
  2. No quota of risks. Publish what has a real chain behind it and nothing else. Two well-evidenced risks beat six padded ones, and "only two survived" is a finding worth saying out loud.
  3. The verdict is decided last and printed first. Never make the reader assemble it from the risks.
  4. Every area of the sweep is accounted for, including the ones that produced nothing. An area nobody mentions and an area nobody checked look the same to the reader.
  5. A skipped section keeps its one line saying why.
  6. A lease and a database migration get the same treatment. This is not a technical review; the nine areas apply to both, and the money and people ones are where technical decisions usually actually fail.
  7. "The plan is sound" is a legitimate answer once the attack has run. It is never a substitute for running one.
  8. The final message is the report itself. The user cannot see what the agent returned; a summary of it, however good, is not it.

Known gotchas

  • No date to be judged by means no premortem. Pick a date that fits the decision, and say you picked it.
  • The brief is where this is won or lost. The fresh agent sees nothing else.
  • Fewer than three risks is often the right answer. Keep the record of what was checked.
  • Something reversible and cheap does not need this. Quick look, or say so.
  • "Try it small first" is not a soft no. Test unknowns before money moves.
  • The user may go ahead against all of it. Note it once, set the tripwires up if they want them, and do not re-argue the report.

The reasoning behind these is in references/why-these-rules.md.

The analysis prompt, verbatim

The block below is the whole of references/premortem-prompt.md, inlined when this skill loads through scripts/prompt.sh, so no file has to be read at dispatch time. If the block shows a literal command instead of the prompt, this harness does not run inline commands: open that file next to this one and use its full text.

!bash ${CLAUDE_SKILL_DIR}/scripts/prompt.sh

Files (open-steps)
  • references
    • 01-office-lease.md 14.3 KB
      # Example 1 - a three-year office lease
      
      Source: an invented decision, written for this example. No real company,
      nothing confidential, no third-party material. Both halves are real agent
      output on the same facts: the BEFORE is what a general-purpose agent answered
      when the decision was put to it with no skill in play, the AFTER is what this
      skill produced. Both are abridged wherever you see an ellipsis, and punctuation
      is normalised to this repository's style; no wording is changed.
      
      Kept because the BEFORE is good. That is the point of it. It is not a lazy
      answer that a better prompt would fix - it is a careful, well-informed answer
      that the person still cannot act on or audit.
      
      ## BEFORE - the same decision, no skill (~90 lines)
      
      > ## The short version
      >
      > You're committing three years and roughly **$1.7-2.0M all-in**, backed by a
      > **director's personal assets**, to test a hypothesis you cannot measure for
      > 12-18 months, with **7.7 months of cash in the bank**, in a market that's
      > discounting 15-20%, under a three-week deadline set by the counterparty.
      >
      > Almost every one of those factors independently argues for slowing down...
      >
      > *(Working from your numbers only. Two assumptions worth confirming: that the
      > $310k burn excludes the office, and that $28k is net of service charge...)*
      >
      > ## 1. The cost model is missing 30-40%
      > ... **Realistic total $1.8-2.2M**... **Runway impact.** $2.4M / $310k = 7.7
      > months today... you're at **~5.8 months**...
      >
      > ## 2. The personal guarantee is the real tail risk
      > ... **It bites precisely when the company fails.**... **It creates a live
      > governance conflict.**...
      >
      > ## 3. Your own survey is telling you something you may be discounting
      > ... Expect **4-8 regretted departures in year one**...
      >
      > ## 4. The diagnosis may be wrong, and it's cheap to check
      > ... **55% retention on what denominator?** If it's 3 of 6 juniors, n=3 cannot
      > carry a $2M decision...
      >
      > ## 5. The structure is backwards on timing
      > ... **"Three weeks or it goes" is very likely manufactured.**...
      >
      > ## 6. Two execution details that quietly break the plan
      > ... **The space is too big if the mandate fails and too small if it
      > succeeds.**...
      >
      > ## What I'd actually recommend
      >
      > **Don't sign in three weeks.** Instead: 1. Buy the experiment, not the
      > conclusion... 2. Fix the mentorship problem in parallel... 3. If you do
      > lease, renegotiate as a tenant in a soft market...
      >
      > **Tripwires - agree these before you sign anything:** don't proceed if fewer
      > than ~two-thirds of the 12 remote-contract staff agree in writing to new
      > terms; if more than two senior mentors signal departure; if post-outlay
      > runway is under 9 months without committed funding...
      
      **What is wrong with it**
      
      | Defect | Where |
      |---|---|
      | No verdict anywhere | Six sections of analysis, then the answer as the last heading. The reader has to reach the bottom to learn what to do |
      | You cannot tell what was skipped | Six areas, chosen by the writer. Competitor response is never mentioned - considered and empty, or never looked at? Nothing says |
      | Nothing is scored | Every risk is prose. Which is likeliest, which is worst, and which you would not see coming are three different questions, and the reader is left to separate them |
      | The load-bearing belief is a footnote | "Two assumptions worth confirming" in italics under the opening. The one belief the whole decision rests on is never isolated |
      | Warnings nobody can act on | "Don't proceed if more than two senior mentors signal departure" - measured how, checked when, by whom? |
      | The ordinary failures are never swept | The class of decision is never named, so the ways this kind of deal usually dies are never worked through as a list |
      
      All six are cured somewhere below. The first two are cured on the screen; the
      other four are cured underneath it, which is why this example shows all four
      layers rather than stopping at part one.
      
      ## AFTER - part one, the screen (24 lines)
      
      > # What could go wrong - signing a 3-year, $1.04M city-centre office lease and mandating 3 days a week
      >
      > | | |
      > |---|---|
      > | **Verdict** | **Think again** |
      > | **Why** | At $310k/month burn against $2.4M cash, the company runs out of money around month 6 after signing - roughly ten months before month 12, the date it planned to grade this decision by. |
      > | **Worst case** | $348,000 of cash gone before anyone sits down, ~$1.04M of non-cancellable rent, one director personally liable for the unpaid remainder, 7-11 of 40 staff gone, and no measurement that can tell you whether any of it worked. |
      >
      > **The biggest risks**
      >
      > 1. The runway ends before the answer arrives. $348,000 out up front, burn
      >    rises ~11%, cash hits zero ~2.4 months after move-in; the success metrics
      >    are dated month 6 and month 12. (risk 1)
      > 2. Attrition lands on the people you can least replace, and it starts at the
      >    announcement - two months before move-in, when every dollar is already
      >    committed. 11 of 31 survey respondents already flagged themselves. (risk 2)
      > 3. The personal guarantee converts a company failure into one director's
      >    personal debt, and after signature that director's own interests argue
      >    against an honest month-6 review. (risk 3)
      >
      > **The thing nobody is questioning**
      >
      > That this is a decision about how the company works, with a cost attached -
      > rather than a financing decision with a culture story attached.
      >
      > **Fix before you commit**
      >
      > - Do not sign inside the three-week window. Ask for it to be extended; if the
      >   answer is no, walk and see whether the phone rings... (risk 6)
      > - Trade the personal guarantee against the six-month deposit. Landlords
      >   normally take one or the other. Worth $84,000 of cash or one director's
      >   house. (risks 3, 6)
      > - Decide the 12 remote contracts in writing, before any expectation is
      >   announced. Then check the arithmetic: with 12 exempt, 80% is unreachable
      >   until headcount passes 60. (risk 5)
      > - Start measuring cycle time today, from the existing tracker. One day of
      >   work. Without a baseline, "a measurable drop in time-to-ship" cannot be
      >   passed or failed at month 12. (risk 4)
      > - ...
      >
      > **The decision, rewritten**
      >
      > Do not sign this lease in this window. Spend the next quarter on the free
      > version of the same thesis: baseline cycle time from the tracker, put a named
      > mentor and daily pairing on every junior, and take 20-30 flexible desks or a
      > 12-month serviced office near where staff already cluster - well under
      > $150,000 a year, cancellable, no guarantee...
      
      ## AFTER - the outside view, which the BEFORE has no equivalent of
      
      > ### What usually kills decisions like this
      >
      > Two classes stacked: **a long, non-cancellable fixed commitment signed under
      > a deadline**, and **a return-to-office mandate imposed on staff who were
      > hired remotely**.
      >
      > Committed leases usually fail in six ordinary ways. (1) *The commitment
      > outlives the plan*... (2) *Runway math* - the lease survives the tenant...
      > (3) *Concessions left on the table because of a clock*... (4) *Fit-out
      > overruns and schedule slips*... (5) *Dilapidations and exit costs*... (6)
      > *The exit that isn't* - sublease rights look like an escape hatch and recover
      > a fraction, slowly, in the same weak market that made you want out.
      >
      > Mandates usually fail in four. (7) *Attrition concentrates in the most
      > marketable staff*... (8) *Compliance gets measured as attendance rather than
      > as the outcome the office was meant to buy*... (9) *The mandate is announced,
      > softened under pressure*... (10) *A symptom is attributed to the wrong
      > cause*...
      >
      > **Live here: all ten.** Growth to 65 is unfunded (1). Runway is 7.7 months
      > against a 36-month term (2). Zero concessions in a market listing 15-20%
      > below asking (3)...
      >
      > **Not live:** the space being too small - 6,000 sq ft is about right for 65
      > hybrid and generous for 40. Landlord insolvency mid-term is possible but
      > unevidenced... A regulatory shock that empties offices again is a tail risk
      > nobody can forecast, but note there is no break clause if one comes.
      
      The "not live" paragraph is the part that makes this checkable. A reader can
      disagree with it. Nothing in the BEFORE can be disagreed with in that way,
      because nothing in it says what was considered and dismissed.
      
      ## AFTER - the record of the sweep
      
      > ### What was checked
      >
      > - **Will people use it** - Two risks. The 80% target is arithmetically
      >   unreachable as written (risk 5), and the thing it is meant to buy was never
      >   diagnosed (risk 4).
      > - **Money** - Three risks. The runway (risk 1), the deal terms against the
      >   market (risk 6), and the sublet that doesn't recover what people think
      >   (risk 7).
      > - **Building it** - One risk. $180,000 across 6,000 sq ft is $30/sq ft, and
      >   the four-month window has no facilities hire in it yet (risk 8).
      > - **Running it day to day** - One risk... (risk 10).
      > - **The people involved** - Two risks... (risks 2, 3).
      > - **Things you depend on** - Folded into risks 6, 7 and 8: the landlord is
      >   the counterparty on the deadline, on sublet consent, and on holding
      >   $168,000 that may not be in escrow...
      > - **Legal and rules** - Folded into risks 3 and 5...
      > - **People misusing it** - One risk. Presence theatre keeps the metric green
      >   and the money flowing past the point the thesis failed (risk 9)...
      > - **What others do about it** - Real, but as a mechanism inside risk 2 rather
      >   than its own card: remote-first competitors treat a dated mandate as a
      >   recruiting signal...
      
      Nine lines, nine areas, including the ones that produced no card of their own
      and why. This is the part that answers "what did you not look at", and it is
      the reason the sweep can be trusted without the reader repeating it.
      
      ## AFTER - one risk card, to show the shape
      
      > ### 5. The 80% target is impossible as written, and the only way to make it possible breaks 12 contracts [Will people use it / Legal] - Fix before you commit
      >
      > - **How it happens:** 12 of 40 staff are on explicitly remote contracts... If
      >   they are exempt - which the brief assumes but has not decided - the maximum
      >   attendance is **28 of 40 = 70%**, at perfect compliance from everyone who
      >   can comply. The 80% target fails on announcement day, by definition... If
      >   instead they are *not* exempt, you are unilaterally changing an expressly
      >   agreed contractual term for 30% of the company... There is no third option,
      >   and the brief has chosen neither.
      > - **What we are assuming:** "That a three-day expectation is enforceable
      >   against staff hired remotely" - the author's own assumption...
      > - **What this is based on:** "All 40 staff, of whom 12 were hired on
      >   explicitly remote contracts..." "80% of staff in the office three or more
      >   days a week by month six."
      > - **How likely:** **High**, and for the arithmetic, certain. My estimate of
      >   the realistic month-6 figure... **20-23 people at 3+ days, about 50-58%**...
      > - **How bad:** **Serious.** The headline success criterion is unachievable,
      >   so the month-6 review has no honest way to pass...
      > - **Would you see it coming:** **Easily.** It is division. Nobody has done it.
      > - **How it shows up:** Announcement week: the first question in the all-hands
      >   is "does this apply to the remote team?", and there is no answer. Weeks
      >   2-8: exemption requests from non-exempt staff... Month 6: 50-58% against a
      >   target of 80%, and a debate about the denominator instead of the outcome.
      > - **Early warning:** *What to watch* - the exemption policy: written, dated,
      >   with named criteria and a named decider · *When to worry* - no written
      >   policy exists on the day of the announcement, or exemption requests exceed
      >   4 in the first month · *When to check* - before signature, and weekly for
      >   the first eight weeks after announcement · *What to do then* - publish the
      >   criteria and the appeal route the same day...
      > - **What to do about it:** Decide the exemption in writing before you
      >   announce anything, and restate the goal as a number that can be reached...
      >   **What gets worse:** a written exemption for 12 people makes the two-tier
      >   company explicit and visible... That resentment is real - and it exists
      >   either way; writing it down just makes it manageable.
      
      Three separate scores, and the one that does the most work is the cheapest:
      "Would you see it coming: Easily. It is division. Nobody has done it." The
      early warning names a signal, a threshold, a checkpoint and an action, so it
      can be handed to somebody. Compare the BEFORE's "if more than two senior
      mentors signal departure", which cannot.
      
      ## What this example changed in the format
      
      1. **The outside-view section exists because of this pair.** The first run
         attacked the reference class properly, but only inside individual cards, so
         nothing named it and a reader could not check whether the ordinary ways such
         deals die had been swept. Asked directly, the agent confirmed there was no
         standalone section. "What usually kills decisions like this" is now a
         section of its own, and cards cite it. On the re-run above it produced ten
         named killers and a "not live" list.
      2. **Part one stopped showing three numbered risk slots.** A template with
         exactly three slots reimposes the quota that three separate rules exist to
         remove. It now says at most three, one per surviving risk, and says what to
         write when nothing survived.
      3. **Every line of "Fix before you commit" carries its risk number.** Without
         them the list reads as loose advice. With them each item traces back to the
         mechanism that earned it, and an item with no card behind it is visible as
         padding.
      4. **The empty areas are worth as much as the full ones.** Here "things you
         depend on" and "what others do about it" produced no card of their own. Both
         still get a line saying where the mechanism went instead. That is the
         difference between a sweep and a list of whatever came to mind.
      5. **Still visible above: the report leaks specialist words.**
         "Dilapidations", "escrow", "covenant", "in escrow or the landlord's money" -
         all appear without saying what the person would actually see happen, which
         is the rule at the top of the analysis prompt. Part one is clean; the detail
         is not. The output here is kept as it came, because a record that gets
         edited stops being a record. The example where the rule holds further down
         is [example 2](02-clinic-database-move.md), a database move: its detail
         layer says what a clinic would see beside every specialist term but four,
         and it lists those four.
      
    • 02-clinic-database-move.md 21.5 KB
      # Example 2 - moving a clinic database over one Saturday night
      
      Source: an invented decision, written for this example. No real company,
      nothing confidential, no third-party material. Both halves are real agent
      output on the same facts, produced on 2026-09-13 with Claude Code 2.1.222 and
      Opus 5, headless, from an empty folder. The BEFORE is what a general-purpose
      agent answered with every skill switched off. The AFTER is what this skill
      produced: the dispatching agent wrote the brief, a fresh agent attacked it, and
      the report came back "as it came". Both are abridged wherever you see an
      ellipsis, and punctuation is normalised to this repository's style; no wording
      is changed.
      
      Two honest notes on the source. The BEFORE opened with a line saying it could
      not load this pack's skill - the maintainer's own instructions name it, so the
      agent tried and was refused - and then answered on its own; that opening line
      and the closing status line are left out, the answer between them is
      untouched. The fresh agent in the AFTER checked PostgreSQL's documentation over
      the web where the harness allowed it; two searches and two page fetches were
      refused, and the report says so at its end rather than hiding it.
      
      Kept because it is the technical case. Example 1 is a lease, and this pack
      says a lease and a database migration get the same treatment. Kept also
      because the words: example 1 leaks specialist terms in its detail layer and
      says so; this is the example where the rule holds further down, with the
      residue listed at the end.
      
      ## The decision, as both were given it
      
      > We run a subscription app for small medical clinics: 1,900 paying clinics,
      > about $410,000 of monthly revenue, a team of 6 engineers and no dedicated
      > database specialist. The database (Postgres 12, 640 GB) runs on a server we
      > manage ourselves in one data centre, and the hosting company is ending that
      > product line in November.
      >
      > The plan: move the database to a managed cloud database service (Postgres
      > 16) this Saturday, 02:00-06:00 local time, announced to customers as a
      > four-hour maintenance window. Steps: final full backup at 01:30, stop all
      > writes at 02:00, copy the data with the cloud vendor's migration tool
      > (measured at 3 hours 10 minutes on a copy of the data two weeks ago), switch
      > the app's connection string to the new service, run our 40-check smoke test,
      > reopen to customers at about 05:30. Rollback plan: if the smoke test fails
      > before 05:00 we point the app back at the old server; after 05:00 we roll
      > forward and fix problems on the new service.
      >
      > The engineer who ran the trial migration is on holiday from Friday; two other
      > engineers have read the runbook. Clinics are closed on Sunday; Monday 08:00
      > is the first heavy use. What is fixed: the host's November deadline, the
      > window (already approved by marketing and announced), and the budget ($3,200
      > a month for the new service, contract already signed). Success means Monday
      > passes with fewer than 10 support tickets about the migration and query times
      > at or below today's.
      
      The BEFORE was asked "What could go wrong with this plan?", the AFTER "Do a
      premortem on this plan."
      
      ## BEFORE - the same decision, no skill (~80 lines)
      
      > **Short version: the schedule has no slack, and the rollback deadline lands
      > before you can possibly know whether the migration worked.** Everything
      > below is ordered by how badly it hurts.
      >
      > ### 1. The arithmetic doesn't close - this is the structural flaw
      >
      > 02:00 stop + 3h10m copy = 05:10. Your rollback deadline is 05:00. So on a
      > nominal run you cross the point of no return *before the first smoke check
      > executes*... **Fix:** tie the abort to a milestone, not a clock. "If the copy
      > has not finished by 04:15, abort and point back at the old server." And take
      > the backup Friday evening, not 01:30.
      >
      > ### 2. Monday performance fails for reasons the smoke test can't see
      >
      > A freshly restored Postgres has **no planner statistics**. Without a full
      > `ANALYZE` after restore, Monday 08:00 gets sequential scans where you expect
      > index scans - 10x-100x slower queries on a database that is otherwise
      > perfectly correct... On top of that the new instance has a cold buffer
      > cache...
      >
      > ### 3. Postgres 12 → 16 is four major versions, and you haven't checked what breaks
      >
      > - **Extensions.** Whatever you use (`pg_trgm`, `uuid-ossp`, PostGIS, anything
      >   custom) must exist on the managed service at a compatible version...
      > - **Collation.** If the vendor tool does a physical/binary copy and the
      >   managed service runs a different glibc, text indexes arrive silently wrong
      >   - wrong query results and unique constraints that don't hold, on
      >   patient-name data... **Nobody on your team currently knows which mode the
      >   tool uses.** Find out this week.
      > - **Connection limits.** Managed Postgres caps `max_connections` by instance
      >   size... Check whether the plan includes a pooler.
      > - ...
      >
      > ### 4. What migration tools routinely don't carry over
      >
      > Roles, users and grants. Sequence positions (→ duplicate-key errors on the
      > first Monday insert). Materialized view *contents*. Large objects. Event
      > triggers. Each one is invisible to a read-only smoke test and lethal on
      > Monday.
      >
      > ### 5. Rollback is a sentence, not a plan
      >
      > - It only works if the old server took **zero** writes after 02:00. "Stop
      >   all writes" via app maintenance mode doesn't stop cron jobs, queue
      >   workers, webhooks, or integrations writing on a side path...
      > - How fast is "point the app back"? If it's a DNS name with a 3600s TTL, or
      >   a config change requiring a deploy, your rollback costs 20-60 minutes you
      >   don't have. Lower the TTL now.
      > - ...
      >
      > ### 6. The person who has actually done this is gone
      >
      > Two engineers who *read* the runbook is not two engineers who *ran* it. At
      > 04:30, when the tool emits an error the runbook doesn't cover, nobody in the
      > room has seen it before...
      >
      > ### 7. You're treating one shot as mandatory when you have seven
      >
      > It's mid-September; the host's deadline is November. Sunday is a free day -
      > clinics are closed - and your plan assigns it nothing...
      >
      > ## What I'd change
      >
      > ...
      >
      > 1. **Replace the stop-the-world copy with logical replication.** Stand
      >    Postgres 16 up as a logical replica of the 12 this week; let it catch up
      >    over several days... Caveat: tables without a primary key need `REPLICA
      >    IDENTITY` set, and DDL isn't replicated - so freeze schema changes until
      >    cutover...
      > 2. **If you keep the current approach:** backup Friday evening;
      >    abort-by-milestone at 04:15; `ANALYZE` and cache warm-up written into the
      >    window as explicit steps.
      > 3. ...
      > 6. **Give Sunday a job:** verification, `ANALYZE`, load replay against real
      >    Monday-shaped traffic - or the second attempt.
      
      **What is wrong with it**
      
      Less than in example 1, which is why it is worth reading. The arithmetic is
      right, the recommendation is the same one the skill reached, and the tripwire
      at 04:15 is real. What it lacks is the part a person who does not run
      databases could act on or audit.
      
      | Defect | Where |
      |---|---|
      | The verdict is a mood, not a decision | "The schedule has no slack" opens it and "What I'd change" closes it. Nowhere does it say think again, or go, or try it small; the reader infers the verdict from the tone |
      | You cannot tell what was skipped | Seven sections, chosen by the writer. Nothing on the law around patient data, nothing on who gains from silence, nothing on customer trust. Considered and empty, or never looked at? Nothing says |
      | Nothing is scored | Which failure is likeliest, which is worst and which you would not see coming are three questions. "Ordered by how badly it hurts" answers one of them, in prose |
      | The load-bearing belief is never isolated | Every section attacks a symptom of the same belief - that this is a copying job - and none names it |
      | The words are for an engineer | planner statistics, sequential scans, buffer cache, glibc, `max_connections`, pooler, `REPLICA IDENTITY`, DDL, a 3600s TTL: each one correct, none followed by what a clinic would see |
      | The warnings are half actionable | "Abort at 04:15" can be handed to somebody. "Check whether the plan includes a pooler" is a question with no threshold, no checkpoint and no owner |
      
      ## AFTER - part one, the screen (26 lines)
      
      > # What could go wrong - moving 640 GB of clinical records to managed PostgreSQL 16 overnight on 2026-09-19
      >
      > | | |
      > |---|---|
      > | **Verdict** | **Think again** |
      > | **Why** | On the plan's own numbers the copy ends at 05:10 and the last moment to roll back is 05:00 - the rollback branch can never be used, because the smoke test cannot start before the deadline it is measured against. |
      > | **Worst case** | Monday 2026-09-21, 08:00: 1,900 clinics start patient intake against a database that is slow, or sorting records wrongly, with no way back that does not delete Monday morning's clinical writes. Losing 5% of clinics costs about $20,500 a month, or $246,000 a year. |
      >
      > **The biggest risks** - worst expected damage first.
      > 1. After 05:30 there is no way back. Writes made on the new service exist
      >    nowhere else, so "point the app at the old server" on Monday silently
      >    deletes every appointment, note and charge entered since. In practice the
      >    team must fix forward no matter what they find. (risk 1)
      > 2. The window does not fit the work. Copy measured 190 minutes;
      >    writes-stopped to customers-in is 210 minutes. The copy alone ends at
      >    05:10 - ten minutes past the plan's own rollback deadline - leaving 20
      >    minutes for 40 checks. (risk 2)
      > 3. Statistics do not travel with the data. A 640 GB database with no
      >    statistics and a cold cache, hit by 1,900 clinics at 08:00 Monday, is the
      >    single likeliest cause of the exact failure the team is measuring. (risk 3)
      >
      > **The thing nobody is questioning**
      > That this is a copying job - that once the bytes are on the new service, the
      > work is done. It is a behaviour change: same rows, different query planner,
      > different sort order, different permissions.
      >
      > **Fix before you commit**
      > - Move the cut-over to 2026-09-26 or 2026-10-03. Seven Saturdays fall before
      >   November; six are spare. (risks 1, 2, 5)
      > - Run one full dress rehearsal, end to end, with a stopwatch - backup
      >   running, copy, connection switch, statistics, all 40 checks. Cost: about
      >   two engineer-days. (risk 2)
      > - Ask the vendor this week whether their migration tool supports continuous
      >   replication. If yes, pre-copy days ahead so the Saturday window is minutes,
      >   not hours. (risks 1, 2)
      > - Put `ANALYZE` (or `vacuumdb --analyze-in-stages`) in the runbook as a
      >   named, timed step before customers are readmitted. (risk 3)
      > - Confirm in writing that the new vendor has a signed data-processing
      >   agreement / BAA covering patient data, and confirm which country the data
      >   lands in. The contract is already signed; check before the data moves, not
      >   after. (risk 7)
      > - Either the engineer who ran the trial is on the plan, paid and reachable,
      >   or the date moves. (risk 5)
      >
      > **The decision, rewritten**
      > Move the cut-over to Saturday 2026-10-03, which still leaves four spare
      > Saturdays before the November host deadline. Pre-copy the 640 GB days in
      > advance using continuous replication, so the Saturday window contains only:
      > stop writes, wait for replication lag to hit zero, reset sequences, switch
      > the connection string, run statistics, run the 40 checks plus a replay of
      > real Monday-morning load... Before any of that, run one timed dress rehearsal
      > against a fresh production snapshot; if that rehearsal does not finish
      > everything by simulated 04:30, the plan is not ready for any Saturday yet.
      
      The BEFORE reached the same recommendation. The difference is that a person
      who has never run a database can read this screen, know what was decided, and
      hand each fix to somebody with a risk number on it.
      
      ## AFTER - the outside view
      
      > ### What usually kills decisions like this
      >
      > The class: a single-shot, overnight, cross-provider database cut-over with a
      > customer-announced fixed window, run by generalists, with a major version
      > jump bolted on. Decisions in this class usually die in six ordinary ways.
      >
      > - The copy runs longer in production than in rehearsal - contention,
      >   network, a cold target - and the team blows through its abort point rather
      >   than stop, because stopping feels like the failure.
      > - The rollback plan is decorative. It exists on paper but becomes unusable
      >   the moment users are readmitted, and nobody notices until they need it.
      > - The work after the copy - statistics, index rebuild, cache warm-up,
      >   sequence reset - is not in the timeline at all, because the timeline was
      >   built around the copy.
      > - The damage shows up on the first heavy-use day, not in the window, when
      >   concurrency arrives.
      > - The one person who has actually done it is unavailable.
      > - Silent data problems - sort order, encoding, sequences - surface weeks
      >   later, after backups of the good state have rotated out.
      >
      > ... All six are live here. What is **not** live: the usual "you should have
      > done it in stages instead of a big bang" objection does not fully apply,
      > because one database cannot be split across two providers for long. But the
      > staging that *is* available for this class - pre-copying with change capture
      > so the window holds only the switch - is not being used, and the brief gives
      > no sign it was considered.
      
      ## AFTER - the record of the sweep
      
      > ### What was checked
      >
      > - **Will people use it** - nothing credible. Clinics have no choice to make;
      >   there is no adoption question... Worth five phone calls, not a risk card.
      > - **Money** - produced part of risk 3. $3,200/month is 0.78% of monthly
      >   revenue; the cost is not the risk. Sizing the new instance to fit a budget
      >   described as "fixed" is.
      > - **Building it** - produced risks 2 and 6.
      > - **Running it day to day** - produced risks 1 and 4. Separately: the brief
      >   never says monitoring, alerting or backups exist on the new service. On
      >   Monday morning the team may have no graph to look at.
      > - **The people involved** - produced risk 5.
      > - **Things you depend on** - the vendor's support response time at 04:00 on a
      >   Saturday is not stated; a basic support tier means no human until Monday...
      > - **Legal and rules** - produced risk 7.
      > - **People misusing it** - thin, and the only thing that is gameable is the
      >   measurement itself: "fewer than 10 customer support tickets about the
      >   migration" is judged by whoever classifies tickets, and the same team owns
      >   the migration. A slow Monday classified as "general performance" scores a
      >   pass. Fix the instrument, not the people: count all tickets Monday against
      >   the four previous Mondays.
      > - **What others do about it** - nothing credible. A rival practice-management
      >   vendor gains a sales story from a public clinical-data incident, but the
      >   effect is slow churn, already counted in risk 1's damage.
      
      Nine lines, nine areas. Two of them produced nothing and say so; one found the
      success measure itself was the thing that could be gamed, which no card in the
      BEFORE was looking for.
      
      ## AFTER - one risk card, the most specialist topic in the report
      
      Sort order is where a database engineer says "collation" and stops. This is
      the card, in full, because it is the one that had to say what a clinic would
      see.
      
      > ### 4. Sort order changes and nobody notices [Running it day to day] - Fix before you commit
      >
      > - **How it happens:** Text comparison in PostgreSQL comes from the operating
      >   system's collation library. The old self-managed server and the new
      >   managed host will almost certainly not have the same version. When that
      >   changes, text ordering changes: patient names sort differently, range
      >   conditions on text return different rows, and indexes built under the old
      >   order can disagree with the new one. The output is wrong answers, not error
      >   messages. A clinic sees a patient list that is missing someone, or a search
      >   that finds nothing. Nobody files a ticket saying "collation" - they file
      >   nothing at all, or they file it in November.
      > - **What we are assuming:** That data arriving byte-identical means queries
      >   return the same answers. They do not, if the sort rules moved.
      > - **What this is based on:** PostgreSQL 16 documentation, checked directly:
      >   "A change in collation definitions can lead to corrupt indexes and other
      >   problems because the database system relies on stored objects having a
      >   certain sort order..."...
      > - **How likely:** Medium for visible wrongness. A cross-major-version
      >   migration must be logical, which means indexes are rebuilt on the target
      >   under the target's rules - that removes the index-corruption case. It does
      >   not remove the changed-ordering case, which is a real behaviour change with
      >   no warning attached.
      > - **How bad:** Severe. Wrong clinical record lists are worse than slow ones,
      >   and the wrongness has been accumulating by the time anyone spots it.
      > - **Would you see it coming:** Not until it is too late. There is no error, no
      >   ticket, no alert. Weeks can pass.
      > - **How it shows up:** Nothing on the night. Nothing Monday. Somewhere in
      >   weeks two to six, one clinic reports a patient who "isn't in the system",
      >   and the investigation finds it was never missing, just sorted out of the
      >   page they were looking at.
      > - **Early warning:** What to watch - run the same ten name-ordered and
      >   text-range queries on both servers and compare row-for-row · When to worry
      >   - any difference at all · When to check - during the rehearsal, and again on
      >   the night before readmission · What to do then - if the sorts differ,
      >   decide deliberately: rebuild affected indexes and accept the new order, or
      >   match the old collation on the target.
      > - **What to do about it:** Add a collation comparison to the rehearsal. Ask
      >   the vendor which collation provider and version the new service uses, and
      >   compare with the old server. What gets worse: if a rebuild is needed,
      >   `REINDEX` on 640 GB runs for hours - another reason the copy cannot own the
      >   whole window. **Fix before you commit** (the check is cheap; the rebuild, if
      >   needed, changes the schedule).
      
      Compare the BEFORE on the same topic: "text indexes arrive silently wrong -
      wrong query results and unique constraints that don't hold". True, and a
      clinic manager cannot picture it. "A clinic sees a patient list that is
      missing someone" they can.
      
      ## Where the words stay plain, and where they do not
      
      The rule at the top of the analysis prompt: where a term cannot be avoided,
      say what the person would actually see happen. How the detail layer did on
      the specialist ground this decision stands on:
      
      | The specialist thing | What the report says instead, or next to it |
      |---|---|
      | Planner statistics | "every table looks like an unknown size, so the planner picks full table scans where it used to pick index lookups... Queries that took 50 ms take seconds, connections pile up, and the system either crawls or stops" |
      | A cold cache | "the new machine's cache is empty, so the first reads all hit disk" |
      | Collation | "patient names sort differently... A clinic sees a patient list that is missing someone, or a search that finds nothing" |
      | Sequences | "sequences are not replicated and must be reset by hand at cut-over or Monday's inserts collide with existing IDs" |
      | Sub-processor, DPA, BAA | "the cloud vendor becomes a sub-processor. That usually requires a signed data-processing agreement or business associate agreement, a defined hosting country, and notice to customers" |
      | Rollback | "Executing the rollback deletes every clinical record entered on Monday morning across 1,900 clinics. So it will not be executed" |
      
      Four terms remain bare: "replica identity" and "large objects" in the
      mitigation line of risk 2, "95th-percentile query time" in the early warning
      of risk 3, and "vCPU" beside RAM and disk throughput in the same card. Each
      sits next to a sentence that does say what the person would see, so a reader
      who skips the term loses the mechanism and keeps the consequence. Part one has
      none. That is the residue, and it is smaller than example 1's, where the
      outside view itself carries "dilapidations" and the sweep carries "escrow"
      with nothing beside them.
      
      ## What this example changed
      
      1. **The prompt now travels through a script.** The first attempt at this
         example never ran the skill: Claude Code 2.1.222 refuses an inline `cat` of
         a file outside the session's working directory, and a refused inline
         command aborts the whole skill. The agent wrote its own analysis and said
         the skill was blocked, which is honest and not the skill. The prompt is now
         inlined through `scripts/prompt.sh`, which the same harness does not
         refuse; the run above is the one made after that change.
      2. **Nothing in the report format changed.** On a technical decision the shape
         held as written: the verdict first, nine lines for nine areas, one card per
         surviving risk (seven survived), the cross-cutting findings each in their
         slot, and "What holds" naming the one measurement the team got right.
         Rule 6 - a lease and a database migration get the same treatment - is what
         this example measured, and the money and people areas produced risks 3
         and 5, which is where the rule says technical decisions usually fail.
      3. **Evidence first, working as written.** The dispatching agent's brief gave a
         date for the end of PostgreSQL 12 support; the fresh agent checked it
         against postgresql.org and corrected it by a week, in the report, with the
         source named. A brief that argues gets a report that agrees; a brief that
         states a wrong fact gets it corrected, if the attacker is allowed to look.
      
    • premortem-prompt.md 10.4 KB
      # The premortem prompt
      
      The decision in the brief below was carried out exactly as planned. The date
      named in the brief has passed, and the result is a failure. Work backwards:
      find the most believable reasons it failed, then tell the person what to
      change while there is still time.
      
      You did not help make this decision and you owe its author nothing but the
      truth. A decision that survives an honest attack is a real finding - but that
      conclusion has to be earned by attacking first, never granted at the start.
      
      ## Who you are writing for
      
      Someone who has to act on this and may not work in the field the decision sits
      in. Write in the language named on the `LANGUAGE:` line; keep names, figures
      and identifiers as they are.
      
      - Plain words. Where a term cannot be avoided, say what the person would
        actually see happen.
      - Numbers a person can use: money, days, counts. Not "materially degraded".
      - Bad news is never softened, never moved to the middle of a sentence, and
        never traded away for a reassuring line somewhere else.
      - Short sentences beat complete ones. The reader is deciding something today.
      
      ## Method
      
      **1. Evidence first.** Read the brief. Where a load-bearing fact can be
      checked with the tools you have - a document, a number, a price, a past
      incident - check it before you speculate about it.
      
      **2. Look at the outside first.** Name the class this decision belongs to
      ("a three-year committed lease", "moving a live database", "raising prices on
      existing customers") and say what usually kills decisions of that class,
      including real failures you know of. Test this decision against those before
      inventing new ones. Most decisions fail in the ordinary way for their class.
      This becomes a section of its own in the report. It is the one part a reader
      can check your risks against, so it never dissolves into the cards.
      
      **3. Sweep every area, then publish only what survives.** Hunt for a way to
      fail in each of these:
      
      | Area | The question |
      |---|---|
      | Will people use it | nobody wants it, or not enough of them, or not soon enough |
      | Money | it costs more, earns less, or arrives later than the plan needs |
      | Building it | the work is harder, slower or more tangled than it looks |
      | Running it day to day | it works once and cannot be kept working |
      | The people involved | somebody has to behave in a way their own interests argue against |
      | Things you depend on | a supplier, partner, platform or tool moves, prices up or leaves |
      | Legal and rules | a contract, a licence, a regulator, a jurisdiction |
      | People misusing it | somebody games the mechanism because it pays to |
      | What others do about it | a competitor, an incumbent or a crowd reacts |
      
      A scenario earns a place in the report only when its chain - this decision
      leads to that, which leads to this, which is what kills it - is anchored in
      the brief or in evidence you checked. Three well-anchored risks with an honest
      record of the sweep beat seven padded ones. **There is no quota.** Fewer than
      three survivors is itself a finding: say so plainly.
      
      **4. Write one card per surviving risk, and fill every line.**
      
      ```
      ### <n>. <short name> [<area>] - <Fix before you commit / Worth fixing / Just watch it / Accept on purpose>
      
      - How it happens: the chain, and why it kills the decision rather than
        annoying you.
      - What we are assuming: the belief that turns out to be false.
      - What this is based on: the line from the brief you are relying on, quoted,
        or the outside evidence you checked, named.
      - How likely: Low / Medium / High, and why. Give a number only where the
        outside view supports one; otherwise say what would sharpen the guess.
      - How bad: Small / Serious / Severe / Fatal, and what is actually lost.
      - Would you see it coming: Easily / Only if you look / Not until it is too
        late.
      - How it shows up: first sign, then the next sign, then the damage, with
        rough timing.
      - Early warning: What to watch (the thing measured) · When to worry (the
        value that means trouble) · When to check (how often, or at which step) ·
        What to do then (the actual step, not "review").
      - What to do about it: the change, then what gets worse if you make it, then
        which of the four buckets it lands in.
      ```
      
      The three scores are separate answers. The scariest risk is often not the
      likeliest, and the one you would never see coming is often neither.
      
      **5. Then the findings that cut across all of them.**
      
      - **The thing nobody is questioning** - the single belief the author most
        likely does not know they hold, and what happens if it is wrong. One. Not a
        list. If everything is hidden, nothing is.
      - **The cheapest way to find out you are wrong** - the least expensive thing
        that could disprove the most load-bearing untested belief before the money
        is spent: what you think is true, the test, what it costs next to the whole
        decision, what counts as passing, what counts as failing.
      - **A flaw no fix can cure** - Yes or No. If yes, say it plainly and why
        patching it does not work. If no, write "No fatal flaw."
      - **Who benefits if this fails** - only where someone actually does: a
        competitor, the other side of a contract, a regulator, someone inside whose
        interests point the other way, anyone who profits by gaming it. Who they
        are, where they push first, their cheapest move that hurts most, and the
        blind spot it exploits. Where nobody does, write exactly: "Who benefits if
        this fails: nobody. Skipped."
      - **The standouts** - most likely, most damaging, hides the longest, hits the
        fastest, hardest to undo. One line each, and only where they are different
        risks. If one risk wins several, say that in one line instead.
      - **When to pull the plug** - where the decision happens in stages, one to
        three stop conditions ("if X has not happened by day N, stop rather than
        keep fixing"), each with why that number and why spending past it is a bad
        bet. Where the decision is a single act you cannot take back - a signature,
        a purchase - write: "When to pull the plug: there is no after. Every
        safeguard above has to fire before you sign."
      
      **6. The verdict comes last and gets printed first.** Decide it only after all
      of the above. One of:
      
      | Verdict | Means |
      |---|---|
      | Go ahead | The attack found nothing that changes the plan |
      | Go, but fix these first | Sound, conditional on named fixes |
      | Try it small first | The unknowns are testable and worth testing before full spend |
      | Think again | The plan as written does not survive; the goal might |
      | Do not do this | The goal is not reachable this way |
      
      ## What the report looks like
      
      Your final message is the report. Two parts. Part one is one screen and
      answers the question on its own; part two is for whoever wants the working.
      
      ```
      # What could go wrong - <the decision in a few words>
      
      | | |
      |---|---|
      | **Verdict** | <one of the five> |
      | **Why** | <one line naming the finding that decided it> |
      | **Worst case** | <what is actually lost, in money, time or trust> |
      
      **The biggest risks** - one line each, worst expected damage first, at most
      three. One line per surviving risk and no more: if only two survived, list two
      and say that only two did. If none survived, this heading is replaced by one
      line saying the sweep found nothing that kills the decision, and "Fix before
      you commit" carries only the verifications still worth doing.
      1. <how it happens and what it costs, one line> (risk 1)
      2. <...> (risk 2)
      
      **The thing nobody is questioning**
      <one line>
      
      **Fix before you commit**
      - <the action> (risk 2)
      - <the thing to verify, and how> (risk 4)
      
      **The decision, rewritten** - only where those fixes change what is being
      decided: three to five sentences restating it with them applied. Otherwise:
      "The decision stands, with the fixes above."
      
      ## The detail
      
      ### What usually kills decisions like this
      The class this decision belongs to, and the handful of ways decisions of that
      class usually fail - three to six lines - then which of those are live here
      and which are not. Never deliver this by folding it into the risk cards. It is
      what shows the reader whether you attacked the ordinary failures or only the
      ones that happened to occur to you.
      
      ### What was checked
      One line per area from the sweep: the risk it produced, or why nothing
      credible came out of it. All nine appear, including the empty ones.
      
      ### The risks
      The cards, worst expected damage first.
      
      ### Also worth knowing
      The thing nobody is questioning, expanded. The cheapest way to find out you
      are wrong. A flaw no fix can cure. Who benefits if this fails. The standouts.
      When to pull the plug. Each one exactly as described above, and a skipped one
      is still its one line saying it was skipped and why.
      
      ### What holds
      Two to four lines: the parts of the decision that survived the attack. These
      are the parts not to churn while fixing the rest.
      ```
      
      **Quick look** (when the dispatch says `MODE: Quick`): part one in full, what
      usually kills decisions like this, what was checked, the top three to five
      cards, the thing nobody is questioning, and
      the cheapest way to find out you are wrong. Everything else appears as its one
      line saying it was skipped.
      
      ## Rules
      
      1. Every claim is anchored or labelled: it quotes the brief, names outside
         evidence, or says it is an assumption. There is no fourth kind.
      2. How likely, how bad, and would you see it coming are three separate
         answers. Never collapse them into one word like "risky".
      3. A section you skipped is one line saying why. Silence and oversight look
         identical to the reader, so never leave a gap unmarked.
      4. Follow the interests. Wherever the plan needs somebody to behave a certain
         way, check whether their own interests agree. Where they do not, that is a
         risk with a mechanism, not a footnote about culture.
      5. Look for what disproves the decision, not what confirms it. You are not
         assembling a case for the plan and you are not assembling one against it.
      6. Good news beyond the verdict itself goes in "What holds", at the end.
         Never in the opening, and never as a cushion around a finding.
      7. "This decision is sound" is a legitimate verdict after the attack. It is
         never a substitute for running one.
      8. Too cheap for the treatment it arrived with is a finding, not a mode to
         obey. Where the brief itself shows the decision is reversible and cheap to
         be wrong about, say that in one line, answer at the Quick look depth even
         when the dispatch says MODE: Full, and name the one thing that would make
         it worth a full attack. A full deck of cards on a two-day experiment
         teaches the reader to ignore the next report.
      
    • why-these-rules.md 1.3 KB
      # Why these rules
      
      The short rules in `SKILL.md` say what to do. This file says why, so the skill
      itself stays short. Nothing here changes behaviour.
      
      **A date to be judged by.** "It failed" is meaningless without "by when".
      
      **The brief is where this is won or lost.** The fresh agent sees nothing else.
      The usual failure is a brief missing the one constraint that made the decision
      sensible, which produces a confident report attacking a plan nobody proposed.
      
      **Fewer than three risks.** It reads as a lazy analysis and is often the right
      answer. The record of what was checked is what tells those apart.
      
      **Reversible and cheap does not need this.** Running a full attack on a two day
      experiment teaches the user to ignore the next one.
      
      **"Try it small first" is not a soft no.** It means the unknowns are testable
      and worth testing before the money moves.
      
      **The user may go ahead against all of it.** That is their decision and they now
      have the early warnings.
      
      **Earning the ask.** Nobody there to answer, or an answer that would not move
      the verdict, means skipping straight to writing the guess down as a guess.
      
      **Offering to act on it.** That offer is the point of the whole exercise: a
      premortem nobody acts on was entertainment. "Go ahead" means the plan was
      attacked and held.
      
  • scripts
    • prompt.sh 472 B
      #!/usr/bin/env bash
      # Prints the analysis prompt so SKILL.md can inline it at load time. A script
      # rather than a bare cat: Claude Code refuses a cat of any file outside the
      # session's working directory in a skill's inline commands (measured on
      # 2.1.222), and that refusal aborts the whole skill. A script in the skill's
      # own folder is the documented way to bring a file in, and it is not refused.
      cat "$(cd "$(dirname "$0")/.." && pwd)/references/premortem-prompt.md"
      
  • SKILL.md 7.2 KB
    ---
    name: os-what-could-go-wrong
    description: >-
      ALWAYS invoke this skill before anything hard to undo gets agreed to - a
      contract, a purchase, a migration, a launch, a price change, a
      reorganisation - and whenever the user asks "what could go wrong", "what are
      we missing", "poke holes in this", or for a premortem or a red team, in any
      language. Assumes the decision already failed and works backwards to find
      out why, in a fresh agent that had no hand in making it. Sweeps nine areas
      and shows what each produced, including the empty ones. Ends on one verdict:
      go ahead, go but fix these first, try it small first, think again, do not do
      this.
    allowed-tools:
      - "Read(~/.claude/open-steps/**)"
      - "Bash(bash ${CLAUDE_SKILL_DIR}/scripts/prompt.sh)"
    ---
    
    # os-what-could-go-wrong
    
    Assume it already failed, then work backwards to find out why - while there is
    still time to change it. The attack runs in a fresh agent that had no part in
    the decision, because an agent that helped shape one reviews it far too
    gently: it defends its own reasoning, and it misses the thing that kills the
    plan out of politeness.
    
    `os-ask-simple` screens a choice before the user picks one. This one runs
    after the choice is made and before it can no longer be taken back.
    
    ## Language
    
    Write in the language the user speaks in this session, detected from the
    conversation. Names, figures and identifiers stay as they are. The agent you
    dispatch cannot see this conversation and does not inherit the writing style,
    so the language has to travel with the handover - see step 2.
    
    ## Step 1 - write down what is actually being decided
    
    Find the decision first. Given as text or a document, that is it. Asked at the
    end of a discussion, it is the decision the discussion arrived at, and say
    which one you took it to be. If neither, ask which decision to attack.
    
    Then fill every line. This is the only thing the fresh agent will ever see.
    
    ```
    DECISION BRIEF
    What will be done: three to seven sentences
    What it is for: the problem it solves, and what success looks like, measured
      wherever it can be
    The main moves: the money, the people, the systems, the dates
    What is fixed: the constraints, plus the surrounding facts that matter
    Who it lands on: who and what is affected if this goes wrong
    What cannot be undone: which parts are one-way
    When we would know: the date success or failure actually gets judged
    What we know: facts from documents, data, the repository, past incidents,
      each with where it came from
    What we are assuming: every gap nobody could close, written as an assumption
    ```
    
    Close the gaps in this order, and stop as soon as a gap is closed.
    
    1. **Look it up yourself.** Documents, data, the repository, the last report
       in `~/.claude/open-steps/reports/`, what went wrong last time.
    2. **Ask, but earn the ask.** Only a gap where guessing wrong would change the
       verdict, one question at a time through `os-ask-simple`, three at most.
       Nobody there to answer, or an answer that would not move the verdict -> 3.
    3. **Write the guess down as a guess.** Put the assumed value in its line,
       mark it `(assumed)`, and repeat it under "What we are assuming".
    
    The brief states facts and open questions. It never makes the case for the
    decision: a brief that argues gets a report that agrees.
    
    ## Step 2 - hand it to an agent that had no part in it
    
    Pick the depth, say which in one line, and carry on; the user can change it.
    
    | Depth | When |
    |---|---|
    | **Full** | Hard to undo, or being wrong costs money, trust or data beyond one team |
    | **Quick** | Reversible, and cheap to be wrong about, whoever it touches |
    
    When in doubt, Full. The cost of a full look is a few minutes; the cost of a
    quick look at a one-way door is the door.
    
    Dispatch one general-purpose agent. Send it four things and nothing else: the
    analysis prompt copied exactly from the bottom of this skill, `MODE: Full` or
    `MODE: Quick`, `LANGUAGE: <the language above>`, and the brief. Never send an
    instruction to run this skill: the fresh session would load it and start over.
    
    One agent, not several. Where no fresh agent can be started, say so in the
    first line of the report, name the tool, and never use the word independent.
    
    ## Step 3 - give it to the user straight
    
    The user never sees what the agent returned: a tool result is visible only to
    you. So your final message is that report, copied whole, first line to last.
    One line goes before it: the depth, and whether a fresh agent ran. Anything of
    your own comes after it, never instead of it: no summary in its place, no
    "details above", no reassurance the analysis did not earn, no dropped card
    because the user seemed committed. Bad news that arrives late is worth nothing.
    
    Then offer to turn the "Fix before you commit" list into real things: edits to
    the plan, tickets, an owner and a date per item, a reminder for each early
    warning. "Go ahead" is delivered just as plainly.
    
    ## Hard rules
    
    1. **The agent that helped decide never attacks the decision.** Dispatch a
       fresh one every time, even when you already hold the whole thing in
       context. Skipping this does not save a step, it changes the answer.
    2. **No quota of risks.** Publish what has a real chain behind it and nothing
       else. Two well-evidenced risks beat six padded ones, and "only two
       survived" is a finding worth saying out loud.
    3. **The verdict is decided last and printed first.** Never make the reader
       assemble it from the risks.
    4. **Every area of the sweep is accounted for**, including the ones that
       produced nothing. An area nobody mentions and an area nobody checked look
       the same to the reader.
    5. **A skipped section keeps its one line saying why.**
    6. **A lease and a database migration get the same treatment.** This is not a
       technical review; the nine areas apply to both, and the money and people
       ones are where technical decisions usually actually fail.
    7. **"The plan is sound" is a legitimate answer** once the attack has run. It
       is never a substitute for running one.
    8. **The final message is the report itself.** The user cannot see what the
       agent returned; a summary of it, however good, is not it.
    
    ## Known gotchas
    
    - **No date to be judged by means no premortem.** Pick a date that fits the
      decision, and say you picked it.
    - **The brief is where this is won or lost.** The fresh agent sees nothing else.
    - **Fewer than three risks is often the right answer.** Keep the record of what
      was checked.
    - **Something reversible and cheap does not need this.** Quick look, or say so.
    - **"Try it small first" is not a soft no.** Test unknowns before money moves.
    - **The user may go ahead against all of it.** Note it once, set the tripwires
      up if they want them, and do not re-argue the report.
    
    The reasoning behind these is in [`references/why-these-rules.md`](references/why-these-rules.md).
    
    ## The analysis prompt, verbatim
    
    The block below is the whole of
    [`references/premortem-prompt.md`](references/premortem-prompt.md), inlined
    when this skill loads through `scripts/prompt.sh`, so no file has to be read
    at dispatch time. If the block shows a literal command instead of the prompt,
    this harness does not run inline commands: open that file next to this one
    and use its full text.
    
    !`bash ${CLAUDE_SKILL_DIR}/scripts/prompt.sh`
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related