{"slug":"gentle-ai-bench-2","title":"gentle-ai-bench","summary":"Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis. Author and verify gentle-ai bench journeys; go test ./bench never proves driven execution.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-08T21:17:37.22641Z","repo":{"url":"https://github.com/Gentleman-Programming/gentle-ai","stars":7321,"forks":800,"license":"MIT","updatedAt":"2026-09-26T22:22:19Z"},"bodyHtml":"<hr>\n<h2>name: gentle-ai-bench\ndescription: \"Trigger: bench, journey, journeys, driven mode, gentle-ai-bench, journey corpus, j-numbers, bench axis. Author and verify gentle-ai bench journeys; go test ./bench never proves driven execution.\"\nlicense: Apache-2.0\nmetadata:\nauthor: \"Gentleman-Programming\"\nversion: \"1.0\"</h2>\n<h2>Activation Contract</h2>\n<p>Load when touching <code>bench/</code> in gentle-ai, adding or changing a journey, changing a product semantic a journey might pin, or diagnosing a bench failure in CI's Unit Tests job.</p>\n<h2>Hard Rules</h2>\n<ul>\n<li><code>go test ./bench</code> validates corpus declarations only. It does NOT execute journeys. The only driven proof is building the harness and the product binary and running the harness against it; a green <code>go test ./bench</code> claims nothing about execution.</li>\n<li>Reproduce CI, do not guess invocations: read the Unit Tests step in <code>.github/workflows/ci.yml</code> and copy its exact build and <code>gentle-ai-bench run --binary ...</code> commands. Use <code>--only &lt;journey-id&gt;</code> to drive one journey.</li>\n<li>Journey IDs are unique across every <code>journeys_*.go</code> file. The collision guard fails loudly naming both files; pick an unused ID by reading the corpus, never reuse a retired one.</li>\n<li>Every journey declares <code>Review:</code> — <code>reviewOptedIn</code> (the runner enables receipt-driven development globally before the first step, uncounted, and fails the journey if the switch does not come on) or <code>reviewUntouched</code> (its subject IS the switch, or it has nothing to do with reviews). The declaration is mandatory; <code>validateCorpus</code> fails the run without it. Never let a journey inherit the product's default: reviews are opt-in, and a journey that assumed otherwise measures a review-refused flow while still reporting <code>completed</code>.</li>\n<li>Every <code>execute</code> transition must carry a runnable command; the dead-execute guard fails the run otherwise.</li>\n<li>When a ratified product semantic changes, grep the corpus for journeys pinning the OLD behavior before shipping. The corpus is a second test surface beyond unit tests; a journey asserting the defect keeps the defect green.</li>\n<li><code>dead_end</code> prints <code>n/a</code> unless the run actually measured one. Never fabricate a value to move the column.</li>\n<li>A <code>by_design</code> exemption costs a shape from the closed vocabulary plus a verified quote of the product's own next-action text. If the quote no longer tells the operator what to do, it is a defect wearing an exemption.</li>\n<li>Prefer a NEW <code>journeys_*.go</code> file when the shared ones are owned by open PRs; bump the core journey-count pin in the same change.</li>\n</ul>\n<h2>Execution Steps</h2>\n<ol>\n<li>Read the corpus area you touch and the CI invocation before writing.</li>\n<li>Author or adapt the journey; update its title, step names, and comment to say WHY the expectation holds (cite the issue or ratified decision).</li>\n<li>Run <code>go test ./...</code> in <code>bench/</code> for declarations, THEN the driven harness for execution; both results go in the PR body.</li>\n<li>On semantic changes, list the journeys you checked for stale pins.</li>\n</ol>\n<h2>Output Contract</h2>\n<p>PR evidence includes the driven-mode summary line (completed / unsupported / failed counts) from a locally built binary, not only <code>go test</code> output.</p>\n","files":[{"path":"SKILL.md","sizeBytes":3181,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-26T23:11:59.107376Z","sha256":"AF252EB2EA4E992B69EA4B8EB8643E273000FB0CAD2386A9FE64D987B023FA9C","sizeBytes":1675},"review":null,"source":{"repositoryUrl":"https://github.com/Gentleman-Programming/gentle-ai","path":"internal/assets/skills/gentle-ai-bench","license":"MIT","commit":"a9e36e9b8a4d7885244466cd9ea6cc3ad330a69b","subtreeSha":"5D9CEA70283FEE71B5FE10A298D5E252DEFB3504CBDB789719AD269493B8B967","lastSyncedAt":"2026-09-26T23:11:39.069062Z"},"reviewedAt":"2026-09-26T23:13:27.86093Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/Gentleman-Programming/gentle-ai/tree/main/internal/assets/skills/gentle-ai-bench"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install gentleman-programming-gentle-ai@llmmart"},{"target":"git","command":"git clone https://github.com/Gentleman-Programming/gentle-ai.git"}]}