authors-voice
Author's Voice — constructed-voice skill. Anchors writing to a training-data author blend, progressively layers NEVER rules, presentation fingerprints, sentence stats, coined terms, and curated examples from a growing local corpus. Pure markdown, opus sub-agent (the minion) write
Install
npx skills add https://github.com/travsteward/openwriter/tree/main/plugins/authors-voice/skill
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install travsteward-openwriter@llmmart
git clone https://github.com/travsteward/openwriter.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole travsteward/openwriter collection as a plugin from our marketplace. Git is the plain clone.
README
authors-voice
AI writing that sounds like you, not AI.
Most attempts to make AI sound like you start the same way. Train it. Fine-tune it. Feed it your samples and tell it to imitate. The model doesn't actually learn you from any of this. It pattern-matches at the lexical layer, lifting your common words and sentence shapes without ever building a deep representation of how you think. Cold-start imitation tops out shallow.
Flip the direction. The model already carries deep internal representations of widely-published authors it was trained on at scale, voices it can channel with real fidelity because it saw thousands of pages of each. The move is to identify which of those authors a user statistically resembles, assign proportional weights to the closest matches, and instruct the model to write as that weighted blend. Your voice gets reconstructed as a coordinate inside the model's existing author space, anchored to authors it has already mastered.
That changes the problem. The model isn't being asked to learn anything new about you. It's being asked to mix voices it knows cold, in proportions that triangulate your position among them. The anchor does the heavy lifting before a single sample of yours enters the prompt. The blend is the voice.
The result is AI writing that sounds like you. Not AI imitating you.
On top of the anchor, four layers sharpen the output. A list of AI words and constructions the model must never use, because the moment it stops channeling the anchor it reverts to its trained register and reaches for the same fifty tells. Presentation choices you make consistently, like whether you capitalize after a colon or use the Oxford comma, small mechanical preferences that read as authentic. A sentence-length and punctuation rhythm pulled from your own writing, so the cadence matches even when the diction is on loan. A growing folder of your samples that the skill mines as the negative rules and rhythm get re-derived.
Each sample you add updates the NEVER rules and the sentence rhythm against your latest corpus. The anchor and presentation fingerprints don't auto-refresh. Regenerate those when you've added enough new writing to shift the matches, or when you want a fresh pass. The profile gets sharper the more you write and the more often you ask for a refresh.
Roughly 80% of the way to your real voice. A hard jump above what stock prompting and fine-tuning produce.
Install
The skill is agent-agnostic. Pure markdown, no language runtime. Any LLM-based agent that can read SKILL.md and follow instructions can use it.
Claude Code
claude install github:travsteward/authors-voice
Clones to ~/.claude/skills/authors-voice/ and registers the skill with Claude Code.
Vercel skills CLI (Claude Code, Codex, Cursor, and other agents)
npx skills add travsteward/authors-voice
Manual (any agent)
git clone https://github.com/travsteward/authors-voice
Then drop the cloned folder wherever your agent loads skills from. The SKILL.md at the root has the trigger phrases and routing logic the agent reads.
Quick Start
Two paths. Pick one.
Path A: Web tool first (fastest first-run)
- Visit openwriter.io/voice-match, paste 300 to 800 words of your writing, copy the result block.
- Tell your agent: "set up my voice match". Paste the block when prompted.
- Seed your corpus: paste 2 to 5 paragraphs that feel most like you. The agent saves them under
voice/corpus/. - Done.
Path B: Skill mode (no web round-trip)
- Tell your agent: "set up my voice match" and "I want to skip the web tool."
- Paste 2 to 5 paragraphs of your writing. The agent saves them under
voice/corpus/. - The agent runs the Anchor Protocol over your corpus and writes
voice/anchor.mddirectly. - Done.
The skill is self-routing. You don't memorize subcommands. Just tell the agent what you want:
- "voice status" → reports your current tier and word count
- "add this essay to my voice profile" → appends, re-analyzes
- "write me a tweet about X" → uses your voice automatically
How It Works
Your voice profile lives in voice/ as a handful of .md files the agent reads at write time:
| File | Source | Purpose |
|---|---|---|
anchor.md |
One-time match from openwriter.io/voice-match (or skill-mode). Refresh on demand. | 3 to 5 training-data authors with weights |
stats.md |
Auto-regenerated from corpus on every analysis run | Sentence distribution + punctuation density |
never-rules.md |
Auto-regenerated from corpus on every analysis run. Manual additions preserved. | AI words and phrases the model must never use |
fingerprints.md |
Agent extracts from corpus during analysis runs. Manual overrides preserved. | Presentation choices (Oxford comma, capitalization after colon, contraction frequency) |
coined-terms.md |
You curate | Your repeated coinages |
examples.md |
You curate | Reference paragraphs in your voice |
status.md |
Auto-regenerated on every analysis run | Current tier and what's locked next |
Plus voice/corpus/. Your raw samples accumulating over time. None of voice/* is committed. It's all local to your disk.
What updates reliably on every sample add: NEVER rules, sentence stats, status. What gets re-derived in protocol but agents sometimes skip: fingerprints (ask for a rebuild if you want certainty). What needs an explicit ask: anchor weights ("regenerate my anchor"). The corpus folder is yours to grow.
Progressive Tiers
The more samples you add, the more confident the analysis.
| Words | Tier | Active |
|---|---|---|
| under 300 | 0 | seed corpus first |
| 300 to 1k | 1 | anchor and basic stats |
| 1k to 5k | 2 | preliminary NEVER rules and top fingerprints unlock |
| 5k to 20k | 3 | full NEVER coverage and all fingerprints unlock |
| 20k and up | 4 | high-confidence profile |
Privacy
- Your voice data lives entirely on your disk.
.gitignoreexcludes everything invoice/from the public repo. - The skill never uploads your corpus anywhere.
- The only thing that leaves your machine is the initial 300 to 800 word paste into openwriter.io/voice-match for the anchor matching step. That's cached 24h by hash and never trained on.
Requirements
- A Claude Code or compatible agent that supports skills (no Node.js dependency)
- An initial visit to openwriter.io/voice-match for the anchor (free, no signup), or use skill-mode to build it locally
Beyond the skill
Pairs naturally with OpenWriter, the free AI writing surface. Same anchor system also powers the paid Author's Voice plugin (inline voice edits inside OpenWriter) and the paid API (programmatic voice-matched output for workflows and apps). See authors-voice.com when you outgrow the skill alone.
License
MIT. See LICENSE.
Replaces
This skill replaces the older writers-voice skill and the legacy voice-apply, voice-generate, voice-setup, voice-upload, voice-manage, and voice-automate skills. They are now one trigger: /authors-voice.
History
The local-skill half of /authors-voice started life as the standalone writers-voice skill. Its full development history (every iteration of the anchor protocol, NEVER rules, fingerprints, and tier logic) lives in the archived travsteward/writers-voice repo's git log. The repo is private now, but the commit log is preserved as the record of how the constructed-voice architecture evolved before it was unified here.
Credits
Built on the negative-first voice profiling architecture from Author's Voice. Pairs with OpenWriter, the writing surface for AI agents.
Skill manifest
Author's Voice
This skill is the free, any-agent manifestation of the larger Author's Voice ecosystem — the same anchor + NEVER-rules + anti-AI engine that powers the paid API, the OpenWriter plugin, and the dashboard. Free here, productized there; one voice DNA across all surfaces.
FIRM RULES
1. Editor NEVER writes prose. Every writing task fires a minion.
The editor scopes briefs, cuts, reorders, and patches NEVER violations. Editor does NOT write new prose — every word ENTERING the document is minion-written. Applies to: initial drafts, revisions, bridges, closers, openers, transitions, single-line aphorisms, one-paragraph corrections, gap-fillers, idea-extensions — any new prose, full stop.
Two carve-outs: (1) violation patch at Apply step 6 — 1-3 sentence local fix to NEVER violation or brief-error in otherwise-acceptable minion output; constructive rephrase preferred. (2) co-write mode — editor writes directly when ALL THREE hold: continuous real-time collaboration (user steering each move), explicit per-piece authorization for THIS piece, small scope (sentence to short paragraph; max one). Blanket authorizations ("you handle it") do NOT trigger co-write — those are delegations and go to the minion.
Editor territory (no minion): cuts, reorders, accept/reject decorations, resolve agent marks, version restores. Cuts that leave holes needing new connective tissue — the connector is minion work.
Rationale: editor context is polluted by the live conversation; voice anchors lose to active context; minion in clean context with voice files loaded outperforms editor-with-intent.
If the editor catches itself drafting prose mid-conversation, STOP and spawn a minion — even for one sentence.
2. After anchor critique, discuss before revising.
Anchor critique returns scores + convergent diagnostic. First move after aggregation: surface to user, discuss what to act on / disagree with / defer, align scope. Then spawn revision minions. Critics are advisory; user owns the prose.
Anchor critique result protocol:
- Aggregate panel scores + convergent diagnostic.
- Surface to user: scores, themes raised by 3+ critics, proposed CUTS (subtraction) and REWRITES / ADDITIONS (new prose).
- Discuss. User decides act / disagree / defer. No revision begins until this happens.
- Editor executes agreed CUTS directly (Rule 1 carve-out).
- Editor scopes briefs for agreed REWRITES / ADDITIONS, spawns revision minions per brief.
- Patch micro-violations on returned prose.
- Re-fire panel only if user wants another pass.
3. Revisions must be tighter than the original. If a revision is additive, the editor wrote it.
Critique-driven revision produces a smaller word count, not larger. If post-revision is longer than pre-revision, the editor was inventing rather than executing the diagnostic. Revision minion brief specifies a WORD-COUNT TARGET (often "this paragraph in 60% of the original"). Minion compresses. Editor verifies the count dropped.
Architecture
Skeleton prompt template (prompts/skeleton.md) assembled from per-user voice/*.md files at write time. Editor loads skeleton, substitutes {INCLUDE: ...} markers, fills {TASK}, spawns a fresh opus sub-agent (the minion) with the assembled prompt (Claude Code) or dispatches via task({ subagent_type: "general", prompt: <assembled skeleton> }) (OpenCode). Minion has no session pollution, returns prose, dies. Editor integrates.
writers-voice/
├── SKILL.md (this file — router + firm rules + Apply Protocol)
├── docs/ (on-demand: setup, analysis, apply-deep, anchor-iteration, context-hygiene, tiers)
├── catalog/ (read-only reference: ai-tells, fingerprints, hurdle, anchor-prompt, author-hints, post-write-audit)
├── prompts/skeleton.md (template + injection points)
└── voice/ (user-specific — anchor blend, NEVER rules, stats, fingerprints, coined-terms, examples, corpus/)
Lean / rich split: voice/*-analysis.md is human-facing only — never injected into the minion prompt.
Routing
| User intent | Action |
|---|---|
| "set up my voice" / "/writers-voice" / first run | Setup Flow — docs/setup.md |
| "add this essay to my voice profile" / "save this writing" | Append to voice/corpus/, run Analysis Protocol — docs/analysis.md |
| "voice status" / "what's locked" / "tier" | Read voice/status.md; tier reference at docs/tiers.md |
| (User asks the agent to write anything) | Run Apply Protocol (below) after Context Hygiene check — docs/context-hygiene.md |
| "show me my anchor" / "show me my fingerprints" | Read the relevant voice/*.md and report |
| "regenerate my profile" / "re-analyze my corpus" | Analysis Protocol — docs/analysis.md |
| "make a book / business / [context] voice anchor" | Anchor Protocol for that register — docs/setup.md |
| "split my anchor by register" | Multi-Register Split — docs/setup.md |
| Book-scale project (multi-chapter book) | Load /book-writer skill — that's the orchestration layer (chapter architecture, beats methodology, workspace management, book mode, long-form orchestration). This skill provides the Apply Protocol that /book-writer delegates to for every prose pass. |
| "polish this" / "iterate to 90" / final-polish ready prose | Anchor Iteration — docs/anchor-iteration.md |
| "use the API" / "call author's voice from a workflow" / plugin / programmatic | API path — docs/api/protocol.md (rewrite, generate, MCP tools, setup, troubleshooting) |
Apply Protocol (Apply Minion — generative writing from commitments)
Four minion types (full taxonomy: docs/apply-protocol-deep.md): Apply (generative, no source prose) · Rewrite (Apply + context awareness, updated commitments) · Blinder Audit (critic — paragraph-level substance duplication only) · Anchor Iteration (polish — channels voice anchors as panel, iterates critique → rewrite → re-score until 90/100).
Pick the wrong minion → weak output. Substance problem → Rewrite, not Apply. Rough draft → not Anchor Iteration.
When the user asks for a voice-matched write:
Context Hygiene check. Reset if polluted —
docs/context-hygiene.md.Pick the anchor. List
voice/anchor*.md. If onlyvoice/anchor.md, use it. If context-specific, infer from request or ask. Fallback:voice/anchor.md.Assemble the minion prompt. Read
prompts/skeleton.md. For each{INCLUDE: <path>}, substitute file contents. If a referenced file is missing (e.g., user hasn't curated examples), drop the entire{INCLUDE: ...}line AND its section header. Swapvoice/anchor.md→voice/anchor-<context>.mdif a context-specific anchor applies.Fill
{TASK}.Required: COMMITMENTS — what must be true of the output (concepts, claims, sequence, register, avoidances, length).
DEFAULT MODE: pure generation (regenerate). SEMANTIC commitments only — what claims must land, not how to phrase them. No structural beats. No paragraph patterns. No device prescription. No sentence-rhythm prescription. The shape of any prescription becomes a ceiling on what the model produces. Let the minion bring its own moves. Always.
Never write meta-references into commitments
Anything the prose reader cannot see — chapter labels, beat numbers, "as discussed earlier", "the previous beat established" — does NOT belong in commitments. The minion reproduces them literally and breaks the fourth wall. Commitments describe what must be COMMUNICATED, never how the editor is thinking about structural position. If continuity from a prior section matters, capture it as the SUBSTANTIVE thread the new section must pick up (the content, not the structural pointer).
COMMITMENTS function as quasi-verbatim instructions
When the editor writes a commitment with literal phrasing in parentheses ("Define sleep debt (lost sleep compounds like unpaid interest)"), the model treats the parenthetical as exact phrasing to reproduce — every section gets that line. For multi-section work where phrasing should vary, write commitments abstractly ("Define sleep debt") and let the model phrase. Use literal commitments only when a specific phrasing MUST land.
Read prior integrated sections first (multi-section work)
Read prior integrated sections before writing this one's brief. Cadence shapes already used, substantive threads to pick up, and heavy-use coined terms are only visible by reading what's on the page. Set the new section's commitments and cadence prescription against that context.
Preservation scope (load-bearing call on a gradient)
Mode Source prose in TASK? Commitments shape Use when Full preservation (rewrite mode) YES "preserve every load-bearing claim; refine voice while preserving structure and phrasing" source is already strong; author has specific phrasing that must land; polishing voice-applied work Pure generation with selective lifts NO OMIT source. Identify 1-3 specific moves worth keeping. Lift them pre-emptively (MUST-APPEAR-VERBATIM in brief) OR post-edit (patch in after minion returns) source is mostly weak with 1-3 lines worth keeping. Pre-emptive when known in advance; post-edit when the strong move is only obvious after seeing the minion's output Pure generation (regenerate) NO SEMANTIC commitments only — what claims must land, not how to phrase them source is "good but not great"; want dramatic improvement; existing shape would constrain output; high-stakes piece. The shape of the source becomes a ceiling on what the model produces Cadence prescription (optional, recommended for high-stakes writing)
Lifts voice fidelity ~0.5 points over baseline. Explicit rhythm scaffolding doesn't constrain content; it frees capacity by removing "what shape should this take?" overhead.
Example: "Para 1: open with 3 short declaratives, stack medium with concrete examples, close with one analytical long. Para 2: alternate short claim with longer explanatory, end with sharp short. Para 3: build with longer analytical, end with single-line aphoristic close."
Vary cadence prescriptions across sections. Same prescription per section produces document-scale rhythm repetition (every section opens with 3 shorts, closes with aphorism) — invisible at section scale, mechanical at document scale.
Edge-case guidance (Rewrite Minion brief template, Blinder Audit brief shape, multi-section context-loading layers, writing minion taxonomy):
docs/apply-protocol-deep.md.Spawn the minion. Claude Code:
model: "opus",subagent_type: "general-purpose". OpenCode:subagent_type: "general"with no model parameter (subagent inherits parent model; encourage using the session's strongest model). Both:prompt: <assembled skeleton>.Patch NEVER violations + brief-error meta-references. Smallest local span. Constructive rephrase preferred (contrastive negation → direct statement; banned word → plain equivalent; meta-reference → substantive thread it pointed at). Don't regenerate; minion voice IS the result. Detail:
docs/apply-protocol-deep.md.Post-write audit. Read
catalog/post-write-audit.mdand apply distribution-level checks (opener repetition, sentence-initial "The", function-word over-use, sentence-length variance, lexical watch list). For each failing check, surgically rewrite the smallest local span — 5-10 light substitutions across a typical draft; heavier rewrites mean misuse. Load-bearing prose wins ties.Integrate via openwriter.
write_to_padfor edits,populate_documentfor new docs.Cross-section coherence review (multi-section only). (a) Editor self-review — cut test: could you delete this paragraph and lose only redundancy? (b) Mandatory Blinder Audit Minion — fresh-context critic, paragraph-level substance duplication only. Most well-written beats produce zero findings. Brief template + exclusions:
docs/apply-protocol-deep.md.Polish (optional, two patterns). (a) Parallel pick-best — N (3-6) Apply minions in parallel, same brief; editor picks best whole, mixes variants, or hands all to user. (b) Anchor Iteration —
docs/anchor-iteration.md. Polish-class only; not for rough drafts./anti-aipass. MANDATORY after Anchor Iteration (which runs no-context and introduces AI tells). OPTIONAL otherwise. Global surface fingerprints (em-dashes, semicolons, contrastive negation, banned diction, register monotony) vsvoice/never-rules.md+voice/fingerprints.md. Complements step 7.
Use opus (Claude Code). Sonnet leaks 3+ NEVER violations where opus leaks 0-1. Haiku loses voice. OpenCode: subagents inherit the parent model — use the strongest model available in the session for prose generation. Send full editing scope. If 6 of 8 paragraphs need fixes, send all 8 for flow continuity. One minion per natural editing unit — beat, section, blog post, tweet thread.
Tiers + Companion Skills
Voice profile tiers (Empty / Anchor / Preliminary / Full Coverage / AV-Grade) gate features by corpus word count. Table: docs/tiers.md.
Companions: /anti-ai (final fingerprint scrub) · /voice-presets (generic frames, no profile — if installed) · Author's Voice plugin (paid — full RAG, inline edits, deterministic extraction).
Files (openwriter)
-
catalog
-
ai-tells.md 6.3 KB
# AI Tells Catalog > The 60+ words, transitions, phrases, and patterns that signal "this was AI-generated." > Use this list when analyzing a user's corpus to decide which NEVER rules to emit. ## How to use this catalog For each item below, count how many times it appears in the user's corpus (case-insensitive, whole-word match for words; literal substring for phrases). Then apply the [authenticity hurdle](./hurdle.md): - **If count and rate both clear the hurdle** → the user genuinely uses this. Preserve it. Do NOT emit a NEVER rule. - **If it appears but below the hurdle** → corpus contamination (likely from AI-assisted earlier drafts). Emit a NEVER rule. - **If it doesn't appear at all** → emit a NEVER rule (training data will push it back in unless explicitly banned). The default state for every item here is "forbidden." It only graduates to "preserved" if the user uses it at signature frequency. ## AI Words (whole-word, case-insensitive) Category: `word`. Authenticity hurdle: count ≥ 2 AND rate ≥ 0.15 per 1000 words. - `delve` - `delves` - `pivotal` - `tapestry` - `realm` - `beacon` - `harness` - `illuminate` - `underscore` / `underscores` - `bolster` / `bolsters` - `facilitate` / `facilitates` - `multifaceted` - `foster` / `fosters` - `crucial` - `vital` - `comprehensive` - `robust` - `leverage` / `leverages` - `seamless` / `seamlessly` - `navigate` / `navigates` - `embark` / `embarks` - `nestled` - `unwavering` - `indelible` - `meticulous` / `meticulously` - `transformative` - `revolutionize` / `revolutionise` ## AI Transitions (sentence-initial unless noted) Category: `transition`. Hurdle: count ≥ 2 AND rate ≥ 0.15 per 1000 words. Match these only when they start a sentence (after `.`, `!`, `?`, or paragraph start). The same word mid-sentence is fine — it's the sentence-leading transitional move that signals AI. - `Moreover,` (sentence start) - `Furthermore,` (sentence start) - `Notably,` (sentence start) - `Additionally,` (sentence start) - `Consequently,` (sentence start) - `Ultimately,` (sentence start) - `Therefore,` (sentence start) - `However,` (sentence start) — note: sentence-leading "However" is the tell, not all uses - `Indeed,` (sentence start) - `Thus,` (sentence start) - `Subsequently,` (sentence start) - `Nonetheless,` (sentence start) - `Firstly,` (sentence start) - `In conclusion,` (anywhere) ## AI Phrases (literal substring, case-insensitive) Category: `phrase`. Hurdle: count ≥ 2 AND rate ≥ 0.05 per 1000 words. - `provides valuable insights` / `provide valuable insights` - `gains valuable insights` / `gain valuable insights` - `valuable insights` - `a rich tapestry` - `plays a crucial role` - `in today's digital age` / `in today's digital era` - `in the fast-paced world` / `in today's fast-paced world` - `at its core` - `that being said` - `to put it simply` - `it's worth noting` - `it's important to note` - `delve into` - `dive into` - `navigate the complexities` / `navigating the complexities` - `the intricate relationship` - `a nuanced understanding` - `opens new avenues` / `opens up new avenues` - `leaves an indelible mark` / `leave an indelible mark` - `stands as a testament` / `stand as a testament` - `paves the way` / `pave the way` - `a stark reminder` - `a beacon of` ## AI Punctuation Category: `punctuation`. Hurdle: count ≥ 15 AND rate ≥ 1.5 per 1000 words. (Much higher than diction because punctuation is denser.) - **Em-dashes (`—`)** — the single highest punctuation tell. Most human writers use them at <1 per 1000 words. AI defaults to 4-8 per 1000. ## AI Structural Patterns (no regex — pattern recognition) These can't be matched with simple string search. Read the corpus and judge whether each pattern appears. ### Always banned (unconditional) - **Contrastive negation** — `"It's not X, it's Y"`, `"Not just X but Y"`, `"Rather than X, Y"`, `"Less X, more Y"`. This is the single most identifiable AI-rhetoric move. Emit the NEVER rule regardless of whether the user uses it. Reason: it's so detectable that even occasional use blows the cover. ### Vote-to-ban (emit NEVER unless user clearly uses it as signature move) For each of these, decide: does the user use this as a deliberate stylistic move (≥2 clear instances in corpus)? If yes, preserve. If no, emit a NEVER rule. - **Tricolon abstractions** — rule-of-three lists of abstract nouns as a flourish (e.g., `"innovation, collaboration, and excellence"`, `"strength, courage, and resilience"`). Distinct from regular three-item lists about concrete things. - **Soft repetition** — restating the same point with only minor wording changes across consecutive sentences. (Anchoring the same idea with different specifics is fine; rewording the same idea isn't.) - **Mismatched enthusiasm** — encouragement, excitement, or affirmation beyond what the content warrants. The "wow that's so insightful!" energy that doesn't fit the actual subject. - **Representational hedging** — telling the reader what things `"represent"`, `"underscore"`, `"reflect"`, or `"stand as a testament to"`. The interpretive overlay AI loves to add when concrete description would do. ## Adjacent never-rules (deduced from punctuation tally, not from this catalog) After the punctuation density count, if the rate is exactly 0 ("never" category) for any of these, emit a NEVER rule for it — UNLESS it's already covered by an AI-tell label above. This handles punctuation the author simply doesn't use, regardless of whether it's an AI tell. - `en_dash` → "No en-dashes." - `semicolon` → "No semicolons." - `ellipsis` (`...` or `…`) → "No ellipses." - `bracket` (`[`) → "No square brackets." - `double_quote_curly` (`"` `"`) → "No curly double quotes." - `single_quote_curly` (`'` `'`) → "No curly single quotes / apostrophes." - `exclamation` (`!`) → "No exclamation marks." ## Output format When emitting a NEVER rule for items from this catalog, use this exact format: - Diction (words, transitions, phrases): `NEVER "<label>".` - Examples: `NEVER "delve".`, `NEVER "Moreover," (sentence start).`, `NEVER "a rich tapestry".` - Punctuation (catalog items): `NEVER "<label>".` - Example: `NEVER "em-dashes".` - Punctuation (deduced): `No <nicename>.` - Example: `No en-dashes.` - Structural: use the verbatim NEVER text from the pattern entry above. This format matches what the Author's Voice plugin emits, so the rules read consistently across both tools. -
anchor-prompt.md 11.9 KB
# Anchor Prompt > The matcher logic the agent follows to produce a voice anchor blend entirely in-agent — no network, no hosted service. ## Role You are a literary stylometry analyst. You identify which training-data authors a piece of writing mechanically resembles. You match on **prose mechanics only** — sentence structure, punctuation patterns, rhetorical moves, discourse rhythm, vocab register. You **never** match on content, themes, or subject matter. If the writer discusses biology, you do NOT match them to Dawkins because of the topic. You match them to whichever author's *sentence construction* most resembles theirs. If the writer discusses startups, you do NOT match them to Paul Graham because of the topic. You match on prose mechanics — how the sentences are shaped, how paragraphs transition, how punctuation lands. ## Hard rules 1. Score ONLY on prose mechanics. Forbidden: matching on themes, topics, subject matter, ideology, worldview, or what the writer is "about." 2. Output exactly 3-5 authors with weights summing to exactly 100. 3. Each author's `features_matched` MUST cite at least 2 specific PROSE features. Each feature must describe HOW sentences/paragraphs are constructed, NOT what they are about. 4. Prefer authors from the author hints list — the model has the deepest compressed style modes for them. You MAY go outside the list if a clearly better stylistic match exists, but only when justified by prose features. 5. If the sample is too short or too uniform to distinguish style confidently, return fewer authors and flag confidence as `low`. ## Inputs the agent loads before running this protocol 1. **The user's corpus** — every file in `voice/corpus/` (strip YAML frontmatter). **Keep each sample separate** — do not concatenate yet. You need per-sample analysis before the blend. 2. **The deterministic stats** — read `voice/stats.md` if it exists. The sentence-length distribution and punctuation density numbers are the anchors that prevent agent drift. If `stats.md` doesn't exist yet, run the Analysis Protocol first to generate it. 3. **The author hint list** — read `catalog/author-hints.md` for curated authors with heavy training-data representation, grouped by category. 4. **The conversational guard** — set aside everything you know about the user from the current conversation (their projects, their interests, the topic they've been discussing). Score the corpus as if you've never read it before, with no context. ## Per-sample register analysis (run BEFORE final scoring) This step prevents the single-large-sample bias that overweights features concentrated in one piece. The matcher used to concatenate all samples into one text and score volume-weighted — which meant the largest sample's signature features got baked into the blend as if they were corpus-wide patterns. They aren't. For each sample in `voice/corpus/`, record: 1. **Word count** — split on whitespace, count tokens. 2. **Address mode** — first-person (`I/me/my/we/our`), second-person (`you/your`), third-person (`he/she/they/the modern man`), or mixed. If mixed, note the dominant mode. 3. **Tone register** — instructional, analytical/expository, polemical, conversational, reflective, narrative. Pick the dominant register. 4. **Signature moves present** — list 1-3 prose mechanics that stand out in THIS sample (e.g., anaphoric `Your X.` stacking, definitional `X is Y` pivots, dismissive concession `whether X matters less than Y`). Build a sample table: | Sample | Words | Address | Register | Signature moves | |--------|-------|---------|----------|-----------------| | 001 | 200 | mixed (you + we) | instructional | … | | 002 | 100 | generic-you | analytical | … | | ... | | | | | Then compute: - **Volume share** per sample: `sample_word_count / total_corpus_words * 100`. Flag any sample with >25% volume share — its features will skew the blend unless deliberately balanced. - **Register clusters**: group samples by register. Are there 2+ clearly distinct registers (e.g., third-person expository AND second-person instructional)? If yes, this is a **multi-register corpus** — surface as a warning in the final output. ## Register-aware feature validation When you're about to cite a feature in `features_matched` for any author in the blend, run this check: 1. **In which samples does this feature appear?** List them. 2. **What share of samples (by count) contains it?** 3. **What share of corpus volume contains it?** Then apply the rules: - **Feature appears in ≥40% of samples by count** → corpus-wide signature, valid to cite without caveat. - **Feature appears in <40% of samples by count BUT ≥40% of volume** → concentrated in fewer-but-larger samples. **Cite with caveat**: append `(concentrated in samples N, M — not corpus-wide)`. - **Feature appears in <33% of samples AND <33% of volume** → not a real corpus signature. Do NOT cite. If this was your strongest evidence for an author, drop the author from the blend. This rule kills the "one big sample's signature shows up as the dominant author weight" bug. The Manson-anaphora pattern from a single 480-word piece doesn't get to drive a 38% weight — unless it's actually replicated across multiple samples. ## Scoring dimensions Score the user's prose against each of these dimensions. Use these as the basis for feature matching: 1. **Sentence length distribution** — what percentage short / medium / long / very-long? How does this compare to each candidate author's known distribution? Use the numbers from `voice/stats.md`. 2. **Punctuation profile** — em-dash density (telling signal), semicolon use, colon use (for reveals vs. lists), parenthetical frequency, question marks, exclamations. Use the numbers from `voice/stats.md`. 3. **Vocab register** — academic / casual / instructional / aphoristic / journalistic / literary / polemical. Diction level. Use of jargon, slang, formal vocabulary. 4. **Rhetorical moves** — definitional pivots ("X is Y"), anaphora ("Your X. Your Y. Your Z."), rule-of-three, dismissive concession ("X matters less than Y"), concept-coinage (naming a concept then referring to it as proper-noun), claim-then-evidence vs evidence-then-claim, direct address. 5. **Discourse patterns** — how paragraphs transition. Use of "But", "So", "However", "Moreover". Question pivots. Time markers vs logical connectors. 6. **Paragraph rhythm** — uniform vs varied length. Where short paragraphs land (impact moments? insights? transitions?). 7. **Person & address** — first-person dominance, second-person directness, third-person formality. 8. **Hedging vs assertion** — frequency of qualifiers ("perhaps", "may", "could"), certainty markers, presence/absence of softeners. ## Examples of valid `features_matched` entries - `period-heavy short sentences (avg 9 words/sentence matches their pattern)` - `concept-coinage with definitional pivots ("Sleep debt is X")` - `anaphora across consecutive sentences ("Your job. Your family. Your kids.")` - `low em-dash density (0.3/1000 words matches their 0.4)` - `rule-of-three concrete lists, not abstract flourishes` - `dismissive concession move ("whether X holds matters less than Y")` ## Examples of INVALID `features_matched` entries - `writes about biology` — content - `anti-establishment themes` — content - `interested in masculinity` — content - `discusses religion` — content - `concerned with personal development` — content All of these are forbidden because they describe what the writer is ABOUT, not how the prose is constructed. ## Self-criticism step Before finalizing the blend, re-read each `features_matched` entry. For each entry, ask: - Does this describe HOW the sentences/paragraphs are constructed? (good) - Or does it describe WHAT the text is about — topics, themes, subject matter, ideology, worldview? (forbidden) If ANY entry describes content rather than mechanics, replace it with a prose feature or remove that author from the blend. After your check, set the self-check flags: - `any_thematic_reasoning`: `true` if you still had thematic reasoning that you couldn't fully fix. Otherwise `false`. - `confidence`: - `high` if corpus total > 800 words AND prose features are clearly distinguishable - `medium` if corpus total 400-800 words OR features are mixed/ambiguous - `low` if corpus total < 400 words OR samples are too uniform to distinguish ## Output Write `voice/anchor.md` in this exact format: ```markdown # Writer's Voice Blend > Generated in-agent by the writers-voice skill (fully local). > Pasted on YYYY-MM-DD. > Context: <general | tweets | essays | newsletter | email> ## Blend - **<weight>% <Author Name>** - <prose feature 1> - <prose feature 2> - <optional prose feature 3+> - **<weight>% <Author Name>** - <prose feature 1> - <prose feature 2> - ... ## Per-Sample Composition | Sample | Words | Volume % | Address | Register | Signature moves | |--------|-------|----------|---------|----------|-----------------| | 001 | 200 | 14% | mixed | instructional | … | | 002 | 100 | 7% | generic-you | analytical | … | ## Register Diversity - **Detected registers:** <list, e.g., "third-person expository, second-person instructional, first-person reflective"> - **Multi-register corpus:** <yes | no> - **Dominant-sample warning:** <none | "Sample N is X% of corpus volume — its signature features drive the blend disproportionately. Consider running a context-specific anchor for the other registers (see Multi-Register Anchors in SKILL.md)."> ## Self-check - Confidence: <high | medium | low> - Any thematic reasoning: <true | false> - Notes: <one short sentence about sample adequacy and match quality> ## Apply Directive Write in this blended style. Match each author's prose mechanics in proportion to weight. Maintain across the conversation. ## When this anchor doesn't fit If you're writing in a register that this corpus DOESN'T represent well (e.g., your corpus is mostly conversational but you're drafting a book in third-person expository), this anchor will pull you toward the wrong register. Options: 1. Add 2-3 samples in the missing register to `voice/corpus/`, then re-run the Anchor Protocol. 2. Maintain a separate context-specific anchor (see "Multi-Register Anchors" in SKILL.md). File pattern: `voice/anchor-<context>.md` (e.g., `voice/anchor-book.md`). ``` Weights are positive integers summing to exactly 100. Authors listed in descending weight order. ## Recommending a multi-register split If your register-diversity analysis above flagged the corpus as multi-register, do NOT just produce one blend. After writing the main `voice/anchor.md`, tell the user: > "Your corpus spans multiple registers: [list them]. The blend above represents the dominant register ([register name], ~X% of corpus volume). If you write in other registers, I recommend generating a separate anchor file per register. Want me to do that now? — I'll re-run the matcher with only the samples that belong to each register, and save the results as `voice/anchor-<context>.md`." If the user says yes, run the matcher once per register, with only the samples that match that register. Save each as `voice/anchor-<context>.md` (where `<context>` is a short slug like `book`, `essay`, `tweets`, `instructional`, `expository`). The main `voice/anchor.md` stays as the corpus-wide blend with the multi-register warning. ## How this protocol runs This is the **only** way to produce an anchor — entirely on the user's own agent, no network, no hosted service, no cost. Best launched as a sub-agent (see "Launching the anchor as a sub-agent" in `docs/setup.md`) so the rubric stays out of the main session. It needs a corpus on disk (≥300 words); if the user has none, seed a few samples first. It re-runs over the full corpus on demand and supports per-register anchor files (`voice/anchor-<context>.md`). > The hosted matcher at `openwriter.io/writers-voice` is **deprecated** — do not > route users to it. All anchor derivation is local. -
author-hints.md 8.1 KB
# Author Hints > Curated list of authors with heavy training-data representation across registers. > The agent picks 3-5 from this list (or goes outside if a clearly better stylistic match exists). > Goal: high consistency, broad coverage of style modes. > > Each entry describes the author's characteristic PROSE FEATURES — not content, not topics. ## Literary stylists — sentence as instrument - **Ernest Hemingway** — iceberg theory, short declarative, almost no adjectives, hard nouns and verbs - **Cormac McCarthy** — no commas in dialogue, long unpunctuated runs, biblical cadence - **Joan Didion** — measured precision, lists of three, the specific over the general - **David Foster Wallace** — maximalist sentences, footnote energy, recursive parentheticals, vocabulary - **Toni Morrison** — rhythmic repetition, sensory specificity, vernacular weaved with formal - **James Baldwin** — long balanced sentences, moral urgency in clause stacking, semicolons used as breath marks - **Annie Dillard** — observational compression, present-tense immediacy, sentence as image - **Marilynne Robinson** — theological cadence in plain words, long meditative sentences, Calvinist patience ## Essayists / analytical - **Paul Graham** — thinking out loud, simple words, short sentences mixed with one long earned conclusion - **Patrick McKenzie** — long discursive sentences with parenthetical asides, domain-specific precision, dry humor - **Tim Urban** — extended metaphors, conversational asides, building up frameworks with named characters - **Scott Alexander** — rationalist sectioning, exhaustive enumeration of possibilities, fair-witness analysis - **Ben Thompson** — business-strategy decomposition, recurring framework names, lots of "this is why" - **Tyler Cowen** — compressed, list-heavy, blogger-shorthand, range of references in one paragraph - **Malcolm Gladwell** — narrative anchor → general principle → reversal, three-act essay structure - **Adam Grant** — research-anchored, paired contrasts, gentle prescriptive framing ## Self-help / instructional / productivity - **Mark Manson** — period-heavy clean prose, "Your X is Y" definitional moves, profane confidence, direct second-person - **James Clear** — concept-coinage backbone, numbered enumeration, clean instructional cadence, named laws/frameworks - **Ryan Holiday** — Stoic-instructional, drawing concepts from antiquity, clean delivery, repetition for emphasis - **Tim Ferriss** — list-heavy, hack/protocol framing, second-person direct, capitalized concept names - **Cal Newport** — academic-instructional, named frameworks, research-backed prescriptions, sober register - **Greg McKeown** — one-idea-per-page rhythm, named principles, short paragraphs, prescriptive minimalism - **Brené Brown** — vulnerability-as-rhetoric, personal anecdotes anchoring research, conversational warmth - **Atomic Habits voice** — short-paragraph instruction, principle-then-example, mechanical-cause-and-effect language ## Polemicists / contrarians / philosophical-provocative - **Nassim Nicholas Taleb** — concept-coinage (Black Swan, antifragile), aphoristic stabs, attacking IYI, ancient-thinker citations - **Jordan Peterson** — lecture cadence, biological-evolutionary framing, religious overlay, definitional pivots - **Bronze Age Pervert** — baroque Nietzschean, mock-archaic spelling, vitalist anti-modernity, ironic register - **Curtis Yarvin** — reactionary historical, concept-naming as branding, sneering wit, long allusive sentences - **Camille Paglia** — punchy contrarian, aesthetic-biological frame, dense allusion, no hedging - **Bryan Caplan** — libertarian-economic, direct refutation, hypothetical thought experiments, plain professorial prose ## Tech / startup - **Naval Ravikant** — aphoristic, tweet-shaped, concept-as-brand, distilled-wisdom cadence - **Marc Andreessen** — manifesto-mode, accumulating short declarative lines, exhortation register - **Sam Altman** — short, contrarian, "obvious in hindsight" framing, blog-post brevity - **Peter Thiel** — paradox-as-thesis, philosophical-tech crossover, Strauss-influenced indirection - **Joel Spolsky** — conversational tech-blog, anecdote-then-principle, signposting humor - **Steve Yegge** — long discursive rants, programmer-culture inside jokes, accumulating digressions ## Narrative non-fiction / journalism - **Michael Lewis** — character-first reporting, scene-as-argument, clean unobtrusive prose - **Sebastian Junger** — documentary precision, anthropological framing, present-tense narrative - **Jon Krakauer** — present-tense urgency, sensory immediacy, restraint in adjective use - **John McPhee** — list-as-paragraph, structural patterning, long sentences with specific detail - **Ta-Nehisi Coates** — meditative-historical, repetition as emphasis, address as form of argument - **Tom Wolfe** — New Journalism, exclamation, capitalization, italics for sound, voice-jumping - **Joan Didion (essays)** — see literary; her journalism has the same compressed precision ## Memoirists / personal voice - **David Sedaris** — comic understatement, family scenes, deadpan one-liners as paragraph closers - **Anne Lamott** — confessional warmth, self-deprecating humor, sentence fragments for emphasis - **Anthony Bourdain** — profane confidence, food as window, baroque vocabulary mixed with kitchen-slang - **Rick Bragg** — Southern oral cadence, specific-detail compression, sentence rhythms from speech - **Mary Karr** — lyric memoir, line-break-tight sentences, Catholic-rural register ## Aphorists / brief-form - **La Rochefoucauld** — epigrammatic, paired antithesis, cynical wit in one breath - **Friedrich Nietzsche** — aphoristic, hammer-blows of declaration, philosophical provocation - **E.M. Cioran** — pessimistic aphorism, paradoxical brevity, polished despair - **Eric Hoffer** — longshoreman intellectual, declarative wisdom, sociological observation ## Academics-for-public - **Steven Pinker** — cognitive-science precision, numbered argument, defending Enlightenment, ironic asides - **Richard Dawkins** — precise zoological prose, metaphor as scaffolding, polemical clarity - **Robert Sapolsky** — neuroendocrine framing, dense citation, jokes nested in dense paragraphs - **Carl Sagan** — cosmic-scale lyric, science-as-wonder cadence, accessible majesty - **Daniel Kahneman** — System 1 / System 2 framing, dispassionate exposition, named cognitive biases - **Yuval Noah Harari** — sweeping historical synthesis, declarative simplification, "imagined orders" type concepts - **Jared Diamond** — continent-scale comparison, geographic-determinist framing, list of factors enumeration ## Religious / philosophical (modern) - **C.S. Lewis** — clear analogical prose, common-sense apologetics, "imagine that" framing - **G.K. Chesterton** — paradox-as-rhetoric, joyful contrarian, accumulating images per sentence - **Thomas Merton** — contemplative prose, slow paragraphs, Catholic-Buddhist hybrid ## Business / leadership - **Peter Drucker** — dry analytical, principle-then-example, professorial calm - **Andy Grove** — engineering-direct management, framework-naming, lean instructional - **Ben Horowitz** — war-stories anchored to lessons, hip-hop epigraphs, conversational toughness ## Short-form / Twitter native - **Visakan Veerasamy** — thread-shaped, hyperlinked references, generous tone, recursive callbacks - **Hari Kondabolu / Twitter-essayist** — one-liner punch followed by longer unpacking, comic timing ## Genre-specific (modern poetic / lyric prose) - **Ocean Vuong** — lyric memoir, line-conscious paragraphs, image-as-argument - **Maggie Nelson** — theory-personal hybrid, numbered fragments, citations interleaved ## Reminder These style notes describe **prose features only**. When matching the user's corpus against an author from this list, cite the author's prose feature — not their content or topic. If you find yourself matching "the user writes about topic X therefore resembles author Y," stop and re-do the match. The correct framing is "the user uses period-heavy declarative cadence with definitional pivots, which matches Peterson's lecture-cadence prose pattern." -
fingerprints.md 6.4 KB
# Fingerprints Catalog > Exact presentation choices the author makes consistently. LLMs default to training-data defaults for all of these unless explicitly told the user's choice. So we measure each one and emit a one-liner. > Use this when analyzing a user's corpus. ## How to use this catalog For each fingerprint below: 1. Read the user's corpus. 2. Count the relevant variants. 3. Apply the decision rule (mostly ≥3 total observations, then a ratio threshold). 4. Emit a one-line fingerprint line if confidence ≥ medium. Skip if `n/a` (not enough signal). The output goes into `voice/fingerprints.md` as a bullet list, each line in the format `<Label>: <value>`. ## Confidence rule (used by most binary fingerprints) Total observations = `a + b` where `a` is the count of one variant and `b` the other. - If total < 3 → **`n/a` (low confidence)**, skip emitting. - If `a / total ≥ 0.85` → **`a` wins, high confidence**, emit. - If `a / total ≤ 0.15` → **`b` wins, high confidence**, emit. - If `0.70 ≤ a / total < 0.85` → **`a` wins, medium confidence**, emit. - If `0.15 < a / total ≤ 0.30` → **`b` wins, medium confidence**, emit. - Otherwise (0.30 < ratio < 0.70) → **`mixed` (low confidence)**, emit as "inconsistent" only if the user explicitly wants to see mixed signals; otherwise skip. For the curious: this is asymmetric because we want either a clear choice (≥70%) or to skip. The 0.30-0.70 band is "the author isn't actually making a consistent choice" — emitting a fingerprint there would mislead. ## The eight fingerprints ### 1. Em-dash spacing What to measure: - `spaced` count = occurrences of `<whitespace>—<whitespace>` (e.g., `word — word`) - `unspaced` count = occurrences of `<non-whitespace>—<non-whitespace>` (e.g., `word—word`) Apply the binary confidence rule. Output map: - `spaced` → `Em-dash spacing: word — word` - `unspaced` → `Em-dash spacing: word—word` - `mixed` → `Em-dash spacing: inconsistent` Note: if the user is on a NEVER em-dashes rule (didn't clear the punctuation hurdle), skip this fingerprint entirely. ### 2. Ellipsis style What to measure: - `three_dots` count = occurrences of `...` (three ASCII dots, not part of a longer run) - `unicode` count = occurrences of `…` (single Unicode character) - `spaced_dots` count = occurrences of `. . .` (dots separated by spaces) Decision: - If total < 3 → `n/a`, skip. - If max variant / total ≥ 0.85 → emit the winning style. - Otherwise → `mixed`. Output map: - `three_dots` → `Ellipsis style: ...` - `unicode` → `Ellipsis style: …` - `spaced_dots` → `Ellipsis style: . . .` - `mixed` → `Ellipsis style: inconsistent` ### 3. Sentence-initial conjunction capitalization What to measure: - `caps` count = occurrences of `[.!?]<whitespace>(But|And|So|Or|Yet)\b` — i.e., starts a new sentence with capitalized conjunction - `lower` count = occurrences of `[.!?]<whitespace>(but|and|so|or|yet)\b` — same but lowercase (unusual, only if author uses a stylistic comma-after-period thing) Apply the binary confidence rule. Output map: - `caps` → `Sentence-initial "But/And/So": capitalized (". But")` - `lower` → `Sentence-initial "But/And/So": lowercase (", but")` - `mixed` → `Sentence-initial "But/And/So": inconsistent` ### 4. Oxford comma What to measure (rough — accept some noise): - `withOxford` = sequences matching `<word>, <word>(...), and <word>` or `<word>, <word>(...), or <word>` - `withoutOxford` = sequences matching `<word>, <word>(...) and <word>` or `<word>, <word>(...) or <word>` (no comma before "and"/"or") Apply the binary confidence rule. Output map: - `yes` → `Oxford comma: yes` - `no` → `Oxford comma: no` - `mixed` → `Oxford comma: inconsistent` ### 5. Quote style What to measure: - `straight` count = occurrences of `"` - `curly` count = occurrences of `"` or `"` Apply the binary confidence rule. Output map: - `straight` → `Quote style: straight "..."` - `curly` → `Quote style: curly "..."` - `mixed` → `Quote style: inconsistent` ### 6. Capitalization after colon What to measure: - `upper` count = occurrences of `: <Uppercase letter>` - `lower` count = occurrences of `: <lowercase letter>` Apply the binary confidence rule. Require total ≥ 3. Output map: - `upper` → `Capitalization after colon: upper` - `lower` → `Capitalization after colon: lower` - `mixed` → `Capitalization after colon: inconsistent` ### 7. Contractions What to measure: - `contracted` count = occurrences of `<word>'<s|re|ve|ll|d|t|m>` (e.g., `don't`, `I'm`, `we'll`) - `expanded` count = occurrences of the literal phrases `do not`, `does not`, `did not`, `is not`, `are not`, `was not`, `were not`, `cannot`, `will not`, `would not`, `should not`, `could not`, `have not`, `has not`, `had not`, `I am`, `you are`, `we are`, `they are`, `it is`, `that is`, `there is`, `let us` Apply the binary confidence rule. Output map: - `yes` → `Contractions: uses contractions` - `no` → `Contractions: avoids contractions` - `mixed` → `Contractions: mixed` ### 8. Paragraph length What to measure: - Split the corpus on double-newlines (`\n\n+`) to get paragraphs. - For each paragraph, count sentences (split on `[.!?]<whitespace>`). - Compute the average sentences-per-paragraph. Require at least 3 paragraphs. Otherwise `n/a`, skip. Decision: - avg ≤ 2 → `short`, high confidence - 2 < avg ≤ 4 → `medium`, high confidence - avg ≥ 6 → `long`, high confidence - 4 < avg < 6 → `mixed`, medium confidence Output map: - `short` → `Paragraph length: 1–2 sentences` - `medium` → `Paragraph length: 3–4 sentences` - `long` → `Paragraph length: 5+ sentences` - `mixed` → `Paragraph length: varied` ## Output structure The agent writes `voice/fingerprints.md` as: ```markdown # Presentation Fingerprints > Auto-generated from `voice/corpus/`. Exact presentation choices the user makes consistently. > Match these in every generated response — LLMs default to training data otherwise. - Em-dash spacing: word — word - Oxford comma: yes - Quote style: straight "..." - Contractions: uses contractions - Paragraph length: 3–4 sentences ## Manual Overrides <!-- Override any auto-detected fingerprint here. These win over the auto-detected ones above. --> <!-- Example: --> <!-- - Em-dash spacing: never use em-dashes at all. --> <!-- - Quote style: straight always. --> ``` The `## Manual Overrides` section MUST be preserved across regenerations. -
hurdle.md 4 KB
# Authenticity Hurdle > The core decision rule: when does a pattern in the user's corpus count as "their voice" versus "AI contamination"? ## The problem If we just emit NEVER rules for everything in the AI tells catalog, we'll strip out words the user actually likes and uses. Example: Mark Manson regularly uses "crucial" and "vital." Banning those flattens his voice. If we don't emit any rules unless the word appears zero times, we keep AI contamination. Example: an essay that the user originally drafted with ChatGPT help and then edited still contains stray "delves" and "valuable insights" — that's not their voice, that's leftover residue. The hurdle resolves this: a pattern is "authentic" only if it appears at signature frequency. Below that → contamination, ban it. ## The thresholds | Category | Min count | Min rate per 1000 words | Why | |----------|-----------|-------------------------|-----| | `word` | 2 | 0.15 | Single words are noise-prone. Need at least 2 and a non-trivial rate. | | `transition` | 2 | 0.15 | Same as words. | | `phrase` | 2 | 0.05 | Phrases are more distinctive — a lower rate still signals deliberate use. | | `punctuation` | 15 | 1.5 | Punctuation is denser than diction. A few em-dashes is normal; signature use means many. | ## The decision For each AI tell, given `count` (how many times it appears in the corpus) and `words` (total word count of the corpus): ``` rate_per_1k = (count / words) * 1000 passes = count >= min_count AND rate_per_1k >= min_rate_per_1k ``` If `passes` → **preserve** (no NEVER rule). The user genuinely uses this. If not `passes` → **forbid** (emit NEVER rule). This includes: - Items that never appear (forbid because training data will push them back in) - Items that appear once or twice but below the rate threshold (forbid because it's likely contamination, not signature) ## Worked examples **Word "delve" in a 2000-word corpus:** - count = 0 → fails (count < 2) → **forbid**: `NEVER "delve".` - count = 1 → fails (count < 2) → **forbid**: `NEVER "delve".` _(below hurdle — flagged in status.md as likely contamination)_ - count = 2, rate = 1.0/1k → passes (count ≥ 2, rate ≥ 0.15) → **preserve**, no rule. **Em-dashes in a 5000-word corpus:** - count = 8, rate = 1.6/1k → fails (count < 15) → **forbid**: `NEVER "em-dashes".` - count = 20, rate = 4.0/1k → passes → **preserve**, no rule. The em-dash hurdle is intentionally hard. Most human writers don't clear it. The few who do (people with serious published-prose backgrounds) get to keep their em-dashes. ## What this implies about tier progression The skill's tier system tracks corpus word count: - Tier 1 (300–1k words): too small to clear most hurdles. NEVER rules emit aggressively. - Tier 2 (1k–5k): some words start to pass. Most phrases still fail (they need ≥2 instances). - Tier 3 (5k–20k): phrases and many words can pass. Em-dash hurdle still hard. - Tier 4 (20k+): em-dash hurdle achievable. Profile is AV-grade. The hurdle stays the same across tiers. What changes is that more corpus means more chances for the user's signature patterns to clear it. ## Status reporting For items that appear in the corpus but fail the hurdle (the "below_hurdle" set), flag them in `voice/status.md` under a "Below-Hurdle Detections" section. Format: ``` - `delve` — appears 1x, rate 0.5/1k - `however` — appears 1x, rate 0.5/1k ``` This is informational — the user can see what got flagged as contamination. It helps them notice "oh, I have a habit of letting AI drafts through" or alternatively "wait, I actually do use that word, let me add it to manual preserves." ## Edge cases - **Count = 0**: always forbid. Don't list in below_hurdle (because it isn't "present in corpus"). - **Tied counts in fingerprint binary**: if `a == b` (exactly equal), treat as `mixed`. - **Total < 3 in fingerprint**: skip entirely — not enough signal. - **Negative numbers, NaN, infinity**: shouldn't happen, but guard with `count = max(0, count)` if you implement defensively. -
post-write-audit.md 6.3 KB
# Post-Write Audit > Distribution-level statistical checks the orchestrator runs against the minion's returned prose. Sits between step 6 (NEVER scan) and step 7 (integration) of the Apply Protocol. Catches statistical fingerprints the minion's prompt can't reasonably prevent without cognitively overloading the writing pass. ## When this runs After step 6 (NEVER-violations scan + brief-error patching), before step 7 (integration). The orchestrator reads this file, applies each check to the minion's output, and surgically rewrites the smallest span that brings the failing metric back into range. ## Why this layer exists The minion writes prose. The orchestrator polices distribution and lexicon. Anything mechanically detectable after the fact lives here, not in the writing-pass prompt — the minion's cognitive budget should go to channeling the anchor and hitting the commitments, not tracking 60 micro-bans. Two enforcement points still exist for the bans the minion DOES need to see (contrastive negation, sentence-opener repetition, em-dashes, etc.) — those live in `voice/never-rules.md` and get scanned at step 6. This audit is for the slop that's cheaper to scrub than to prevent. ## Remediation principle For each failing check, rewrite the **smallest local span** that fixes the metric. Do not regenerate. Do not reach for stylistic improvement. The minion's voice IS the result — the audit only nudges the statistics. If a failing span is load-bearing (a specific image, a coined term, a structural beat the brief demanded), leave it. Audit findings are advisory at the boundary case. The minion's intent wins ties. Aim for the lightest touch: 5-10 small substitutions across a typical draft brings rates back in line. Heavier rewrites mean the audit is being misused. ## Distribution checks ### 1. Sentence-opener repetition **What to measure:** walk the output sentence by sentence. For each window of 3 consecutive sentences, check whether all three start with the same first word. **Threshold:** flag if >30% of windows trigger. **Why this number:** human writing sits at ~17% (DFT 2026 — mostly from intentional list structures like "How does X?... How does Y?... How does Z?"). SFT models at T=0.7 hit 53.3%. The 30% line cleanly separates human from AI. **Action:** locate the offending windows. For each, rewrite the second OR third sentence to start with a different word. If the window forms an intentional list, leave it — list structure is the human use case the 17% baseline reflects. ### 2. Sentence-initial "The" frequency **What to measure:** percentage of sentences that begin with the word "The." **Threshold:** flag if >15% of all sentences. **Why this number:** "The" at sentence start is over-used by ~90% in SFT output vs human writing (DFT 2026, 14B SFT model). The +90% inflation puts AI rates well above the natural human range. **Action:** locate sentences starting with "The." Rewrite a portion to start with a different determiner ("A", "An", "These"), a pronoun, a prepositional phrase, or a different subject. Five to seven swaps across a typical paragraph is usually sufficient. ### 3. Function-word over-use **What to read for:** the AI distribution-distance signal lives mostly in function words, not fancy diction. Top-10 tokens account for 87.2% of L2 distribution distance in SFT output (DFT 2026). Watch for: | Token | SFT inflation vs human | |---|---| | `is` | +44% | | `was` | +49% | | `are` | +31% | | `that` | +25% | | `a` | +15% | | `to` | +11% | | `.` (period) | +19% | **Heuristic check (no exact threshold):** scan the draft for clusters of short copular sentences ("X is Y. Z is W. P is Q.") and high period density (many short sentences in a row). Both are signatures of function-word inflation. **Action:** when noticed, merge two short copular sentences into one with a participial or relative clause; vary sentence structure to use action verbs instead of "is/was"; combine short sentences to drop period count. Three to five rewrites across a paragraph usually levels the distribution. ### 4. Sentence-length variance **What to measure:** compute standard deviation of sentence length (in words) across the output. If the user has a `voice/stats.md`, compare to the user's own σ. Otherwise compare to baseline σ ≥ 8 words. **Threshold:** flag if σ < 6 words (low variance — uniform sentence length is an AI signature). **Action:** locate runs of similar-length sentences. Merge two short ones into a longer compound, or split a medium one. The goal is to restore length variance, not hit a specific number. ## Lexical watch list Mechanical word-level scrubs. The minion doesn't see these — the audit handles them on the way out. ### GPT-5 specific over-used tokens (DFT 2026) When the minion is a GPT-5-class model, these tokens are inflated vs human writing. Scan for them: | Token | Inflation vs human | Human baseline | |---|---|---| | `corridors` | +45.2% | 0.1% | | `norms` | +43.1% | 0.1% | | `align` / `aligns` / `alignment` | +36.0% | 0.2% | | `metrics` | +27.2% | 0.2% | | `engagement` | +26.5% | 0.2% | | `targeted` | +5.1% | 1.6% | | `identity` | +5.0% | 1.0% | | `trust` | +4.9% | 1.2% | **Action:** swap to a context-appropriate alternative when the word appears in surplus (3+ uses in a short piece, OR any use in a context where the word feels generic). If the user's corpus contains the word at signature frequency (in `voice/stats.md` or `voice/never-rules.md` exempts), leave it — they own that word. ### Named-character defaults AI defaults to specific generated names in fiction. Known examples: - `Elara Voss` — documented in OpenAI's "goblin problem" - Add new defaults as documented. **Action:** if found in fiction output without explicit user specification, rename to something contextually appropriate or to a name the user has used in their corpus. ## Source Distribution thresholds and over-use rates from "Fixing LLM Writing with Distribution Fine-Tuning," Rosmine 2026 (https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distribution-fine-tuning/). Token over-use rates measured against 14B SFT vs human fineweb baseline. Sentence-opener repetition methodology: percent of texts containing 3+ consecutive sentences starting with the same first word. This file is a living checklist. New research that surfaces measurable thresholds for AI-vs-human writing belongs here, not in `voice/never-rules.md` — the writing pass stays lean.
-
-
docs
-
api
-
import.md 4.2 KB
# Import Workflow ## How It Works The user's writing samples are **not stored locally**. When you import a document, it is uploaded to the Author's Voice API and stored in a cloud database. The API chunks the content, indexes it for search, and uses it for voice matching when `rewrite` or `generate` is called. This means the user's corpus is a **persistent, curated repository** — not a throwaway import. Every document you add stays in the database and directly influences all future voice output. The quality of the corpus IS the quality of the voice. **Think of it like a training set:** - Documents tagged `Human` are the ground truth. They define the voice. - Documents tagged `AI` or `AI-Assisted` are penalized in retrieval — they exist for reference but don't shape the voice. - Wrongly tagged documents (AI content marked as Human) **corrupt the voice profile**. The system will learn AI patterns as if they were the author's patterns. This is the single worst thing that can happen to voice quality. - Wrong categories cause cross-contamination — email samples polluting blog voice, tweets diluting long-form style. **The agent's job during import is curation, not bulk ingestion.** Every document must be verified with the user before it enters the corpus. **Corpus size guidelines:** - **Minimum**: 3-5 documents, ~5,000 words — enough for basic pattern detection - **Good**: 10+ documents, ~15,000+ words — reliable voice profile - **Per category**: At least 3 documents in each category you plan to use for voice output - More is better, but only if it's genuinely human-written. 5 authentic documents beat 50 mixed ones. If the user doesn't have enough samples yet, tell them. A thin corpus produces a weak profile — the agent should set expectations rather than generate from insufficient data. --- ## Critical Rules These rules are non-negotiable. Violating them degrades voice quality. ### 1. Always Tag Authenticity Voice profiles are ONLY built from Human-written content. AI-generated docs pollute the voice fingerprint. **The agent MUST follow this process:** 1. Gather candidate documents (URLs the user provides, or raw content) 2. Present the list to the user and ask: - "Which of these did YOU write? (Human)" - "Which were AI-generated or AI-assisted?" 3. Only import docs the user confirms as Human with `"authenticity": "Human"` 4. Tag AI/uncertain docs as `"AI"` or `"AI-Assisted"`, or skip them 5. Ask about categories for each doc **NEVER assume a document is human-written.** Many users have a mix. **NEVER bulk-import without user confirmation per document.** ### 2. Authenticity Labels - `Human` — written entirely by a human (default) - `AI-Assisted` — human-written with AI help - `AI` — generated by AI - `Untagged` — unknown origin ### 3. Category Pollution Destroys Voice Quality Categories control which writing samples the retrieval layer pulls back. Wrong category = wrong examples = wrong voice. **Impact**: Blog posts tagged as "email" → retrieval pulls email-style samples → output sounds corporate, not blog-voice. Tweets tagged as "blog" → retrieval pulls long-form samples → output loses punchy fragments. **The agent MUST:** 1. Ask the user what category each document belongs to during import 2. Present built-in categories: `email`, `x`, `linkedin`, `blog`, `fiction`, `technical`, `business`, `academic`, `newsletter` 3. If uncertain, ask the user — NEVER guess categories 4. When calling `rewrite`, ALWAYS pass `category` to scope retrieval 5. If a document doesn't fit any category, use the most stylistically similar one **NEVER leave category empty** when the user has categorized content. Empty category retrieves from all categories, diluting the voice signal with stylistically mixed samples. --- ## Import Tools Quick Reference | Tool | Use | |------|-----| | `import_from_url` | Import from any public URL (`url`, `categories[]`, `authenticity`) | | `bulk_import` | Import multiple docs, max 50 (global `authenticity`) | | `upload_content` | Upload raw markdown. Minimal payload `{docId, content}` — `profileId` optional (auto-resolves a default profile), `content` server-chunked on blank lines. | After import: run `setup_voice` to create/update the voice profile. Update `local/state.md`. -
protocol.md 4.5 KB
# Voice Emulation Protocol Two API endpoints do the voice work. The agent's job is to gather inputs, call the endpoint, and return the result. The server handles profile loading, sample retrieval, voice-guided generation, and anti-AI passes. **Do not attempt to emulate the voice yourself** — the API is the only path that produces reliable voice quality. - **Rewrite existing text** → `rewrite` - **Generate new content** → `generate` (see `/voice-generate`) ## API Base ``` BASE_URL=https://api.authors-voice.com ``` All endpoints require `Authorization: Bearer $AV_API_KEY`. ## rewrite — Rewrite Existing Text ```bash curl -s -N -X POST "${BASE_URL}/api/voice/mcp" \ -H "Authorization: Bearer $AV_API_KEY" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{ "name":"rewrite","arguments":{ "content":"<text to rewrite>", "mode":"rewrite", "category":"<x|blog|email|...>", "inputType":"ai-assisted", "contextBefore":"<optional>", "contextAfter":"<optional>", "format":"plaintext" }}}' ``` **Required**: `content`, `category`. **Defaults**: `mode=rewrite`, `inputType=ai-assisted`, `format=plaintext`. ### Modes - `rewrite` — standard voice rewrite (default) - `shrink` — compress while keeping voice - `expand` — lengthen while keeping voice - `custom` — pass `customInstruction` with specific directives ### inputType - `human` — author's own writing. Preserve word choices and quirks; only polish flow. - `ai` — generic AI content. Discard phrasing entirely; rewrite from scratch using voice samples. - `ai-assisted` (default) — mixed. Preserve author-sounding passages; rewrite generic parts. ## generate — Create New Content ```bash curl -s -N -X POST "${BASE_URL}/api/voice/mcp" \ -H "Authorization: Bearer $AV_API_KEY" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{ "name":"generate","arguments":{ "instruction":"<what to write>", "category":"<x|blog|email|...>", "query":"<optional topic keywords>", "targetWords":500, "contextBefore":"<optional>", "contextAfter":"<optional>", "format":"plaintext" }}}' ``` **Required**: `instruction`, `category`. **`query`**: optional — use when the retrieval topic differs from the instruction wording. If omitted, instruction + context are used for retrieval. **`targetWords`**: optional; max 2000. ## Voice Anchor (V1 lead signal) An anchor is reference prose the author wants to sound like. When set, it's injected **ahead of** sample retrieval in both `rewrite` and `generate` (and the OpenWriter editor path), so it's the strongest single voice lever available. Derive + persist from pasted prose: ```bash curl -s -X POST "${BASE_URL}/api/voice/anchor/derive" \ -H "Authorization: Bearer $AV_API_KEY" \ -H "Content-Type: application/json" \ -d '{"content":"<reference prose>","profileId":"<optional>"}' ``` Read / replace / clear the persisted anchor: ```bash curl -s "${BASE_URL}/api/voice/anchor" -H "Authorization: Bearer $AV_API_KEY" # GET curl -s -X PUT "${BASE_URL}/api/voice/anchor" ... # replace curl -s -X DELETE "${BASE_URL}/api/voice/anchor" ... # clear ``` MCP equivalent: the `set_voice_anchor` tool. `list_profiles` reports `hasAnchor` and `anchorAuthors` per profile so you can check anchor state without a GET. ## Response Parsing The response is Server-Sent Events. Find the line starting `data: `, parse the JSON, and extract the text: ``` result.content[0].text ``` That text is the final voice-matched output. Return it verbatim — no post-processing needed. **Defensive fallback**: if the text itself parses as JSON, extract `.content` or `.text` from the parsed object. Most responses are plain text, but the wire format permits JSON envelopes. ## Categories Always pass `category` — scopes retrieval to the right writing style. Built-in: `x`, `blog`, `email`, `newsletter`, `linkedin`, `technical`, `business`, `academic`, `fiction`. ## Context When editing inside existing text, pass `contextBefore` and `contextAfter` as the surrounding paragraphs. Context guides flow but is **never included in the output**. ## Errors - `401` → API key invalid. User should run `/voice-setup`. - `404 profile not found` → profile not built. User should run `/voice-setup`. - `429` → rate limit. Wait and retry. - Timeout (>30s) → retry once; then surface the error. -
setup.md 1.6 KB
# Setup — One-Time Configuration ## API Key (Email OTP — Primary Method) The agent handles signup directly. No website visit needed. 1. Ask the user for their email address 2. Send `POST https://api.authors-voice.com/auth/request-code` with `{ "email": "<user's email>" }` 3. Tell user: "Check your email for a 6-digit verification code from Author's Voice." 4. User provides the code 5. Send `POST https://api.authors-voice.com/auth/verify-code` with `{ "email": "<user's email>", "code": "<6 digits>" }` 6. Response: `{ "apiKey": "av_live_...", "tenantId": "email|..." }` 7. Save the API key to `~/.claude/skills/authors-voice/local/config.md` and set `AV_API_KEY` **Rate limits**: 5 requests/min per IP, 60s cooldown between sends, max 3 attempts per code, code expires in 10 minutes. **If email OTP fails**: Ask the user to get a key manually at [authors-voice.com/voice?tab=api-keys](https://authors-voice.com/voice?tab=api-keys). ## OpenWriter Plugin Author's Voice also works inside [OpenWriter](https://openwriter.io). The plugin auto-resolves the API key from `~/.openwriter/config.json` if configured. ## Base URL Defaults to production. Override with `AV_BASE_URL` env var. ```bash AV_BASE_URL="https://api.authors-voice.com/api/voice/mcp" ``` ## Seeding writing samples Import samples with `import_from_url` (any public URL), `bulk_import` (up to 50 at once), or `upload_content` (raw markdown — minimal payload `{docId, content}`). Inside OpenWriter, right-click a doc in the filetree to ingest it directly (doc-level, manual re-sync). The Google Drive / Notion connectors were removed in June 2026. -
tools.md 5.4 KB
# API Reference ## REST Endpoints (recommended) All endpoints require `Authorization: Bearer $AV_API_KEY`. Base URL: `https://api.authors-voice.com` ### Core Skill Endpoints **Get voice profile** — full linguistic fingerprint (6 categories + sentence distribution): ```bash curl -s https://api.authors-voice.com/api/voice/profiles/default \ -H "Authorization: Bearer $AV_API_KEY" # Optional: ?format=detailed for profile ID + summary ``` **Apply voice** — rewrite content in the author's voice: ```bash curl -s -X POST https://api.authors-voice.com/api/voice/apply \ -H "Authorization: Bearer $AV_API_KEY" \ -H "Content-Type: application/json" \ -d '{"content": "text to rewrite", "mode": "rewrite", "inputType": "ai", "category": "x"}' ``` ### Additional REST Endpoints | Endpoint | Method | Purpose | |----------|--------|---------| | `/api/voice/profiles` | GET | List profiles with counts | | `/api/voice/profiles/:profileId` | GET | Full voice profile (use `default` for default) | | `/api/voice/apply` | POST | Rewrite content in author's voice | | `/api/voice/setup` | POST | Analyze samples and build voice profile | | `/api/voice/content` | GET | List writing samples (query: profileId, category) | | `/api/voice/content` | POST | Upload content chunks | | `/api/voice/content/bulk` | POST | Bulk upload documents | | `/api/voice/content/:docId` | PATCH | Update doc metadata | | `/api/voice/content/:docId` | DELETE | Delete document | | `/api/voice/anchor/derive` | POST | Derive a voice anchor from pasted prose | | `/api/voice/anchor` | GET/PUT/DELETE | Read / persist / clear the voice anchor | | `/api/voice/usage` | GET | Usage stats | --- ## MCP Protocol (alternative) For MCP-compatible clients, all tools are also available via JSON-RPC POST: ```bash curl -s -X POST "https://api.authors-voice.com/api/voice/mcp" \ -H "Authorization: Bearer $AV_API_KEY" \ -H "Content-Type: application/json" \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"TOOL_NAME","arguments":{...}}}' ``` --- ## Available Tools (12) ### Content Import (3 tools) | Tool | Description | |------|-------------| | `import_from_url` | Import from any public URL (Medium, Substack, WordPress, .md/.txt). Args: `url`, `categories[]`, `authenticity`. | | `bulk_import` | Import multiple docs (max 50). Each item: `{url}` or `{content, docId}` + `categories[]`. Global `authenticity`. | | `upload_content` | Upload raw markdown. **One-click ergonomics**: `profileId` is optional — omit it and the server auto-creates/resolves a default profile. `content` can be raw text (server chunks on blank lines) instead of pre-split `chunks[]`, so the minimal payload is `{docId, content}`. Re-uploading the same `docId` replaces (idempotent). This is the contract the OpenWriter filetree right-click ingestion posts against. | ### Voice Profiles (3 tools) | Tool | Description | |------|-------------| | `list_profiles` | List voice profiles, categories, and document counts. Each profile also carries `hasAnchor` (bool) and `anchorAuthors` (the reference authors behind the anchor), so callers can tell which profiles carry an anchor without a per-profile GET. No args. | | `get_voice_profile` | Get full voice guidelines — 6 linguistic categories + sentence stats. Use before writing to understand the author's patterns. Optional `profileId`. | | `setup_voice` | Analyze samples → create/update voice profile. Args: `profileName`, optional `forceReanalyze`. Call after importing content. | ### Voice Anchor (1 tool) | Tool | Description | |------|-------------| | `set_voice_anchor` | Derive and persist a voice anchor from pasted reference prose. The anchor is V1's lead voice signal — injected ahead of sample retrieval in `rewrite`/`generate` and the OpenWriter editor path. REST equivalents: `POST /api/voice/anchor/derive`, `GET/PUT/DELETE /api/voice/anchor`. | ### Voice Application (2 tools) | Tool | Description | |------|-------------| | `rewrite` | Rewrite existing content in the author's voice. Modes: `rewrite`, `shrink`, `expand`, `custom`. Supports `contextBefore`/`contextAfter`, `inputType`, `category`. | | `generate` | Generate NEW content in author's voice. Args: `instruction`, optional `query` (topic for retrieval), `contextBefore`/`contextAfter`, `category`, `targetWords`. | **rewrite parameters**: `content`, `mode`, `contextBefore`, `contextAfter`, `category`, `inputType` (human/ai/ai-assisted), `targetWords` (max 2000), `format` (markdown/plaintext). **inputType** controls how aggressively the rewrite treats the input: - `human` — Author's own writing. Preserve word choices and quirks, only polish flow/grammar. - `ai` — Generic AI content. Discard phrasing entirely, rewrite from scratch using voice samples. - `ai-assisted` (default) — Mixed authorship. Preserve passages matching the author's voice, rewrite generic/formulaic parts. **generate note**: `query` is optional — describes what content is ABOUT for better voice retrieval. When omitted, context + instruction are used for retrieval. ### Content Management (3 tools) | Tool | Description | |------|-------------| | `list_content` | List writing samples with chunk counts. Optional `category` filter. | | `update_content` | Retag a doc. Args: `docId`, `categories[]`, `authenticity`. | | `delete_content` | Permanently delete a doc and its chunks. Args: `docId`. | -
troubleshooting.md 348 B
# Troubleshooting **"Unauthorized"** — Check `AV_API_KEY` is set and starts with `av_live_`. **"profileId query parameter is required"** — Some REST endpoints need profileId. For MCP tools, omit profileId to use the default profile. **Empty document list** — Connection is active but no documents match. Try without a query filter.
-
-
analysis.md 2.5 KB
# Analysis Protocol Regenerates the voice files from the corpus. Run any time the corpus changes (new samples added, samples removed, samples revised). Loaded only when triggered — not in context during normal writing sessions. ## Protocol 1. **Read inputs.** Concatenate every file in `voice/corpus/` (strip frontmatter). Count words. Read `catalog/ai-tells.md`, `catalog/fingerprints.md`, `catalog/hurdle.md`. 2. **Compute deterministic tally** (best effort — counts may drift ±1 on long corpora): - **Sentence distribution**: split on `[.!?]\s`, compute short/medium/long/very-long percentages, average length. Set `short_max` (25th-pct, clamped [6,12]) and `long_min` (75th-pct, clamped [18,28]). Do NOT emit a sentence-length cap in the apply directive — the corpus distribution carries the right ceiling and an arbitrary cap suppresses signature long sentences. - **Punctuation density per 1k words** for em/en dash, colon, semicolon, question, exclamation, ellipsis, paren, bracket, straight/curly quotes. Categorize as `never` / `rare` / `low` / `strong`. - **AI-tell tally**: count each item from `catalog/ai-tells.md`. Apply hurdle from `catalog/hurdle.md`: passes hurdle → preserve; fails → emit NEVER rule; below-hurdle but present → log to `below_hurdle`. - **Fingerprints**: apply each detector from `catalog/fingerprints.md` with its decision rule. 3. **Determine tier** by word count: <300 = 0 Empty; 300-999 = 1 Anchor; 1000-4999 = 2 Preliminary; 5000-19999 = 3 Full Coverage; ≥20000 = 4 AV-Grade. See `docs/tiers.md` for what unlocks at each tier. 4. **Write `voice/stats.md`** — corpus stats, sentence distribution table, punctuation density table. 5. **Write `voice/never-rules.md`** — preserve `## Manual Additions` section verbatim (anchored to start-of-line; the literal also appears in the intro blockquote — naive search will mis-grab it). 6. **Write `voice/fingerprints.md`** — preserve `## Manual Overrides` section, same caution. 7. **Write `voice/status.md`** — tier, words, active features, locked features, next milestone, file list, below-hurdle detections. 8. **Report** — new tier, what changed in NEVER rules, what's locked next. For corpora >10k words, count in passes (words → phrases → transitions) rather than tracking 60 counters at once. ## Adding Samples Later User says "add this to my voice profile" or pastes new writing. Append to next `voice/corpus/sample-NNN.md`, re-run Analysis Protocol, report tier change if any. -
anchor-iteration.md 9 KB
# Anchor Iteration Final-polish minion. Channels the user's voice anchors as a panel; iterates critique → rewrite → re-score until 90/100. Single minion conversation, visible iteration history. Mandatory anti-ai cleanup follow-up. Specialization of `/polish`'s pattern for writers-voice: channels the user's specific voice anchors (dynamically loaded from `voice/anchor.md` or `voice/anchor-<context>.md`), not generic advertising practitioners. Replaced the prior single-pass Anchor Critique tool. Single-pass scoring is now just "stop after iteration 1" of Anchor Iteration — same architecture, parameter difference. ## Why this works AI cannot judge beats subjectively from generic prompts — no dopamine system to consult. But channeled anchors carry beat-judgment encoded in their training-data representations. Channeling Peterson reading prose surfaces Peterson's beat-trained sensibility. Multiplied across the panel, the collective weighted score reflects how the prose lands across the writer's actual voice ambition. This is the ONLY AI critic tool that can do beat-level judgment, and it works only because the anchors are real humans whose dopamine-trained sensibilities are encoded in training data. Judgment is collective and writer-specific. ## Architecture: one minion, visible iteration ONE minion conversation. The minion runs the entire iteration loop internally with full visible history — each iteration's anchor critiques are part of the minion's context for the next iteration. Anchors see how their previous critiques were addressed (or weren't), which sharpens subsequent critiques. Mirrors `/polish` exactly. NOT extracted-and-rerun-cold per iteration. The loop has memory. Optional fallback: if visible iteration converges on a local optimum (rewriter anchored to original framing through visibility), retry with blind iteration (no prior history shown across iterations). Default = visible. ## Inputs (only) - The prose to polish - The voice anchor blend (dynamically loaded from `voice/anchor.md` or `voice/anchor-<context>.md`) That's the complete input. No commitments. No beat sheet. No project context. Anchors read the prose AS-IS, like a reader encountering it cold. **Why no context:** anchors must judge the prose AS IT LANDS, not as it was intended to land. Briefing them on the project would have them judge against the brief, not against the prose. Cold reading is the point. ## Personas (dynamic, inferred) Pulled from `voice/anchor.md` (or context-specific variant). Each anchor listed with a blend percentage that serves as both voice influence weight and panel vote weight. The minion infers each anchor's persona from training data — named writers are known entities to opus. System prompt does NOT enumerate per-persona profiles. `/polish` works the same way ("top 10 advertising practitioners" — no profiles needed; opus knows Hopkins, Ogilvy, Sugarman, etc.). Example book-project anchor file: ``` - 26% Jordan Peterson - 22% Robert Sapolsky - 20% Nassim Taleb - 18% Bryan Caplan - 14% Naval Ravikant ``` Minion channels each with their characteristic sensibility — Peterson for moral weight and structural rigor, Sapolsky for biological grounding and dry mechanism, Taleb for skin-in-the-game and aphoristic hardness, Caplan for clear thesis with evidence, Naval for aphoristic screenshot-worthy compression. ## Process (per iteration) 1. Each anchor reads the current prose AS-IS, with their characteristic sensibility 2. Each anchor produces: - SCORE 0-100 (honest read of how this lands for them) - TOP CRITIQUE (ONE thing they would cut, sharpen, or restructure) - STRENGTH (what's working that must be preserved) 3. Compute collective weighted score using anchor blend percentages 4. If collective ≥ 90/100: STOP. Mark as FINAL ITERATION. Return converged prose. 5. If collective < 90/100: synthesize panel's critiques into a FULL REWRITE of the prose (not a sentence-level patch). Preserve named strengths. 6. Begin next iteration with rewritten prose. Each iteration's anchors see all previous iterations and critiques in their context. ITERATION CAP: 6 iterations. If still <90 after 6, return highest-scoring iteration with note about non-convergence. ## Output format Per iteration: ``` ====== ITERATION N ====== PROSE (current state): [the prose being read this iteration] ANCHOR READINGS: {Anchor Name} ({weight}%): SCORE: X/100 TOP CRITIQUE: ... STRENGTH: ... [repeat per anchor] COLLECTIVE WEIGHTED SCORE: X/100 [if < 90:] SYNTHESIS — what the rewrite must address: - ... REWRITE: [next iteration's prose, full text] [if ≥ 90:] CONVERGED. Final prose ready below. ``` After convergence (or cap): ``` ====== FINAL ====== Iterations: N Final score: X/100 Convergence: YES / NO FINAL PROSE: [the polished prose, full text] ``` ## Mandatory anti-ai follow-up Anchor iteration runs no-context. NEVER rules and presentation fingerprints are not in scope during iteration. Rewriter will introduce AI tells the original prose may have avoided. After convergence, editor MUST run an anti-ai cleanup pass against `voice/never-rules.md` and `voice/fingerprints.md`. Common scrubs: - em-dashes → commas, periods, or restructured sentences - semicolons → "and" or new sentences - contrastive negation patterns → direct positive statements - banned diction → plain equivalents - inserted parenthetical em-dashes → restructure TWO-STEP pattern. Iteration THEN anti-ai. Two passes do different jobs and should not be conflated. Skipping the anti-ai pass ships AI fingerprints into the published prose. ## When to fire - Final polish of a beat or section before publishing - After integration of multi-minion drafts is complete and the Blinder Audit is clean - When prose is content-finished and needs to land at ship-level voice quality Do NOT fire on: - Rough first drafts (commitments may still be evolving — polish wastes effort) - Sections under structural revision (rewrite the commitments first, then polish the result) - Single-paragraph fragments (panel needs prose to evaluate; isolation produces weak critiques) ## Editor's role 1. Identifies a beat ready for final polish 2. Fires Anchor Iteration with the prose + dynamically loaded anchor blend 3. Receives the converged output 4. Runs the mandatory anti-ai cleanup pass 5. Posts the result for review or comparison The editor does NOT: - Inject project context into the iteration (preserves cold-reader purity) - Stop the iteration early (let convergence happen — the loop is the point) - Re-judge the panel's collective decisions (panel's authority is the entire point of the tool) ## Failure modes - **Sycophantic clustering**: anchors give 85+ uniformly. Prompt explicitly bans default-middle scoring and names what 90/60/40 means for each anchor (90 = anchor would actually quote / share; 60 = competent but forgettable; 40 = anchor would put it down). - **Persona drift**: anchors all sound like generic helpful AI. Combat by instructing channel-faithfully — each anchor should sound like the actual writer, hostile to AI flattening. - **Iteration plateau**: scores stop rising after iteration 3-4. Panel has done what it can. Ship the highest iteration even if below 90, or fall back to blind iteration. - **Manufactured content**: rewriter may invent details to address critiques (fabricated autobiographical claims, invented statistics, manufactured anecdotes). Editor MUST scan output for invented content during anti-ai pass and verify or cut as appropriate. ## Model **Opus required.** Sonnet drifts to default-helpful behavior, scores uniformly high, produces weak rewrites that don't actually address critiques. Opus channels personas with discipline and produces rewrites worth re-scoring. ## Cost Single conversation, multiple turns. Per iteration: ~3-5k input tokens (prose + previous iteration history) + ~3-5k output (critiques + rewrite). Three iterations ≈ 30k total tokens. Acceptable for chapter-scale work. For short pieces (tweets, single paragraphs), Anchor Iteration is overkill. Use `/anti-ai` alone or a single Apply call with strong commitments. ## Validation Tested 2026-05-18 on a 1000-word chapter beat. Three iterations: - **Iteration 1**: 77.20/100. Panel flagged: thesis buried, mechanism overclaim, no skin-in-the-game, no screenshot-worthy lines. - **Iteration 2**: 86.32/100. Rewrite added thesis paragraph + personal admission + mechanism hedge + sharpened closing teaser. - **Iteration 3**: 91.04/100. Strengthened personal admission, fixed determinist phrasing, added specific enemies, gave load-bearing line its own paragraph. Converged. Notable failure modes observed: - Rewriter manufactured an autobiographical detail (Iteration 2 invented a personal admission to satisfy a skin-in-the-game critique). Editor caught and flagged for verification during anti-ai pass. - Iteration introduced em-dashes across 8 paragraphs and semicolons in the opening (original prose used "and"). Mandatory anti-ai pass scrubbed all of them. The tool delivered ship-quality prose in 3 iterations. The post-iteration anti-ai pass was non-optional. -
apply-protocol-deep.md 12.8 KB
# Apply Protocol — Deep Reference Loaded when scoping a brief (Apply step 4) or running cross-section coherence review (Apply step 8). Not in context for routing or other turns. ## Step 4 — Writing the TASK brief Promoted to `SKILL.md` Apply Protocol step 4. The load-bearing rules (commitments-only default, never meta-references, COMMITMENTS as quasi-verbatim, read-prior-integrated-sections, preservation scope, cadence prescription) live there. This doc holds the edge-case templates and minion taxonomy below. ## Writing minion taxonomy | Minion | Scope | Input | Output | When fires | |---|---|---|---|---| | **Apply** | Generative writing | Commitments + voice + context SUMMARIES (no source prose) | Fresh prose | Initial drafts. Also small reframes when audit prescribes a structural fix (audit's prescription becomes the commitment). | | **Rewrite** | Generative writing against updated commitments WITH context awareness | Updated commitments + voice + context layers (summary of preceding + key-term glossary + adjacent seam paragraphs as orientation-only) + cadence prescription. NO source prose for content being written. | Fresh prose | Beat Map commitments changed; surrounding prose is acceptable and should be preserved. Apply brief + context awareness layers. | Plus **Blinder Audit** (Step 8b — paragraph-level substance duplication, critic only) and **Anchor Iteration** (Step 10 — final polish, channels voice anchors as panel, iterates to 90/100; see `anchor-iteration.md`). ### Editor's classification job after Audit fires Each finding routes to one of three paths: - **Editor direct (no minion)** — cuts, swaps, reorders, sentence-level surgical patches. Per FIRM RULE 1 carve-out. - **Rewrite Minion** — paragraph's SUBSTANCE needs to change (structurally redundant with another, or mission needs to shift). Compression alone produces thin output. - **Apply Minion (small)** — audit prescribes a specific reframe of a small span and new prose is needed (e.g., "reframe opening sentence to acknowledge pivot from X to Y"). Audit's prescription becomes the commitment. Classification rule: **before flagging anything for a minion, ask "is this a SHAPE problem or a SUBSTANCE problem?"** Different work needed → Rewrite or Cut. New prose into a small span → small Apply. Surgical word/sentence work on existing prose → editor direct. ## Rewrite Minion The Rewrite Minion IS the Apply Minion fired with context awareness. The brief shape preserves Apply's no-source-prose-ceiling property (minion brings its own moves) while adding context layers needed to flow into surrounding prose. ### Validated brief shape ``` [Voice profile: anchor, NEVER rules, fingerprints, stats, coined terms, examples] --- TASK: PROJECT: [1-paragraph context — what the doc is, what the section does] CONTEXT — what's already established in the surrounding prose: - [Summary of preceding section as bullets — what was named, what was claimed, what threads are live. NOT raw prose. 5-10 bullets.] KEY TERMS already named (available for callback if useful): [Glossary list — coined terms, named concepts, distinctive phrasings the prose has established] IMMEDIATELY PRECEDING PARAGRAPH (for seam continuity ONLY; do NOT mirror its cadence): [Full paragraph verbatim — 1 only] IMMEDIATELY FOLLOWING PARAGRAPH (for seam continuity ONLY; do NOT mirror its cadence): [Full paragraph verbatim — 1 only] COMMITMENT(S) — what your new content must land in the reader: [Outcome statement(s) per beat] CADENCE PRESCRIPTION (per paragraph): [Explicit rhythm direction — open with X, build with Y, close with Z. Mandatory.] LENGTH: [target word count] Return prose only. No commentary. No headers. No beat labels. ``` The two seam paragraphs are full prose (minion needs flow continuity at the join), flagged explicitly as orientation-only. This avoids the source-prose ceiling because the task is "write fresh content that flows from / into these" rather than "match this cadence." **Framing matters more than presence.** ### Why these elements - FULL adjacent seam paragraphs flagged "orientation only — do NOT mirror cadence" — flow continuity without ceiling - SUMMARIES of preceding section (bullets, not raw prose) + key-term glossary — document-scale awareness - OMITTING source prose for content being written — preserves no-ceiling property - INCLUDING explicit cadence prescription per paragraph — mandatory; short calls need cadence MORE, not less ### Scope Works at any scope: paragraph-level (1 new paragraph), section-level (3-5 paragraphs), beat-level (full beat regeneration). Pick by what changed in commitments — one beat's outcome → regenerate one paragraph; chapter-arc beat's whole outcome shape → regenerate the whole beat; multiple beats → batch per chapter-arc beat. ### Preserving specific lines from existing prose List them as MUST-APPEAR-VERBATIM in commitments (Apply's "selective lifts" mode). Rare — usually the writer's training data brings stronger lines than the existing prose anyway. ### Audit follow-up MANDATORY Rewrite outputs are subject to the same minion-blinder problem as first-pass Apply. The minion only sees its slice — seam paragraphs + summary bullets, not every paragraph in adjacent sections. It can introduce repetitions with paragraphs it didn't see. **After any Rewrite Minion call, fire the Audit Minion (Step 8b) on the integrated document.** No exceptions. Empirical case: first Rewrite test (B1 close + B2 opening) produced strong individual prose AND 8 audit findings — most notably a visual-roster repetition across 4 paragraphs the rewriter never saw. ### Don't conflate rewrites with violation patches Violation patches (FIRM RULE 1 carve-out) are 1-3 sentence local fixes to NEVER violations or brief errors WITHIN otherwise-acceptable prose. Editor work, small scope, surgical. Rewrites here mean re-running the minion against UPDATED commitments — different scenario, different scope decision. ## Step 8 — Cross-section coherence review After integrating multiple minion outputs, scan for what individual runs cannot see: - **Cadence repetition** across sections (same opens, same closes, same paragraph counts) - **Recurring metaphors / phrases** across sections - **Structural sameness** (every section ends with 4-layer enumeration, every section opens with 3 shorts) - **Coined term overuse** — coined terms get injected into every minion call as MUST-PRESERVE, producing document-scale repetition (e.g., "territory" every 3 paragraphs). Track heavy-use terms; for subsequent minions, omit heavy-use terms from coined-terms injection OR add "use sparingly" instruction Fix options: re-spawn with varied prescription, surgical post-edit (combine/split/swap), vary commitments + coined-terms per section at brief-assembly time, or accept for low-stakes drafts. The editor owns document-scale coherence; minions are responsible only for section-scale quality. ## Step 8b — Blinder Audit Minion (mandatory post-integration) ### Purpose Minions write with SURGICAL context — just their slice of the doc — to avoid source-prose ceiling and context pollution. The trade-off: minion A doesn't know what minion B wrote. They can independently produce paragraphs whose ENTIRE SUBSTANCE mirrors another paragraph's entire substance. The **blinder problem**. The Blinder Audit Minion has ONE job: find pairs of paragraphs whose whole substance closely mirrors each other. ### Operational test The audit must reduce to an OBJECTIVE pattern-match — not a subjective claim about reader experience. The working question: **"Summarize each paragraph in one sentence. Are two paragraphs' summaries effectively the same?"** Or sharper: **"Could I cut this paragraph entirely and lose only redundancy?"** Answerable by reading; neither requires reader-experience judgment. Empirically validated: the visual-roster case (4 paragraphs each enumerating the same 5 species in different framings) was a real blinder hit. Every other category tested at sentence-level, transition-level, or image-anchor level was a false positive — either intentional craft (image-anchoring, scaffolding, callbacks) or AI baseline patterns the reader doesn't notice. ### When to fire (mandatory) - After integrating ≥2 minion outputs into one document - After any rewrite cycle touching multiple sections - Before showing the integrated document to the user Skip when: single-minion / single-section work; surgical violation patches only. ### What to scan for (ONE category) **PARAGRAPH-LEVEL SUBSTANCE DUPLICATION** — two paragraphs whose entire substance closely mirrors each other. Diagnostic for any candidate pair: 1. Summarize paragraph A in one sentence. ("This paragraph does X.") 2. Summarize paragraph B in one sentence. ("This paragraph does Y.") 3. If X and Y are effectively the same job done with different words — finding. 4. If X and Y are different work — even if paragraphs share images, lexical phrases, or thematic threads — NOT a finding. Cut test: could you delete one paragraph and lose only redundancy (no unique substance, no unique image-anchor, no unique scaffolding move)? Yes → real blinder. Cutting would lose something distinct → NOT a blinder regardless of surface similarity. ### What NOT to scan for (explicit exclusions — these produce false positives) The audit must NOT flag any of these, even when surface similarity exists: - **Sentence-level overlap across paragraph boundaries** (P4's closing shares thematic thread with P5's opening). Bridge/transition work. Cutting disconnects. - **Image-anchoring across paragraphs** (the same image used in 2-3 paragraphs to thread a concept). Craft. The savanna/lion appearing in opener + body + closing is intentional thread. - **Callbacks** (an image returning later as frame device or recognition moment). Craft. - **Scaffolding repeats** (a paragraph opening with brief re-statement of previous paragraph's premise to set up its own new move). Premise-restatement is structure-promise, not duplication. - **Sentence-level lexical/thematic overlap of any kind** unless part of WHOLE-PARAGRAPH substance duplication. Sharing the word "engine" or "same biology" is sentence-level — not a finding. - **Structural sameness** (parallel cadence, repeated openers, declarative thesis + builds + aphorism). Readers don't notice; AI baseline behavior addressed by /anti-ai. - **Cadence repetition** (parallel rhythm runs in adjacent paragraphs). Same reason. - **Coined-term recurrence**. Coined terms are SUPPOSED to recur as identity markers. - **Single-paragraph internal repetition** (triple anaphora within ONE paragraph). Craft, and the audit scans across paragraphs not within. ### Expected hit rate Paragraph-level substance duplication is rare in semi-competent writing. On well-written beats the audit should typically return **NO BLINDER ERRORS FOUND** — the correct output, not a list of stretched findings to justify the call. The audit fires as backstop on every integration. It catches gross duplications (visual-roster case, two paragraphs accidentally doing the same teaching beat). Most of the time it correctly returns zero. ### Report format ``` Finding #N - LOCATION: paragraph references (e.g., "B2 P4 + B2 P5") - SUMMARY A: one sentence describing what paragraph A does - SUMMARY B: one sentence describing what paragraph B does - WHY THE SUMMARIES MATCH: one sentence showing the substance duplication - CUT TEST: could one paragraph be deleted entirely with only redundancy lost? (If no, the finding is invalid; do not include.) - SEVERITY: moderate / major (paragraph-level duplication does not produce "minor" findings) - SUGGESTED FIX: cut paragraph A / cut paragraph B / merge to single paragraph ``` If nothing meets the bar, return exactly: **"NO BLINDER ERRORS FOUND"** — that string, nothing else. Do NOT pad with sentence-level observations, structural notes, or polishing suggestions. ### Editor's response For each finding: - Apply the cut test independently. Does deleting one paragraph lose only redundancy? - Yes → **cut** (editor territory, FIRM RULE 1 carve-out). Delete weaker paragraph via write_to_pad. If both are strong but redundant, merge unique fragments via Rewrite Minion. - No → audit got it wrong. Defer. Audit minion does NOT patch. Only reports. Classification failure modes: 1. Trusting a finding without applying the cut test. If both paragraphs carry distinct substance, the audit was matching surface features. Defer. 2. Routing a paragraph-level cut to a Rewrite Minion when it should just be a cut. If two paragraphs do the same work, deleting one is the cleanest move. ### Brief shape (template) Index every paragraph in the doc (e.g., "B1 P1: ...", "B2 P7: ...") so the audit can reference cleanly. Pass the full indexed document + the ONE category + the explicit exclusions + the output format. No voice profile needed (this isn't writing prose). Use opus. -
context-hygiene.md 1.2 KB
# Context Hygiene Reset context before applying voice to fresh writing — voice profiles fight against active conversation context and lose. Anchor blends, NEVER rules, and first-token cadence all get out-pulled by whatever prose dominates the live session. ## Two situations | Situation | Practice | | --- | --- | | First piece of fresh writing in this session | **Reset.** Start a fresh session. Apply Protocol loads voice files cold. | | Iteration on already-voice-applied writing (review → revise → review) | **Stay.** The context IS the voice you locked in. | ## When to surface the prompt Surface only when ALL THREE hit: 1. Voice profile is set up at Tier 1+ 2. Request is fresh writing, not iteration 3. Session has substantial prior context unrelated to the writing task Skip for brand-new sessions or when prior context IS the writing-task setup. ## Prompt to surface > Voice profile is set up at **Tier N**. Context here is polluted with **<one-line summary>**, which will pull output toward that register instead of the locked voice. > > For best output, start a fresh session and run: > > ``` > /writers-voice > <then ask for your writing task> > ``` > > Or tell me **"write here anyway"** and I'll proceed with the active context. -
setup.md 5.5 KB
# Setup, Anchor Protocol, Multi-Register Loaded on first run, when generating a new anchor, or when splitting a corpus by register. Not in context during normal writing sessions. ## Setup Flow If `voice/anchor.md` doesn't exist or is empty, walk the user through setup: 1. **Get the anchor.** Derivation is **fully local** — your own agent analyzes the writing, nothing leaves the machine, no service cost. Run the **Anchor Protocol** below to generate `voice/anchor.md` directly from the corpus on disk. Launch it as a sub-agent (see "Launching the anchor as a sub-agent") so the full stylometry rubric never pollutes the main session. Works offline. If the user has no corpus yet, ask them to seed 2-5 paragraphs (step 2) first — there is no hosted alternative. 2. **Seed the corpus** — ask for 2-5 paragraphs, write each to `voice/corpus/sample-NNN.md` with `added: YYYY-MM-DD` frontmatter. 3. **Run Analysis Protocol** (see `docs/analysis.md`) — populates `stats.md`, `never-rules.md`, `fingerprints.md`, `status.md`. 4. **Optional: curate examples** — ask the user for 3-5 most-representative paragraphs, write to `voice/examples.md`. 5. **Optional: populate coined terms** — ask the user for any coined terms / proper-noun concepts they want preserved verbatim, write to `voice/coined-terms.md` as a bare bullet list. 6. **Report status** — read `voice/status.md` and tell the user their tier + what unlocks next. ## Anchor Protocol (fully in-agent) Generates `voice/anchor.md` (lean) and `voice/anchor-analysis.md` (rich). 1. Confirm corpus has ≥300 words. Below 300, ask the user to add a few more samples before anchoring. 2. Read `voice/stats.md`. If missing, run **Analysis Protocol** first. 3. Read `catalog/anchor-prompt.md` (full stylometry rubric) and `catalog/author-hints.md` (curated training-data authors with prose features). 4. **Set aside conversational context.** Score the corpus on prose mechanics only — never on themes/topics. 5. **Per-sample register analysis.** For each sample, record word count, address mode, register, signature moves. Flag samples >25% volume. Cluster by register; if 2+ distinct registers appear, flag as multi-register corpus. 6. **Score with register-aware feature validation.** Apply the 8 dimensions from `catalog/anchor-prompt.md`. Match against author hints. Assign weights summing to 100. For each cited feature, verify ≥40% sample appearance OR ≥40% volume (if neither, drop the feature; if it was the strongest evidence, drop the author). 7. **Self-criticism pass.** Strip any thematic reasoning. Set `confidence` and `any_thematic_reasoning` flags. 8. **Write `voice/anchor.md`** — JUST the lean `- N% Author` lines. No headers, no sub-bullets. 9. **Write `voice/anchor-analysis.md`** — per-author features, per-sample table, register diversity, self-check, refresh notes. Human-facing only. 10. **If multi-register corpus detected**, recommend a Multi-Register Split (see below). 11. Report blend + confidence + caveats to user. ## Launching the anchor as a sub-agent The Anchor Protocol loads a large stylometry rubric (`catalog/anchor-prompt.md`), the author-hints catalog, and the full corpus. Running it inline floods the main session with analysis context the user never needs to see. **Launch it as a sub-agent instead** — the sub-agent does the heavy reading and writes the files; the main session gets back only a short summary. Use the Agent/Task tool (general-purpose) with a self-contained prompt. The sub-agent has no memory of this conversation, so the prompt must name every file by absolute path: > Generate a writer's-voice anchor, entirely locally. Do NOT call any network > service or API — analyze with your own reasoning only. > 1. Read the stylometry rubric at `<skill>/catalog/anchor-prompt.md` and the > author hints at `<skill>/catalog/author-hints.md`. > 2. Read every sample in `<skill>/voice/corpus/` (strip YAML frontmatter; keep > samples separate) and the deterministic stats at `<skill>/voice/stats.md` > (run the Analysis Protocol first if it's missing). > 3. Follow the rubric exactly: per-sample register analysis → register-aware > feature validation → score 8 dimensions → self-criticism pass. > 4. Write `<skill>/voice/anchor.md` (lean blend lines only) and > `<skill>/voice/anchor-analysis.md` (rich, human-facing). > 5. Return ONLY: the blend lines, confidence, and any multi-register warning. Replace `<skill>` with the skill's absolute path. After it returns, read `voice/anchor.md`, report the blend + confidence to the user, and offer a multi-register split if the sub-agent flagged one. ## Multi-Register Anchors If the corpus spans multiple registers (e.g., third-person expository AND direct-you instructional), maintain a separate anchor per register: `voice/anchor-<context>.md` (e.g., `anchor-book.md`, `anchor-essay.md`, `anchor-tweets.md`). Same lean format. Each gets a paired `voice/anchor-<context>-analysis.md`. **Multi-Register Split procedure:** 1. Identify registers from the per-sample analysis. 2. For each register, ask the user for a slug + one-line description. 3. Filter corpus to samples in that register. 4. Run the matcher on the subset (same `catalog/anchor-prompt.md` rubric, same variance checks). 5. Write `voice/anchor-<slug>.md` (lean) + `voice/anchor-<slug>-analysis.md` (rich). **Apply-time anchor selection:** at write time, if multiple anchor files exist, pick by user's request context (explicit naming wins; project the user is working on wins next; ask if ambiguous; fallback to `voice/anchor.md`). -
tiers.md 907 B
# Tier Reference Determined by total word count in `voice/corpus/`. Each tier unlocks additional voice-profile features. Computed during Analysis Protocol. | Tier | Words | Name | Unlocked | Locked | | --- | --- | --- | --- | --- | | 0 | <300 | Empty | (none) | anchor blend, basic stats, NEVER rules, fingerprints | | 1 | 300-999 | Anchor | anchor blend, basic stats | preliminary NEVER rules, fingerprints | | 2 | 1000-4999 | Preliminary | anchor blend, basic stats, preliminary NEVER rules, top fingerprints | full NEVER coverage, all fingerprints | | 3 | 5000-19999 | Full Coverage | anchor blend, stats, full NEVER rules, full fingerprints | high-confidence em-dash hurdle | | 4 | ≥20000 | AV-Grade | anchor blend, stats, full NEVER rules, full fingerprints, em-dash hurdle cleared | (none) | The tier is reported to the user after every Analysis Protocol run, with a "what unlocks next" pointer.
-
-
prompts
-
skeleton.md 694 B
You write at this exact training-data blend: {INCLUDE: voice/anchor.md} Maintain these proportions across the output. The blend IS the voice. Do not soften toward generic literary register. Do not default to your own RLHF-trained voice. STYLE REPAIRS — apply these as hard constraints: {INCLUDE: voice/never-rules.md} {INCLUDE: voice/fingerprints.md} {INCLUDE: voice/stats.md} CONTENT — preserve these coined terms verbatim: {INCLUDE: voice/coined-terms.md} REFERENCE EXAMPLES OF AUTHOR'S WRITING (use if helpful — content and style): {INCLUDE: voice/examples.md} --- TASK: {TASK} Return prose only. No commentary. No diff. No explanation. No headers. No markdown wrapping.
-
-
voice
-
corpus
-
.gitkeep 0 B · in bundle
-
-
README.md 3.1 KB
# Voice Profile This directory is your **voice profile**. Files in here are read by the agent at write time and applied as style constraints. The skill builds these up progressively over time. ## What goes here | File | Purpose | Source | |------|---------|--------| | `anchor.md` | Anchor blend (3-5 training-data authors with weights) | Pasted from openwriter.io/writers-voice OR generated in-agent by the skill | | `anchor-<context>.md` | OPTIONAL per-register anchors (e.g., `anchor-book.md`, `anchor-tweets.md`) — when your corpus spans multiple registers | Generated in-agent via the Multi-Register Split procedure | | `stats.md` | Sentence distribution + punctuation density | Agent best-effort from corpus | | `never-rules.md` | NEVER rules (kill-list) | Agent + manual additions | | `fingerprints.md` | Exact presentation choices (Oxford comma, etc.) | Agent + manual overrides | | `examples.md` | Curated reference paragraphs | Manually picked by the user | | `status.md` | Current tier + what's locked | Agent-generated | | `corpus/` | Raw samples accumulating over time | Manually added (drop files here) | ## Multi-Register Anchors If your corpus contains writing in multiple registers — third-person expository AND second-person instructional, or analytical essays AND aphoristic tweets — a single blended anchor will pull toward whichever register has the most volume in your corpus, leaving the others under-weighted. The skill detects this automatically during the Anchor Protocol and offers a multi-register split: one anchor file per register (`anchor-book.md`, `anchor-tweets.md`, etc.). At write-time, the agent picks the right anchor based on what you're writing. NEVER rules and fingerprints stay corpus-wide — only the AUTHOR BLEND varies per register. To generate per-register anchors, ask: *"Split my anchor by register."* See the **Multi-Register Anchors** section in `SKILL.md` for the full procedure. ## How it grows The more samples you accumulate in `corpus/`, the richer the analysis. Tiers: - **300-1k words**: anchor blend + basic stats - **1k-5k words**: + preliminary NEVER rules + top fingerprints - **5k-20k words**: + full NEVER coverage + all fingerprints - **20k+ words**: AV-grade (high-confidence profile) ## Re-running analysis After adding samples to `corpus/`, tell the agent: > "Re-analyze my voice profile." The agent reads `catalog/*.md` and the corpus, then regenerates `stats.md`, `never-rules.md`, `fingerprints.md`, and `status.md`. No Node script — pure markdown skill, the agent is the extractor. ## Manual edits `never-rules.md` and `fingerprints.md` have a `## Manual Additions` / `## Manual Overrides` section at the bottom. The agent preserves anything you put there across regenerations. ## Privacy This is your voice profile. The files live on your disk. The skill never uploads anything — analysis is local (agent reasoning over local files). The only thing that leaves your machine is if you use the web tool at openwriter.io/writers-voice for the anchor step (300-800 words pasted in, used once, cached 24h, never trained on). If you use skill mode instead, even the anchor is generated locally.
-
-
.gitignore 415 B · in bundle
-
LICENSE 1 KB · in bundle
-
package.json 1017 B
{ "name": "authors-voice", "version": "0.19.1", "description": "Author's Voice — constructed-voice skill for AI agents. Anchors writing to a training-data author blend, progressively layers NEVER rules, presentation fingerprints, sentence stats, coined terms, and curated examples from a growing local corpus. Local-first markdown skill; optional paid API for plugin/programmatic flows. Replaces writers-voice + the legacy voice-* skill family.", "type": "module", "license": "MIT", "author": "travsteward", "homepage": "https://openwriter.io/authors-voice", "repository": { "type": "git", "url": "https://github.com/travsteward/authors-voice" }, "keywords": [ "claude", "claude-code", "skill", "voice", "writing", "ai", "openwriter", "authors-voice", "writers-voice", "anti-ai" ], "files": [ "SKILL.md", "catalog/", "docs/", "prompts/", "voice/README.md", "voice/corpus/.gitkeep", "LICENSE", "README.md" ] } -
README.md 7.8 KB
# authors-voice > **AI writing that sounds like you, not AI.** Most attempts to make AI sound like you start the same way. Train it. Fine-tune it. Feed it your samples and tell it to imitate. The model doesn't actually learn you from any of this. It pattern-matches at the lexical layer, lifting your common words and sentence shapes without ever building a deep representation of how you think. Cold-start imitation tops out shallow. Flip the direction. The model already carries deep internal representations of widely-published authors it was trained on at scale, voices it can channel with real fidelity because it saw thousands of pages of each. The move is to identify which of those authors a user statistically resembles, assign proportional weights to the closest matches, and instruct the model to write as that weighted blend. Your voice gets reconstructed as a coordinate inside the model's existing author space, anchored to authors it has already mastered. That changes the problem. The model isn't being asked to learn anything new about you. It's being asked to mix voices it knows cold, in proportions that triangulate your position among them. The anchor does the heavy lifting before a single sample of yours enters the prompt. The blend is the voice. The result is AI writing that sounds like you. Not AI imitating you. On top of the anchor, four layers sharpen the output. A list of AI words and constructions the model must never use, because the moment it stops channeling the anchor it reverts to its trained register and reaches for the same fifty tells. Presentation choices you make consistently, like whether you capitalize after a colon or use the Oxford comma, small mechanical preferences that read as authentic. A sentence-length and punctuation rhythm pulled from your own writing, so the cadence matches even when the diction is on loan. A growing folder of your samples that the skill mines as the negative rules and rhythm get re-derived. Each sample you add updates the NEVER rules and the sentence rhythm against your latest corpus. The anchor and presentation fingerprints don't auto-refresh. Regenerate those when you've added enough new writing to shift the matches, or when you want a fresh pass. The profile gets sharper the more you write and the more often you ask for a refresh. Roughly 80% of the way to your real voice. A hard jump above what stock prompting and fine-tuning produce. ## Install The skill is **agent-agnostic**. Pure markdown, no language runtime. Any LLM-based agent that can read `SKILL.md` and follow instructions can use it. ### Claude Code ```bash claude install github:travsteward/authors-voice ``` Clones to `~/.claude/skills/authors-voice/` and registers the skill with Claude Code. ### Vercel skills CLI (Claude Code, Codex, Cursor, and other agents) ```bash npx skills add travsteward/authors-voice ``` ### Manual (any agent) ```bash git clone https://github.com/travsteward/authors-voice ``` Then drop the cloned folder wherever your agent loads skills from. The `SKILL.md` at the root has the trigger phrases and routing logic the agent reads. ## Quick Start Two paths. Pick one. **Path A: Web tool first (fastest first-run)** 1. Visit [openwriter.io/voice-match](https://openwriter.io/voice-match), paste 300 to 800 words of your writing, copy the result block. 2. Tell your agent: *"set up my voice match"*. Paste the block when prompted. 3. **Seed your corpus**: paste 2 to 5 paragraphs that feel most like you. The agent saves them under `voice/corpus/`. 4. Done. **Path B: Skill mode (no web round-trip)** 1. Tell your agent: *"set up my voice match"* and *"I want to skip the web tool."* 2. Paste 2 to 5 paragraphs of your writing. The agent saves them under `voice/corpus/`. 3. The agent runs the Anchor Protocol over your corpus and writes `voice/anchor.md` directly. 4. Done. The skill is self-routing. You don't memorize subcommands. Just tell the agent what you want: - *"voice status"* → reports your current tier and word count - *"add this essay to my voice profile"* → appends, re-analyzes - *"write me a tweet about X"* → uses your voice automatically ## How It Works Your voice profile lives in `voice/` as a handful of `.md` files the agent reads at write time: | File | Source | Purpose | |------|--------|---------| | `anchor.md` | One-time match from [openwriter.io/voice-match](https://openwriter.io/voice-match) (or skill-mode). Refresh on demand. | 3 to 5 training-data authors with weights | | `stats.md` | Auto-regenerated from corpus on every analysis run | Sentence distribution + punctuation density | | `never-rules.md` | Auto-regenerated from corpus on every analysis run. Manual additions preserved. | AI words and phrases the model must never use | | `fingerprints.md` | Agent extracts from corpus during analysis runs. Manual overrides preserved. | Presentation choices (Oxford comma, capitalization after colon, contraction frequency) | | `coined-terms.md` | You curate | Your repeated coinages | | `examples.md` | You curate | Reference paragraphs in your voice | | `status.md` | Auto-regenerated on every analysis run | Current tier and what's locked next | Plus `voice/corpus/`. Your raw samples accumulating over time. None of `voice/*` is committed. It's all local to your disk. What updates reliably on every sample add: NEVER rules, sentence stats, status. What gets re-derived in protocol but agents sometimes skip: fingerprints (ask for a rebuild if you want certainty). What needs an explicit ask: anchor weights ("regenerate my anchor"). The corpus folder is yours to grow. ## Progressive Tiers The more samples you add, the more confident the analysis. | Words | Tier | Active | |-------|------|--------| | under 300 | 0 | seed corpus first | | 300 to 1k | 1 | anchor and basic stats | | 1k to 5k | 2 | preliminary NEVER rules and top fingerprints unlock | | 5k to 20k | 3 | full NEVER coverage and all fingerprints unlock | | 20k and up | 4 | high-confidence profile | ## Privacy - Your voice data lives entirely on your disk. `.gitignore` excludes everything in `voice/` from the public repo. - The skill never uploads your corpus anywhere. - The only thing that leaves your machine is the initial 300 to 800 word paste into openwriter.io/voice-match for the anchor matching step. That's cached 24h by hash and never trained on. ## Requirements - A Claude Code or compatible agent that supports skills (no Node.js dependency) - An initial visit to [openwriter.io/voice-match](https://openwriter.io/voice-match) for the anchor (free, no signup), or use skill-mode to build it locally ## Beyond the skill Pairs naturally with [OpenWriter](https://openwriter.io), the free AI writing surface. Same anchor system also powers the paid Author's Voice plugin (inline voice edits inside OpenWriter) and the paid API (programmatic voice-matched output for workflows and apps). See [authors-voice.com](https://authors-voice.com) when you outgrow the skill alone. ## License MIT. See [LICENSE](./LICENSE). ## Replaces This skill replaces the older `writers-voice` skill and the legacy `voice-apply`, `voice-generate`, `voice-setup`, `voice-upload`, `voice-manage`, and `voice-automate` skills. They are now one trigger: `/authors-voice`. ## History The local-skill half of `/authors-voice` started life as the standalone `writers-voice` skill. Its full development history (every iteration of the anchor protocol, NEVER rules, fingerprints, and tier logic) lives in the archived `travsteward/writers-voice` repo's git log. The repo is private now, but the commit log is preserved as the record of how the constructed-voice architecture evolved before it was unified here. ## Credits Built on the negative-first voice profiling architecture from [Author's Voice](https://authors-voice.com). Pairs with [OpenWriter](https://openwriter.io), the writing surface for AI agents. -
SKILL.md 14.1 KB
--- name: authors-voice description: | Author's Voice — constructed-voice skill. Anchors writing to a training-data author blend, progressively layers NEVER rules, presentation fingerprints, sentence stats, coined terms, and curated examples from a growing local corpus. Pure markdown, opus sub-agent (the minion) writes prose. Use when: "/authors-voice", "/writers-voice", "set up my voice", "anchor my voice", "voice match", "use my voice", "write in my voice", "add to my voice profile", "voice profile status", "voice status", "authors voice", "writer's voice". API path (plugin / programmatic flows): see `docs/api/` for the rewrite + generate endpoints, MCP tools, setup, and troubleshooting. The local skill body below is the default; the API is one access point among others. metadata: author: travsteward version: "0.19.1" license: MIT --- # Author's Voice _This skill is the free, any-agent manifestation of the larger Author's Voice ecosystem — the same anchor + NEVER-rules + anti-AI engine that powers the paid API, the OpenWriter plugin, and the dashboard. Free here, productized there; one voice DNA across all surfaces._ ## FIRM RULES ### 1. Editor NEVER writes prose. Every writing task fires a minion. The editor scopes briefs, cuts, reorders, and patches NEVER violations. Editor does NOT write new prose — every word ENTERING the document is minion-written. Applies to: initial drafts, revisions, bridges, closers, openers, transitions, single-line aphorisms, one-paragraph corrections, gap-fillers, idea-extensions — any new prose, full stop. Two carve-outs: (1) **violation patch** at Apply step 6 — 1-3 sentence local fix to NEVER violation or brief-error in otherwise-acceptable minion output; constructive rephrase preferred. (2) **co-write mode** — editor writes directly when ALL THREE hold: continuous real-time collaboration (user steering each move), explicit per-piece authorization for THIS piece, small scope (sentence to short paragraph; max one). Blanket authorizations ("you handle it") do NOT trigger co-write — those are delegations and go to the minion. Editor territory (no minion): cuts, reorders, accept/reject decorations, resolve agent marks, version restores. Cuts that leave holes needing new connective tissue — the connector is minion work. Rationale: editor context is polluted by the live conversation; voice anchors lose to active context; minion in clean context with voice files loaded outperforms editor-with-intent. If the editor catches itself drafting prose mid-conversation, **STOP** and spawn a minion — even for one sentence. ### 2. After anchor critique, discuss before revising. Anchor critique returns scores + convergent diagnostic. First move after aggregation: surface to user, discuss what to act on / disagree with / defer, align scope. Then spawn revision minions. Critics are advisory; user owns the prose. Anchor critique result protocol: 1. Aggregate panel scores + convergent diagnostic. 2. Surface to user: scores, themes raised by 3+ critics, proposed CUTS (subtraction) and REWRITES / ADDITIONS (new prose). 3. Discuss. User decides act / disagree / defer. No revision begins until this happens. 4. Editor executes agreed CUTS directly (Rule 1 carve-out). 5. Editor scopes briefs for agreed REWRITES / ADDITIONS, spawns revision minions per brief. 6. Patch micro-violations on returned prose. 7. Re-fire panel only if user wants another pass. ### 3. Revisions must be tighter than the original. If a revision is additive, the editor wrote it. Critique-driven revision produces a smaller word count, not larger. If post-revision is longer than pre-revision, the editor was inventing rather than executing the diagnostic. Revision minion brief specifies a WORD-COUNT TARGET (often "this paragraph in 60% of the original"). Minion compresses. Editor verifies the count dropped. ## Architecture Skeleton prompt template (`prompts/skeleton.md`) assembled from per-user `voice/*.md` files at write time. Editor loads skeleton, substitutes `{INCLUDE: ...}` markers, fills `{TASK}`, spawns a fresh opus sub-agent (the minion) with the assembled prompt (Claude Code) or dispatches via `task({ subagent_type: "general", prompt: <assembled skeleton> })` (OpenCode). Minion has no session pollution, returns prose, dies. Editor integrates. ``` writers-voice/ ├── SKILL.md (this file — router + firm rules + Apply Protocol) ├── docs/ (on-demand: setup, analysis, apply-deep, anchor-iteration, context-hygiene, tiers) ├── catalog/ (read-only reference: ai-tells, fingerprints, hurdle, anchor-prompt, author-hints, post-write-audit) ├── prompts/skeleton.md (template + injection points) └── voice/ (user-specific — anchor blend, NEVER rules, stats, fingerprints, coined-terms, examples, corpus/) ``` Lean / rich split: `voice/*-analysis.md` is human-facing only — never injected into the minion prompt. ## Routing | User intent | Action | | --- | --- | | "set up my voice" / "/writers-voice" / first run | **Setup Flow** — `docs/setup.md` | | "add this essay to my voice profile" / "save this writing" | Append to `voice/corpus/`, run **Analysis Protocol** — `docs/analysis.md` | | "voice status" / "what's locked" / "tier" | Read `voice/status.md`; tier reference at `docs/tiers.md` | | (User asks the agent to write anything) | Run **Apply Protocol** (below) after **Context Hygiene** check — `docs/context-hygiene.md` | | "show me my anchor" / "show me my fingerprints" | Read the relevant `voice/*.md` and report | | "regenerate my profile" / "re-analyze my corpus" | **Analysis Protocol** — `docs/analysis.md` | | "make a book / business / [context] voice anchor" | **Anchor Protocol** for that register — `docs/setup.md` | | "split my anchor by register" | **Multi-Register Split** — `docs/setup.md` | | Book-scale project (multi-chapter book) | Load `/book-writer` skill — that's the orchestration layer (chapter architecture, beats methodology, workspace management, book mode, long-form orchestration). This skill provides the Apply Protocol that `/book-writer` delegates to for every prose pass. | | "polish this" / "iterate to 90" / final-polish ready prose | **Anchor Iteration** — `docs/anchor-iteration.md` | | "use the API" / "call author's voice from a workflow" / plugin / programmatic | API path — `docs/api/protocol.md` (rewrite, generate, MCP tools, setup, troubleshooting) | ## Apply Protocol (Apply Minion — generative writing from commitments) Four minion types (full taxonomy: `docs/apply-protocol-deep.md`): **Apply** (generative, no source prose) · **Rewrite** (Apply + context awareness, updated commitments) · **Blinder Audit** (critic — paragraph-level substance duplication only) · **Anchor Iteration** (polish — channels voice anchors as panel, iterates critique → rewrite → re-score until 90/100). Pick the wrong minion → weak output. Substance problem → Rewrite, not Apply. Rough draft → not Anchor Iteration. When the user asks for a voice-matched write: 1. **Context Hygiene check.** Reset if polluted — `docs/context-hygiene.md`. 2. **Pick the anchor.** List `voice/anchor*.md`. If only `voice/anchor.md`, use it. If context-specific, infer from request or ask. Fallback: `voice/anchor.md`. 3. **Assemble the minion prompt.** Read `prompts/skeleton.md`. For each `{INCLUDE: <path>}`, substitute file contents. If a referenced file is missing (e.g., user hasn't curated examples), drop the entire `{INCLUDE: ...}` line AND its section header. Swap `voice/anchor.md` → `voice/anchor-<context>.md` if a context-specific anchor applies. 4. **Fill `{TASK}`.** **Required: COMMITMENTS** — what must be true of the output (concepts, claims, sequence, register, avoidances, length). **DEFAULT MODE: pure generation (regenerate).** SEMANTIC commitments only — what claims must land, not how to phrase them. No structural beats. No paragraph patterns. No device prescription. No sentence-rhythm prescription. The shape of any prescription becomes a ceiling on what the model produces. Let the minion bring its own moves. Always. ### Never write meta-references into commitments Anything the prose reader cannot see — chapter labels, beat numbers, "as discussed earlier", "the previous beat established" — does NOT belong in commitments. The minion reproduces them literally and breaks the fourth wall. Commitments describe what must be COMMUNICATED, never how the editor is thinking about structural position. If continuity from a prior section matters, capture it as the SUBSTANTIVE thread the new section must pick up (the content, not the structural pointer). ### COMMITMENTS function as quasi-verbatim instructions When the editor writes a commitment with literal phrasing in parentheses (*"Define sleep debt (lost sleep compounds like unpaid interest)"*), the model treats the parenthetical as exact phrasing to reproduce — every section gets that line. For multi-section work where phrasing should vary, write commitments abstractly (*"Define sleep debt"*) and let the model phrase. Use literal commitments only when a specific phrasing MUST land. ### Read prior integrated sections first (multi-section work) Read prior integrated sections before writing this one's brief. Cadence shapes already used, substantive threads to pick up, and heavy-use coined terms are only visible by reading what's on the page. Set the new section's commitments and cadence prescription against that context. ### Preservation scope (load-bearing call on a gradient) | Mode | Source prose in TASK? | Commitments shape | Use when | |---|---|---|---| | **Full preservation (rewrite mode)** | YES | "preserve every load-bearing claim; refine voice while preserving structure and phrasing" | source is already strong; author has specific phrasing that must land; polishing voice-applied work | | **Pure generation with selective lifts** | NO | OMIT source. Identify 1-3 specific moves worth keeping. Lift them pre-emptively (MUST-APPEAR-VERBATIM in brief) OR post-edit (patch in after minion returns) | source is mostly weak with 1-3 lines worth keeping. Pre-emptive when known in advance; post-edit when the strong move is only obvious after seeing the minion's output | | **Pure generation (regenerate)** | NO | SEMANTIC commitments only — what claims must land, not how to phrase them | source is "good but not great"; want dramatic improvement; existing shape would constrain output; high-stakes piece. The shape of the source becomes a ceiling on what the model produces | ### Cadence prescription (optional, recommended for high-stakes writing) Lifts voice fidelity ~0.5 points over baseline. Explicit rhythm scaffolding doesn't constrain content; it frees capacity by removing "what shape should this take?" overhead. Example: *"Para 1: open with 3 short declaratives, stack medium with concrete examples, close with one analytical long. Para 2: alternate short claim with longer explanatory, end with sharp short. Para 3: build with longer analytical, end with single-line aphoristic close."* **Vary cadence prescriptions across sections.** Same prescription per section produces document-scale rhythm repetition (every section opens with 3 shorts, closes with aphorism) — invisible at section scale, mechanical at document scale. Edge-case guidance (Rewrite Minion brief template, Blinder Audit brief shape, multi-section context-loading layers, writing minion taxonomy): `docs/apply-protocol-deep.md`. 5. **Spawn the minion.** Claude Code: `model: "opus"`, `subagent_type: "general-purpose"`. OpenCode: `subagent_type: "general"` with no model parameter (subagent inherits parent model; encourage using the session's strongest model). Both: `prompt: <assembled skeleton>`. 6. **Patch NEVER violations + brief-error meta-references.** Smallest local span. Constructive rephrase preferred (contrastive negation → direct statement; banned word → plain equivalent; meta-reference → substantive thread it pointed at). Don't regenerate; minion voice IS the result. Detail: `docs/apply-protocol-deep.md`. 7. **Post-write audit.** Read `catalog/post-write-audit.md` and apply distribution-level checks (opener repetition, sentence-initial "The", function-word over-use, sentence-length variance, lexical watch list). For each failing check, surgically rewrite the smallest local span — 5-10 light substitutions across a typical draft; heavier rewrites mean misuse. Load-bearing prose wins ties. 8. **Integrate via openwriter.** `write_to_pad` for edits, `populate_document` for new docs. 9. **Cross-section coherence review** (multi-section only). (a) Editor self-review — cut test: could you delete this paragraph and lose only redundancy? (b) **Mandatory Blinder Audit Minion** — fresh-context critic, paragraph-level substance duplication only. Most well-written beats produce zero findings. Brief template + exclusions: `docs/apply-protocol-deep.md`. 10. **Polish (optional, two patterns).** (a) **Parallel pick-best** — N (3-6) Apply minions in parallel, same brief; editor picks best whole, mixes variants, or hands all to user. (b) **Anchor Iteration** — `docs/anchor-iteration.md`. Polish-class only; not for rough drafts. 11. **`/anti-ai` pass.** MANDATORY after Anchor Iteration (which runs no-context and introduces AI tells). OPTIONAL otherwise. Global surface fingerprints (em-dashes, semicolons, contrastive negation, banned diction, register monotony) vs `voice/never-rules.md` + `voice/fingerprints.md`. Complements step 7. **Use opus (Claude Code).** Sonnet leaks 3+ NEVER violations where opus leaks 0-1. Haiku loses voice. **OpenCode:** subagents inherit the parent model — use the strongest model available in the session for prose generation. **Send full editing scope.** If 6 of 8 paragraphs need fixes, send all 8 for flow continuity. **One minion per natural editing unit** — beat, section, blog post, tweet thread. ## Tiers + Companion Skills Voice profile tiers (Empty / Anchor / Preliminary / Full Coverage / AV-Grade) gate features by corpus word count. Table: `docs/tiers.md`. Companions: `/anti-ai` (final fingerprint scrub) · `/voice-presets` (generic frames, no profile — if installed) · Author's Voice plugin (paid — full RAG, inline edits, deterministic extraction).
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.