Claude Cursor Skill

thesis-control

Use when AI-assisted thesis or manuscript edits risk claim drift, scope creep, loss of intended use, experiment-role promotion, or repeated revisions that fail to converge; provides author-intent control, lightweight or strict contracts, drift audits, revision escalation, and hum

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download yha9806-academic-writing-toolkit-archive_skills_thesis-control-184e482.zip · 35 KB
Part of yha9806/academic-writing-toolkit — 21 skills

Install

skills CLI npx skills add https://github.com/yha9806/academic-writing-toolkit/tree/main/archive/skills/thesis-control
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install yha9806-academic-writing-toolkit@llmmart
Git git clone https://github.com/yha9806/academic-writing-toolkit.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole yha9806/academic-writing-toolkit collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

/thesis-control - Thesis Drift Control

Purpose

Prevent AI-assisted writing from becoming fluent but distorted. Use this before and after substantive thesis or manuscript edits when the risk is not spelling or style, but loss of author control: project-level reframing, deletion of the real-world task or intended use, a primary domain becoming a secondary example, a changed research object or question, auxiliary analyses becoming primary, widened claims, blurred section purpose, missing caveats, unsynchronised adjacent paragraphs, or local edits that weaken the paper spine.

Trigger Words

This skill activates on: thesis control, drift audit, edit contract, spine card, claim drift, author control, loss of control, scope creep, rewrite risk, /thesis-control.

Core Rule

Do not edit thesis prose until the project intent, current manuscript contract, global thesis audit, section spine, and intended local change form one explicit and traceable contract chain.

The contract must answer:

This edit is allowed to change [specific local issue] in [specific unit], while preserving [spine sentence], [scope boundary], [core claims], and [do-not-change items].

If this sentence cannot be written, stop and diagnose the section instead of rewriting it.

The control hierarchy is:

Author-approved Project Intent
→ Author-approved Manuscript Contract
→ Passed Global Thesis Audit
→ Section Spine Card
→ Edit Contract
→ Post-edit Drift Audit

A lower layer cannot amend a higher one. If the title, abstract, primary domain, research object, research question, contribution scope, or manuscript structure no longer matches the approved intent, stop. Revise or roll back the manuscript, or create a new explicitly approved intent version that preserves the earlier row as history.

Control Files

Choose one control profile and name it as canonical.

For a lightweight single-manuscript workflow, read references/author_control_lightweight.md and use:

  • 00_AUTHOR_INTENT.md
  • 01_EVIDENCE_AND_CLAIMS.md
  • 02_REVISION_LOG.md

Create and check the bundled templates with:

python scripts/scaffold-author-control.py <project_root>
python scripts/check-author-control.py <project_root> --strict

The lightweight checker validates structure, approval state, and unresolved placeholders. It does not infer semantic alignment. Do not use the lightweight profile to bypass a gate that already requires the durable packet.

Use a thesis_control/ directory when the project needs durable tracking:

  • project_intent.csv
  • manuscript_contracts.csv
  • global_thesis_audits.csv
  • spine_cards.csv
  • edit_contracts.csv
  • drift_audits.csv
  • revision_escalations.csv

Run the optional validator when Python is available:

python {skill_dir}/scripts/check_thesis_control.py <project_root> --strict

Strict validation requires the project-intent layer. It blocks approved or applied edit contracts unless they reference a passed global thesis audit for the active author-approved intent and manuscript contract. Non-strict mode can still inspect legacy packets that do not yet have this layer.

The validator checks packet structure and recorded gate consistency. It does not infer semantic alignment or judge scholarly truth. The author or reviewer must compare the manuscript with the intent and record each alignment field honestly; the validator then prevents an unresolved or drifted audit from being used as authorisation.

Strict validation requires revision-tracking schema v3. Upgrade a complete legacy packet without guessing historical revision families:

python {skill_dir}/scripts/upgrade_thesis_control_revision_tracking.py <project_root>

To create a draft packet from a real Markdown unit before editing prose:

python {skill_dir}/scripts/scaffold_thesis_control.py <project_root> \
  --source chapters/ch1_introduction.md \
  --start-line 71 \
  --end-line 104 \
  --revision-issue-id ri-ch1-gap-clarity \
  --attempt-no 1 \
  --copy-source

The scaffold writes schema v4 draft project-intent and manuscript contracts, a pending global thesis audit, human_approved=false, status=draft, and AUTHOR_REVIEW_REQUIRED fields. Replace those fields with concrete author judgement before applying a substantive edit. A scaffolded packet may be structurally valid while remaining non-executable. Its default contract id includes the attempt number, so attempts 1 and 2 become ec-<unit>-001 and ec-<unit>-002. Reuse an explicit revision_issue_id for retries.

The migration helper stops without writing when revision metadata is partial or when a legacy escalation cannot be classified from current contracts and resolved audits. It preserves named extension columns and converts one- or two-trigger legacy rows to early_diagnostic; a three-trigger row becomes a cycle_gate only when it already matches one completed failure group.

Upgrade a complete schema-v3 packet into a deliberately blocked schema-v4 draft without guessing author intent:

python {skill_dir}/scripts/upgrade_thesis_control_project_intent.py \
  <project_root> --json

The helper adds manuscript_id and global_audit_id links, preserves named extension columns, and creates AUTHOR_REVIEW_REQUIRED draft intent, manuscript, and global-audit rows through one atomic batch. Previously approved or applied edits remain blocked. Replace the draft fields with real author judgement, approve the active intent and manuscript contract, and resolve the global audit before strict validation can pass. A partial project-intent schema stops without mutation.

Workflow

0. Establish The Project Intent And Manuscript Contract

Before section-level planning, record:

  • the real-world problem and intended user or beneficiary
  • the intended application and the present method or software task
  • the primary scholarly domain
  • the research object
  • the core research question
  • the primary experiment that directly answers that question
  • supporting, robustness, exploratory, failed-development, and out-of-scope analyses
  • the strongest evidence-licensed headline claim
  • the current validation and evidence boundaries
  • the target venue or audience
  • concepts that must remain visible in the title or abstract
  • reframes that require fresh author approval
  • the current title, abstract focus, contribution scope, and structure
  • concrete approval evidence and the active version ids

Keep one active author-approved project intent and one active author-approved manuscript contract. A later intent version must identify the immediately previous version in supersedes_intent_id, record the amendment reason, and leave the earlier version as superseded. Do not overwrite the original row.

Run a global thesis audit whenever the title, abstract, primary domain, research object, research question, contribution scope, or overall structure changes. Record each dimension as aligned, drifted, or not_assessed. Only a fully aligned audit with detected_reframe=false can have status=passed and human_decision=accept.

If any dimension is drifted, set human_review_required=true and use needs_review or failed. The author must choose to revise the manuscript, roll back, or approve a versioned intent amendment. Merely accepting the audit cannot authorise the reframe.

Keep intended use and current validation separate. Narrow evidence may narrow the empirical or headline claim, but it must not silently delete a legitimate application problem or recast the evidence boundary as the paper's research object. Record future application as intended use and untested hardware, clinical, causal, deployment, or transfer outcomes as unvalidated boundaries.

Before admitting a completed analysis into the paper, record whether it directly answers the core question, whether the main conclusion survives its removal, its one-sentence argumentative function, its role, destination, and author decision. Completion alone does not make an analysis a main contribution. Use the role and placement defaults in references/author_control_lightweight.md.

1. Establish Or Read The Spine Card

Before editing a chapter, section, or paragraph cluster, identify:

  • unit id
  • source path
  • section title
  • spine sentence
  • scope boundary
  • core claims
  • do-not-change items
  • the active manuscript contract id

The spine sentence should be narrow:

This unit argues that [specific claim] by showing [specific basis], so the chapter can [specific function].

If the current text does not support a clear spine sentence, produce a diagnosis and ask for author direction before changing prose.

2. Create The Edit Contract

For every substantive edit, state:

  • target unit and file range
  • change scope: local_patch, section_restructure, or full_reframe
  • allowed changes
  • forbidden changes
  • evidence baseline: the ref, artifact, data, configuration, table, or frozen numbers inherited
  • argument baseline: the author-approved intent and manuscript version inherited
  • adjacent context that must be checked
  • acceptance checks
  • whether human approval is required before editing
  • the passed global thesis audit id that covers the spine card's manuscript contract

When using the scaffold helper, treat its output as a draft control packet, not as approval. A generated contract becomes actionable only after the author has replaced the AUTHOR_REVIEW_REQUIRED fields and explicitly approved the scope.

Always require human approval for:

  • changing the section spine
  • adding or broadening claims
  • deleting caveats or limitations
  • moving evidence between sections
  • rewriting more than one paragraph
  • merging or splitting sections
  • changing the title, abstract, primary domain, research object, research question, contribution scope, or manuscript structure
  • changing the real-world problem, intended use, primary experiment, analysis prominence, headline claim, or evidence boundary

Treat any change to the title, abstract thesis, research object, core question, primary experiment, contribution order, evidence chain, application purpose, or paper-wide structure as a full_reframe, even when the request calls it polishing. Before editing, show an old-versus-proposed spine comparison and obtain explicit author approval.

3. Apply Only Approved Changes

After approval, edit only the approved scope. Do not apply the edit if the linked global thesis audit is pending, failed, stale, drifted, or attached to a different manuscript contract.

Keep mechanical fixes separate from argument changes. Do not bundle style, structure, evidence, and claim changes into one patch unless the contract explicitly allows it.

4. Run The Drift Audit

After editing, compare the new prose against the contract and report:

  • changed claims
  • changed boundaries or caveats
  • new unsupported claims
  • deleted evidence anchors
  • missed adjacent updates
  • section-spine change
  • research-object or core-question change
  • deleted, generalised, or demoted real-world task or intended use
  • promoted auxiliary analysis
  • evidence boundary rewritten as the paper topic
  • loss of application meaning caused by over-cautious wording
  • title, abstract, Introduction, Results, and Conclusion alignment
  • decision: accept, partial accept, revise, or rollback

If any claim, boundary, or caveat changed, the result needs human review even if the prose is smoother.

Use audit status=needs_review only while the author's post-edit decision is pending. Strict validation blocks an applied contract in that state. After the author decides, record status=passed for accept or partial_accept, and status=failed for revise or rollback. Do not treat a pending audit as a completed unsuccessful attempt.

5. Record Human Gate Outcome

The author decides whether to accept, partially accept, revise, or rollback. Do not mark a high-risk edit as accepted without explicit human approval.

6. Run Post-Spine Readability Gates When Relevant

Only after the research spine is stable, use /logic-review to audit repeated argument functions and /self-review to prepare the unfamiliar-reader packet. Do not solve repetitive AI prose by generating synonyms. Remove duplicated problem, gap, evidence, interpretation, or boundary functions while preserving essential local qualifiers. A model simulation cannot pass a human unfamiliar- reader gate; record it as advisory or not_run until an actual reader responds.

Revision Escalation Rule

Treat three unsuccessful attempts on the same revision issue as an operational escalation threshold, not as evidence that every task fails after three turns. Use revision_issue_id to keep successive contract versions attached to that issue. Only count an attempt when its drift decision is revise or rollback and its audit status is failed. Only applied contracts count as unsuccessful attempts. Record author rejection as one of those decisions. Multiple failed audits of one contract still count as one attempt; contradictory passed and failed resolved audits are invalid. Clarifying discussion, pending human reviews, and unexecuted proposals do not count.

After three unsuccessful attempts, stop. Do not apply a fourth prose patch. Record a row in revision_escalations.csv; a later contract may become approved or applied only after the matching escalation has human_approved=true and status=approved.

An approved escalation closes only that group of three unsuccessful contracts. If three later contracts also receive revise or rollback, require a new escalation before another contract can proceed.

Only a cycle_gate whose three triggers exactly match one completed group of unsuccessful contracts, in attempt order, may close that group. Set approved_after_attempt to the final attempt number in that group. The gate is effective only with human_approved=true and status=approved. One gate cannot close more than one group. Do not repeat a trigger contract within a row or create multiple rows for the same issue and trigger set.

Only an escalation whose trigger set exactly matches one completed group of three unsuccessful contracts may close that group. One escalation cannot close more than one group.

Record an earlier warning as early_diagnostic with one or two unique triggers and an empty approved_after_attempt. It may be author-approved as a diagnosis, but it never closes or pre-authorises a later completed group.

An earlier escalation with fewer than three trigger contracts does not close or pre-authorise a later completed group.

Escalate earlier than three attempts when any of these signals is already visible:

  • the section spine cannot be stated consistently
  • the requested claim lacks supporting evidence
  • a revision changes a claim, caveat, or scope boundary outside the contract
  • the latest author-approved version cannot be identified
  • old assumptions, duplicated explanations, or conflicting requirements indicate version contamination

Required Escalation Check

Before editing again:

  1. Consolidate the currently valid requirements into one brief.
  2. Compare that brief with the spine card, evidence boundaries, current contract, and latest author-approved version.
  3. Classify the failure as one primary category:
    • underspecified or conflicting intent — the target, audience, venue, constraint, or acceptance condition is missing or inconsistent, or the feedback is evaluative but not operational, such as “weak”, “unclear”, or “still not right” without a concrete change target
    • local execution failure — the contract is clear, but the edit did not implement it correctly
    • structural mismatch — the problem affects the section purpose, research question, gap, contribution, evidence chain, or manuscript structure
    • evidence gap — the requested claim is not supported by the available sources, data, experiments, or files
    • version contamination — accumulated patches mix incompatible assumptions, duplicate reasoning, or obscure which prose the author approved
  4. Classify the writing scope:
    • local patch — wording or presentation changes that preserve the spine, claims, evidence, and adjacent-section relationships
    • section-level restructure — changes confined to one section without changing the research question, contribution, or evidence chain
    • full reframing — changes to the title, abstract, research question, gap, contribution, methods-results alignment, evidence chain, or discussion framing
  5. Recommend the smallest valid next action and wait for author approval.

Use these default actions:

  • For a local execution failure, create a corrected local contract.
  • For underspecified or conflicting intent, ask for the missing decision before editing.
  • For a structural mismatch, propose a section-level restructure or full reframing plan before editing.
  • For an evidence gap, narrow, qualify, or remove the unsupported claim unless the author supplies more evidence.
  • For version contamination, restore or copy the latest author-approved version, then apply a consolidated contract. Create a separate branch or manuscript version only when the approved scope requires structural work.

For full reframing, hand off a brief that states the target venue, old and proposed real-world problem, intended use, research object, research question, primary experiment, contribution order, headline claim, evidence boundary, evidence baseline, argument baseline, available evidence, claims that must not be made, and proposed new structure. Do not rewrite the manuscript until the author approves that brief.

Return the escalation check in this form:

## Revision Escalation Check

Revision issue:
Contract:
Unsuccessful attempts:
Trigger contracts:
Primary category:
Writing scope:
Why the revisions did not converge:
Valid requirements:
Missing or conflicting information:
Latest author-approved version:
Recommended next action:
Author decision required:

Output Patterns

Audit Only

Return:

  • current spine diagnosis
  • likely drift risks
  • control gaps
  • recommended edit contracts
  • blocked items needing author decision

Pre-Edit Contract

Return:

## Edit Contract

Unit:
Spine sentence:
Scope: local_patch / section_restructure / full_reframe
Evidence baseline:
Argument baseline:
Allowed changes:
Forbidden changes:
Adjacent context to check:
Acceptance checks:
Human approval required:
Proceed only after approval:

Post-Edit Drift Audit

Return:

## Drift Audit

Contract:
Changed claims:
Changed boundaries:
New unsupported claims:
Missed adjacent updates:
Research object or question changed:
Intended use deleted or demoted:
Auxiliary analysis promoted:
Cross-section paper identity aligned:
Decision:
Human review required:
Recommended next action:

Stop Conditions

Stop and ask for author direction if:

  • the section spine cannot be stated clearly
  • the requested edit would broaden a claim without evidence
  • a local edit requires adjacent updates outside the approved scope
  • the user asks for a full-chapter rewrite without a spine map
  • previous AI edits cannot be distinguished from author-approved text
  • the edit would remove caveats, limitations, or uncertainty language without explicit approval
  • the active project intent or manuscript contract cannot be identified
  • the real-world task, intended use, primary experiment, evidence baseline, or argument baseline cannot be identified
  • a global thesis audit is missing, unresolved, stale, or records project-level drift
  • a proposed local contract would preserve a section spine that conflicts with the author-approved project intent
  • a full reframe lacks an approved old-versus-proposed spine comparison
Files (academic-writing-toolkit)
  • scripts
    • check_thesis_control.py 38.9 KB
      #!/usr/bin/env python3
      """Validate thesis-control packets.
      
      The checker is intentionally structural. It verifies that a project has
      spine cards, edit contracts, drift audits, revision-escalation records, and
      human-gate consistency for AI-assisted thesis edits. It does not judge whether
      the academic argument is true or well written.
      """
      
      from __future__ import annotations
      
      import argparse
      import json
      import re
      import sys
      from pathlib import Path
      from typing import Dict, Iterable, List, Mapping, Optional, Sequence, Set, Tuple
      
      from project_intent_control import validate_project_intent_layer
      from thesis_control_io import CsvShapeError, read_csv_table
      
      
      REQUIRED_FILES = {
          "spine_cards.csv": [
              "unit_id",
              "path",
              "section_title",
              "spine_sentence",
              "scope_boundary",
              "core_claims",
              "do_not_change",
          ],
          "edit_contracts.csv": [
              "contract_id",
              "unit_id",
              "change_scope",
              "allowed_changes",
              "forbidden_changes",
              "adjacent_context",
              "acceptance_checks",
              "human_approved",
              "status",
          ],
          "drift_audits.csv": [
              "audit_id",
              "contract_id",
              "changed_claims",
              "changed_boundaries",
              "new_unsupported_claims",
              "missed_adjacent_updates",
              "drift_decision",
              "human_review_required",
              "status",
          ],
      }
      
      REVISION_CONTRACT_COLUMNS = ["revision_issue_id", "attempt_no"]
      REVISION_ESCALATION_V2_COLUMNS = [
          "escalation_id",
          "revision_issue_id",
          "trigger_contracts",
          "primary_category",
          "writing_scope",
          "valid_requirements",
          "missing_or_conflicting_information",
          "latest_author_approved_version",
          "recommended_next_action",
          "human_approved",
          "status",
      ]
      REVISION_ESCALATION_V3_COLUMNS = [
          "escalation_id",
          "revision_issue_id",
          "escalation_kind",
          "trigger_contracts",
          "approved_after_attempt",
          "primary_category",
          "writing_scope",
          "valid_requirements",
          "missing_or_conflicting_information",
          "latest_author_approved_version",
          "recommended_next_action",
          "human_approved",
          "status",
      ]
      
      CONTRACT_STATUSES = {"draft", "approved", "applied", "rejected"}
      DRIFT_DECISIONS = {"accept", "partial_accept", "revise", "rollback"}
      AUDIT_STATUSES = {"passed", "needs_review", "failed"}
      ESCALATION_CATEGORIES = {
          "underspecified_or_conflicting_intent",
          "local_execution_failure",
          "structural_mismatch",
          "evidence_gap",
          "version_contamination",
      }
      WRITING_SCOPES = {"local_patch", "section_level_restructure", "full_reframing"}
      ESCALATION_STATUSES = {"draft", "approved", "rejected"}
      ESCALATION_KINDS = {"early_diagnostic", "cycle_gate"}
      EMPTY_MARKERS = {"", "none", "n/a", "na", "no", "not applicable"}
      IDENTIFIER_RE = re.compile(r"^[A-Za-z0-9](?:[A-Za-z0-9._-]{0,118}[A-Za-z0-9])?$")
      
      
      def is_empty(value: str) -> bool:
          return value.strip().lower() in EMPTY_MARKERS
      
      
      def parse_bool(value: str) -> bool | None:
          lowered = value.strip().lower()
          if lowered in {"true", "yes", "y", "1"}:
              return True
          if lowered in {"false", "no", "n", "0"}:
              return False
          return None
      
      
      def is_valid_identifier(value: str) -> bool:
          stripped = value.strip()
          return bool(IDENTIFIER_RE.fullmatch(stripped)) and ".." not in stripped
      
      
      def validate_source_path(
          root: Path,
          value: str,
          prospective_files: Optional[Set[Path]] = None,
      ) -> str | None:
          path_text = value.strip()
          if not path_text:
              return None
          source_path = Path(path_text)
          if source_path.is_absolute():
              return "source path must be relative to the packet root"
          if ".." in source_path.parts:
              return "source path must not contain '..'"
          candidate = (root / source_path).resolve()
          try:
              candidate.relative_to(root.resolve())
          except ValueError:
              return "source path must stay inside the packet root"
          if candidate not in (prospective_files or set()) and not candidate.is_file():
              return "source path does not exist"
          return None
      
      
      def read_fieldnames(path: Path) -> List[str]:
          if not path.is_file():
              return []
          try:
              fieldnames, _ = read_csv_table(path)
              return fieldnames
          except (CsvShapeError, OSError):
              return []
      
      
      def validate_columns(
          root: Path,
          filename: str,
          required_columns: Iterable[str],
          allow_empty: bool = False,
          table: Optional[Tuple[Sequence[str], Sequence[Mapping[str, str]]]] = None,
      ) -> Tuple[List[Dict[str, str]], List[dict]]:
          path = root / "thesis_control" / filename
          issues: List[dict] = []
          if table is None:
              if not path.is_file():
                  return [], [{"kind": "missing-file", "location": str(path), "message": f"missing {filename}"}]
      
              try:
                  fieldnames, rows = read_csv_table(path)
              except CsvShapeError as exc:
                  issues.append({"kind": exc.kind, "location": exc.location, "message": exc.message})
                  return [], issues
              except OSError as exc:
                  issues.append({"kind": "csv-error", "location": str(path), "message": f"cannot read file: {exc}"})
                  return [], issues
          else:
              fieldnames = list(table[0])
              rows = [dict(row) for row in table[1]]
      
          missing = [column for column in required_columns if column not in fieldnames]
          for column in missing:
              issues.append(
                  {
                      "kind": "missing-column",
                      "location": str(path),
                      "message": f"{filename} missing column: {column}",
                  }
              )
      
          if not rows and not allow_empty:
              issues.append({"kind": "empty-file", "location": str(path), "message": f"{filename} has no data rows"})
      
          return rows, issues
      
      
      def row_location(filename: str, index: int) -> str:
          return f"thesis_control/{filename}:row {index + 2}"
      
      
      def add_issue(issues: List[dict], kind: str, location: str, message: str) -> None:
          issues.append({"kind": kind, "location": location, "message": message})
      
      
      def validate_packet(
          root: Path,
          strict: bool = False,
          require_project_intent: Optional[bool] = None,
          table_overrides: Optional[
              Mapping[str, Tuple[Sequence[str], Sequence[Mapping[str, str]]]]
          ] = None,
          prospective_files: Optional[Iterable[Path]] = None,
      ) -> dict:
          issues: List[dict] = []
          overrides = table_overrides or {}
          prospective = {Path(path).resolve() for path in (prospective_files or [])}
          control_dir = root / "thesis_control"
          contract_path = control_dir / "edit_contracts.csv"
          escalation_path = control_dir / "revision_escalations.csv"
          contract_fields = (
              list(overrides["edit_contracts.csv"][0])
              if "edit_contracts.csv" in overrides
              else read_fieldnames(contract_path)
          )
          escalation_fields = (
              list(overrides["revision_escalations.csv"][0])
              if "revision_escalations.csv" in overrides
              else read_fieldnames(escalation_path)
          )
          escalation_v3 = all(
              column in escalation_fields for column in ["escalation_kind", "approved_after_attempt"]
          )
          enforce_revision_gate = strict or escalation_v3
          revision_tracking = strict or escalation_path.is_file() or all(
              column in contract_fields for column in REVISION_CONTRACT_COLUMNS
          )
      
          intent_state, intent_issues = validate_project_intent_layer(
              root,
              required=strict if require_project_intent is None else require_project_intent,
              table_overrides=overrides,
          )
          intent_tracking = intent_state["enabled"]
      
          spine_columns = list(REQUIRED_FILES["spine_cards.csv"])
          if intent_tracking:
              spine_columns.append("manuscript_id")
          spine_rows, spine_issues = validate_columns(
              root,
              "spine_cards.csv",
              spine_columns,
              table=overrides.get("spine_cards.csv"),
          )
          contract_columns = list(REQUIRED_FILES["edit_contracts.csv"])
          if revision_tracking:
              contract_columns.extend(REVISION_CONTRACT_COLUMNS)
          if intent_tracking:
              contract_columns.append("global_audit_id")
          contract_rows, contract_issues = validate_columns(
              root,
              "edit_contracts.csv",
              contract_columns,
              table=overrides.get("edit_contracts.csv"),
          )
          audit_rows, audit_issues = validate_columns(
              root,
              "drift_audits.csv",
              REQUIRED_FILES["drift_audits.csv"],
              allow_empty=True,
              table=overrides.get("drift_audits.csv"),
          )
          escalation_rows: List[Dict[str, str]] = []
          escalation_issues: List[dict] = []
          if revision_tracking:
              escalation_columns = (
                  REVISION_ESCALATION_V3_COLUMNS
                  if enforce_revision_gate
                  else REVISION_ESCALATION_V2_COLUMNS
              )
              escalation_rows, escalation_issues = validate_columns(
                  root,
                  "revision_escalations.csv",
                  escalation_columns,
                  allow_empty=True,
                  table=overrides.get("revision_escalations.csv"),
              )
          issues.extend(
              intent_issues
              + spine_issues
              + contract_issues
              + audit_issues
              + escalation_issues
          )
      
          spine_ids = set()
          spine_manuscripts: Dict[str, str] = {}
          for index, row in enumerate(spine_rows):
              location = row_location("spine_cards.csv", index)
              unit_id = row.get("unit_id", "").strip()
              if not unit_id:
                  add_issue(issues, "missing-unit-id", location, "spine card has no unit_id")
              elif not is_valid_identifier(unit_id):
                  add_issue(
                      issues,
                      "invalid-unit-id",
                      location,
                      "unit_id must be 1-120 safe ASCII characters and must not contain path segments",
                  )
              elif unit_id in spine_ids:
                  add_issue(issues, "duplicate-unit-id", location, f"duplicate unit_id: {unit_id}")
              else:
                  spine_ids.add(unit_id)
      
              for column in ["path", "section_title", "spine_sentence", "scope_boundary", "core_claims", "do_not_change"]:
                  if is_empty(row.get(column, "")):
                      add_issue(issues, "empty-spine-field", location, f"spine card field is empty: {column}")
              path_issue = validate_source_path(
                  root,
                  row.get("path", ""),
                  prospective_files=prospective,
              )
              if path_issue:
                  add_issue(issues, "invalid-source-path", location, path_issue)
              if intent_tracking:
                  manuscript_id = row.get("manuscript_id", "").strip()
                  if not manuscript_id:
                      add_issue(
                          issues,
                          "missing-manuscript-id",
                          location,
                          "spine card has no manuscript_id",
                      )
                  elif not is_valid_identifier(manuscript_id):
                      add_issue(
                          issues,
                          "invalid-manuscript-id",
                          location,
                          "manuscript_id must be a safe identifier",
                      )
                  elif manuscript_id not in intent_state["manuscripts_by_id"]:
                      add_issue(
                          issues,
                          "unknown-manuscript-id",
                          location,
                          f"spine card references unknown manuscript_id: {manuscript_id}",
                      )
                  if unit_id and manuscript_id:
                      spine_manuscripts[unit_id] = manuscript_id
      
          contract_ids = set()
          applied_contracts = set()
          contract_issues_by_id: Dict[str, str] = {}
          contract_attempts: Dict[str, int] = {}
          contract_statuses: Dict[str, str] = {}
          contract_locations: Dict[str, str] = {}
          issue_attempts: Dict[str, Dict[int, str]] = {}
          for index, row in enumerate(contract_rows):
              location = row_location("edit_contracts.csv", index)
              contract_id = row.get("contract_id", "").strip()
              unit_id = row.get("unit_id", "").strip()
              status = row.get("status", "").strip().lower()
              approved = parse_bool(row.get("human_approved", ""))
              revision_issue_id = row.get("revision_issue_id", "").strip()
              attempt_text = row.get("attempt_no", "").strip()
              global_audit_id = row.get("global_audit_id", "").strip()
      
              if not contract_id:
                  add_issue(issues, "missing-contract-id", location, "edit contract has no contract_id")
              elif not is_valid_identifier(contract_id):
                  add_issue(
                      issues,
                      "invalid-contract-id",
                      location,
                      "contract_id must be 1-120 safe ASCII characters and must not contain path segments",
                  )
              elif contract_id in contract_ids:
                  add_issue(issues, "duplicate-contract-id", location, f"duplicate contract_id: {contract_id}")
              else:
                  contract_ids.add(contract_id)
      
              if not unit_id:
                  add_issue(issues, "missing-unit-id", location, "edit contract has no unit_id")
              elif not is_valid_identifier(unit_id):
                  add_issue(
                      issues,
                      "invalid-unit-id",
                      location,
                      "unit_id must be 1-120 safe ASCII characters and must not contain path segments",
                  )
              elif unit_id not in spine_ids:
                  add_issue(issues, "unknown-unit-id", location, f"contract references unknown unit_id: {unit_id}")
      
              if status not in CONTRACT_STATUSES:
                  add_issue(issues, "invalid-contract-status", location, f"invalid contract status: {status}")
      
              if approved is None:
                  add_issue(issues, "invalid-human-approved", location, "human_approved must be true or false")
              if status in {"approved", "applied"} and approved is not True:
                  add_issue(issues, "missing-human-approval", location, "approved/applied contracts require human_approved=true")
      
              if intent_tracking:
                  if not global_audit_id:
                      add_issue(
                          issues,
                          "missing-global-audit-id",
                          location,
                          "edit contract has no global_audit_id",
                      )
                  elif not is_valid_identifier(global_audit_id):
                      add_issue(
                          issues,
                          "invalid-global-audit-id",
                          location,
                          "global_audit_id must be a safe identifier",
                      )
                  elif global_audit_id not in intent_state["audits_by_id"]:
                      add_issue(
                          issues,
                          "unknown-global-audit-id",
                          location,
                          f"edit contract references unknown global_audit_id: {global_audit_id}",
                      )
                  else:
                      audit_record = intent_state["audits_by_id"][global_audit_id]
                      spine_manuscript = spine_manuscripts.get(unit_id)
                      if spine_manuscript and audit_record["manuscript_id"] != spine_manuscript:
                          add_issue(
                              issues,
                              "contract-global-audit-mismatch",
                              location,
                              "edit contract global audit does not cover the spine card's manuscript contract",
                          )
                      if (
                          status in {"approved", "applied"}
                          and global_audit_id not in intent_state["ready_audit_ids"]
                      ):
                          add_issue(
                              issues,
                              "global-thesis-gate-required",
                              location,
                              "approved/applied edits require a passed global thesis audit for the active project intent and manuscript contract",
                          )
      
              for column in ["change_scope", "allowed_changes", "forbidden_changes", "adjacent_context", "acceptance_checks"]:
                  if is_empty(row.get(column, "")):
                      add_issue(issues, "empty-contract-field", location, f"edit contract field is empty: {column}")
      
              if status == "applied" and contract_id:
                  applied_contracts.add(contract_id)
      
              if contract_id:
                  contract_statuses[contract_id] = status
                  contract_locations[contract_id] = location
      
              if revision_tracking:
                  if not revision_issue_id:
                      add_issue(issues, "missing-revision-issue-id", location, "edit contract has no revision_issue_id")
                  elif not is_valid_identifier(revision_issue_id):
                      add_issue(
                          issues,
                          "invalid-revision-issue-id",
                          location,
                          "revision_issue_id must be a safe identifier",
                      )
      
                  attempt_no = None
                  try:
                      attempt_no = int(attempt_text)
                  except ValueError:
                      pass
                  if attempt_no is None or attempt_no < 1:
                      add_issue(issues, "invalid-attempt-no", location, "attempt_no must be a positive integer")
                  elif revision_issue_id and is_valid_identifier(revision_issue_id):
                      attempts = issue_attempts.setdefault(revision_issue_id, {})
                      if attempt_no in attempts:
                          add_issue(
                              issues,
                              "duplicate-revision-attempt",
                              location,
                              f"revision issue {revision_issue_id} repeats attempt_no {attempt_no}",
                          )
                      else:
                          attempts[attempt_no] = contract_id
                      if contract_id:
                          contract_issues_by_id[contract_id] = revision_issue_id
                          contract_attempts[contract_id] = attempt_no
      
          if revision_tracking:
              for revision_issue_id, attempts in issue_attempts.items():
                  numbers = sorted(attempts)
                  if numbers and numbers != list(range(1, numbers[-1] + 1)):
                      add_issue(
                          issues,
                          "nonsequential-revision-attempt",
                          "thesis_control/edit_contracts.csv",
                          f"revision issue {revision_issue_id} attempt numbers must be sequential from 1",
                      )
      
          audit_ids = set()
          audited_contracts = set()
          unsuccessful_contracts = set()
          resolved_audit_outcomes: Dict[str, Set[str]] = {}
          for index, row in enumerate(audit_rows):
              location = row_location("drift_audits.csv", index)
              audit_id = row.get("audit_id", "").strip()
              contract_id = row.get("contract_id", "").strip()
              decision = row.get("drift_decision", "").strip().lower()
              status = row.get("status", "").strip().lower()
              review_required = parse_bool(row.get("human_review_required", ""))
              changed_claims = row.get("changed_claims", "")
              changed_boundaries = row.get("changed_boundaries", "")
              new_claims = row.get("new_unsupported_claims", "")
              missed_adjacent = row.get("missed_adjacent_updates", "")
      
              if not audit_id:
                  add_issue(issues, "missing-audit-id", location, "drift audit has no audit_id")
              elif not is_valid_identifier(audit_id):
                  add_issue(
                      issues,
                      "invalid-audit-id",
                      location,
                      "audit_id must be 1-120 safe ASCII characters and must not contain path segments",
                  )
              elif audit_id in audit_ids:
                  add_issue(issues, "duplicate-audit-id", location, f"duplicate audit_id: {audit_id}")
              else:
                  audit_ids.add(audit_id)
      
              if contract_id:
                  audited_contracts.add(contract_id)
                  if not is_valid_identifier(contract_id):
                      add_issue(
                          issues,
                          "invalid-contract-id",
                          location,
                          "contract_id must be 1-120 safe ASCII characters and must not contain path segments",
                      )
                  if contract_id not in contract_ids:
                      add_issue(issues, "unknown-contract-id", location, f"audit references unknown contract_id: {contract_id}")
              else:
                  add_issue(issues, "missing-contract-id", location, "drift audit has no contract_id")
      
              if decision not in DRIFT_DECISIONS:
                  add_issue(issues, "invalid-drift-decision", location, f"invalid drift_decision: {decision}")
              if status not in AUDIT_STATUSES:
                  add_issue(issues, "invalid-audit-status", location, f"invalid audit status: {status}")
              if decision in DRIFT_DECISIONS and status in AUDIT_STATUSES:
                  resolved_decisions = {
                      "passed": {"accept", "partial_accept"},
                      "failed": {"revise", "rollback"},
                  }
                  if status in resolved_decisions and decision not in resolved_decisions[status]:
                      add_issue(
                          issues,
                          "invalid-audit-outcome",
                          location,
                          f"audit status={status} is inconsistent with drift_decision={decision}",
                      )
                  if (
                      decision in {"revise", "rollback"}
                      and status == "failed"
                      and contract_statuses.get(contract_id) == "applied"
                  ):
                      unsuccessful_contracts.add(contract_id)
                  if contract_id and (
                      (status == "passed" and decision in {"accept", "partial_accept"})
                      or (status == "failed" and decision in {"revise", "rollback"})
                  ):
                      resolved_audit_outcomes.setdefault(contract_id, set()).add(status)
              if review_required is None:
                  add_issue(issues, "invalid-human-review-required", location, "human_review_required must be true or false")
      
              high_risk = any(
                  not is_empty(value)
                  for value in [changed_claims, changed_boundaries, new_claims, missed_adjacent]
              )
              if high_risk and review_required is not True:
                  add_issue(issues, "missing-human-review", location, "claim/boundary/adjacent drift requires human_review_required=true")
              if high_risk and decision == "accept":
                  add_issue(issues, "unsafe-accept", location, "high-risk drift cannot be accepted without revision or partial acceptance")
              if (
                  strict
                  and contract_statuses.get(contract_id) == "applied"
                  and status == "needs_review"
              ):
                  add_issue(
                      issues,
                      "pending-human-review",
                      location,
                      "applied contract requires the author to resolve this drift audit before strict validation can pass",
                  )
              if not is_empty(new_claims) and status == "passed":
                  add_issue(issues, "unsupported-claim-passed", location, "new unsupported claims cannot have status=passed")
              if not is_empty(missed_adjacent) and status == "passed":
                  add_issue(issues, "missed-adjacent-passed", location, "missed adjacent updates cannot have status=passed")
      
          for contract_id, outcomes in resolved_audit_outcomes.items():
              if outcomes == {"passed", "failed"}:
                  add_issue(
                      issues,
                      "conflicting-resolved-audits",
                      "thesis_control/drift_audits.csv",
                      f"contract {contract_id} has both passed and failed resolved audits",
                  )
      
          cycle_gates_by_issue: Dict[str, List[dict]] = {}
          escalation_ids = set()
          escalation_trigger_sets = set()
          if revision_tracking:
              for index, row in enumerate(escalation_rows):
                  location = row_location("revision_escalations.csv", index)
                  escalation_id = row.get("escalation_id", "").strip()
                  revision_issue_id = row.get("revision_issue_id", "").strip()
                  trigger_contracts = [
                      value.strip()
                      for value in row.get("trigger_contracts", "").split(";")
                      if value.strip()
                  ]
                  trigger_set = set(trigger_contracts)
                  escalation_kind = row.get("escalation_kind", "").strip().lower()
                  approved_after_text = row.get("approved_after_attempt", "").strip()
                  category = row.get("primary_category", "").strip().lower()
                  writing_scope = row.get("writing_scope", "").strip().lower()
                  approved = parse_bool(row.get("human_approved", ""))
                  status = row.get("status", "").strip().lower()
      
                  if not escalation_id:
                      add_issue(issues, "missing-escalation-id", location, "revision escalation has no escalation_id")
                  elif not is_valid_identifier(escalation_id):
                      add_issue(issues, "invalid-escalation-id", location, "escalation_id must be a safe identifier")
                  elif escalation_id in escalation_ids:
                      add_issue(issues, "duplicate-escalation-id", location, f"duplicate escalation_id: {escalation_id}")
                  else:
                      escalation_ids.add(escalation_id)
      
                  if not revision_issue_id:
                      add_issue(issues, "missing-revision-issue-id", location, "revision escalation has no revision_issue_id")
                  elif not is_valid_identifier(revision_issue_id):
                      add_issue(issues, "invalid-revision-issue-id", location, "revision_issue_id must be a safe identifier")
                  elif revision_issue_id not in issue_attempts:
                      add_issue(issues, "unknown-revision-issue-id", location, f"unknown revision_issue_id: {revision_issue_id}")
      
                  if not trigger_contracts:
                      add_issue(issues, "missing-trigger-contracts", location, "revision escalation has no trigger contracts")
                  elif len(trigger_contracts) != len(trigger_set):
                      add_issue(
                          issues,
                          "duplicate-trigger-contract",
                          location,
                          "revision escalation trigger_contracts must not repeat a contract",
                      )
                  for trigger_contract in trigger_contracts:
                      if trigger_contract not in contract_ids:
                          add_issue(
                              issues,
                              "unknown-trigger-contract",
                              location,
                              f"unknown trigger contract: {trigger_contract}",
                          )
                      elif contract_issues_by_id.get(trigger_contract) != revision_issue_id:
                          add_issue(
                              issues,
                              "trigger-contract-issue-mismatch",
                              location,
                              f"trigger contract {trigger_contract} is not part of {revision_issue_id}",
                          )
      
                  if category not in ESCALATION_CATEGORIES:
                      add_issue(issues, "invalid-escalation-category", location, f"invalid primary_category: {category}")
                  if writing_scope not in WRITING_SCOPES:
                      add_issue(issues, "invalid-writing-scope", location, f"invalid writing_scope: {writing_scope}")
                  if approved is None:
                      add_issue(issues, "invalid-human-approved", location, "human_approved must be true or false")
                  if status not in ESCALATION_STATUSES:
                      add_issue(issues, "invalid-escalation-status", location, f"invalid escalation status: {status}")
                  if status == "approved" and approved is not True:
                      add_issue(issues, "missing-human-approval", location, "approved escalation requires human_approved=true")
      
                  approved_after_attempt = None
                  if enforce_revision_gate:
                      if escalation_kind not in ESCALATION_KINDS:
                          add_issue(
                              issues,
                              "invalid-escalation-kind",
                              location,
                              f"invalid escalation_kind: {escalation_kind}",
                          )
                      elif escalation_kind == "early_diagnostic":
                          if not 1 <= len(trigger_contracts) <= 2:
                              add_issue(
                                  issues,
                                  "invalid-early-diagnostic-trigger-count",
                                  location,
                                  "early_diagnostic requires one or two unique trigger contracts",
                              )
                          if approved_after_text:
                              add_issue(
                                  issues,
                                  "unexpected-approved-after-attempt",
                                  location,
                                  "early_diagnostic must leave approved_after_attempt empty",
                              )
                      elif escalation_kind == "cycle_gate":
                          if len(trigger_contracts) != 3:
                              add_issue(
                                  issues,
                                  "invalid-cycle-gate-trigger-count",
                                  location,
                                  "cycle_gate requires exactly three unique trigger contracts",
                              )
                          try:
                              approved_after_attempt = int(approved_after_text)
                          except ValueError:
                              approved_after_attempt = None
                          if approved_after_attempt is None or approved_after_attempt < 1:
                              add_issue(
                                  issues,
                                  "invalid-approved-after-attempt",
                                  location,
                                  "cycle_gate approved_after_attempt must be a positive integer",
                              )
      
                  for column in [
                      "valid_requirements",
                      "missing_or_conflicting_information",
                      "latest_author_approved_version",
                      "recommended_next_action",
                  ]:
                      if is_empty(row.get(column, "")):
                          add_issue(issues, "empty-escalation-field", location, f"revision escalation field is empty: {column}")
      
                  if revision_issue_id and trigger_set:
                      trigger_key = (revision_issue_id, frozenset(trigger_set))
                      if trigger_key in escalation_trigger_sets:
                          add_issue(
                              issues,
                              "duplicate-escalation-trigger-set",
                              location,
                              "revision issue must not repeat an escalation trigger set",
                          )
                      else:
                          escalation_trigger_sets.add(trigger_key)
                      if enforce_revision_gate and escalation_kind == "cycle_gate":
                          cycle_gates_by_issue.setdefault(revision_issue_id, []).append(
                              {
                                  "triggers": trigger_set,
                                  "ordered_triggers": trigger_contracts,
                                  "approved": approved is True and status == "approved",
                                  "approved_after_attempt": approved_after_attempt,
                                  "location": location,
                                  "effective": False,
                              }
                          )
      
              if enforce_revision_gate:
                  unsuccessful_by_issue: Dict[str, List[Tuple[int, str]]] = {}
                  for contract_id in unsuccessful_contracts:
                      revision_issue_id = contract_issues_by_id.get(contract_id)
                      attempt_no = contract_attempts.get(contract_id)
                      if revision_issue_id and attempt_no is not None:
                          unsuccessful_by_issue.setdefault(revision_issue_id, []).append((attempt_no, contract_id))
      
                  groups_by_issue: Dict[str, List[dict]] = {}
                  for revision_issue_id, attempts in unsuccessful_by_issue.items():
                      ordered = sorted(attempts)
                      complete_groups = len(ordered) // 3
                      for group_index in range(complete_groups):
                          trigger_group = ordered[group_index * 3 : (group_index + 1) * 3]
                          groups_by_issue.setdefault(revision_issue_id, []).append(
                              {
                                  "triggers": {contract_id for _, contract_id in trigger_group},
                                  "threshold_attempt": max(attempt_no for attempt_no, _ in trigger_group),
                                  "ordered": trigger_group,
                              }
                          )
      
                  for revision_issue_id, cycle_gates in cycle_gates_by_issue.items():
                      groups = groups_by_issue.get(revision_issue_id, [])
                      for cycle_gate in cycle_gates:
                          matching_group = next(
                              (group for group in groups if group["triggers"] == cycle_gate["triggers"]),
                              None,
                          )
                          if matching_group is None:
                              add_issue(
                                  issues,
                                  "invalid-cycle-gate-trigger-group",
                                  cycle_gate["location"],
                                  "cycle_gate triggers do not match a completed unsuccessful group",
                              )
                              if cycle_gate["approved"]:
                                  add_issue(
                                      issues,
                                      "premature-cycle-gate-approval",
                                      cycle_gate["location"],
                                      "approved cycle_gate requires a completed unsuccessful group",
                                  )
                              continue
                          expected_order = [
                              contract_id for _, contract_id in matching_group["ordered"]
                          ]
                          if cycle_gate["ordered_triggers"] != expected_order:
                              add_issue(
                                  issues,
                                  "misordered-cycle-gate-triggers",
                                  cycle_gate["location"],
                                  "cycle_gate trigger_contracts must follow revision attempt order",
                              )
                              continue
                          if cycle_gate["approved_after_attempt"] != matching_group["threshold_attempt"]:
                              add_issue(
                                  issues,
                                  "invalid-approved-after-attempt",
                                  cycle_gate["location"],
                                  "cycle_gate approved_after_attempt must equal the group's final attempt",
                              )
                              continue
                          cycle_gate["effective"] = cycle_gate["approved"]
      
                  for revision_issue_id, groups in groups_by_issue.items():
                      for group in groups:
                          required_triggers = group["triggers"]
                          threshold_attempt = group["threshold_attempt"]
                          trigger_group = group["ordered"]
                          matching = [
                              cycle_gate
                              for cycle_gate in cycle_gates_by_issue.get(revision_issue_id, [])
                              if required_triggers == cycle_gate["triggers"]
                          ]
                          trigger_list = ", ".join(contract_id for _, contract_id in trigger_group)
                          if not matching:
                              add_issue(
                                  issues,
                                  "missing-revision-escalation",
                                  "thesis_control/revision_escalations.csv",
                                  f"revision issue {revision_issue_id} needs a cycle_gate for {trigger_list}",
                              )
                          if any(cycle_gate["effective"] for cycle_gate in matching):
                              continue
                          for contract_id, attempt_no in contract_attempts.items():
                              if contract_issues_by_id.get(contract_id) != revision_issue_id:
                                  continue
                              if attempt_no <= threshold_attempt:
                                  continue
                              if contract_statuses.get(contract_id) in {"approved", "applied"}:
                                  add_issue(
                                      issues,
                                      "revision-escalation-required",
                                      contract_locations.get(contract_id, "thesis_control/edit_contracts.csv"),
                                      f"contract {contract_id} cannot proceed before revision issue {revision_issue_id} has an approved cycle_gate for {trigger_list}",
                                  )
      
          executable_contracts = [
              contract_id
              for contract_id, status in contract_statuses.items()
              if status in {"approved", "applied"}
          ]
          if intent_tracking and executable_contracts:
              if len(intent_state["active_intent_ids"]) != 1:
                  add_issue(
                      issues,
                      "active-project-intent-required",
                      "thesis_control/project_intent.csv",
                      "approved/applied edits require exactly one active author-approved project intent",
                  )
              if len(intent_state["active_manuscript_ids"]) != 1:
                  add_issue(
                      issues,
                      "active-manuscript-contract-required",
                      "thesis_control/manuscript_contracts.csv",
                      "approved/applied edits require exactly one active author-approved manuscript contract",
                  )
      
          if strict:
              missing_audits = sorted(applied_contracts - audited_contracts)
              for contract_id in missing_audits:
                  add_issue(
                      issues,
                      "missing-drift-audit",
                      "thesis_control/drift_audits.csv",
                      f"applied contract has no drift audit: {contract_id}",
                  )
      
          return {
              "schema_version": 1,
              "base_dir": str(root),
              "strict": strict,
              "summary": {
                  "spine_cards": len(spine_rows),
                  "edit_contracts": len(contract_rows),
                  "drift_audits": len(audit_rows),
                  "revision_escalations": len(escalation_rows),
                  "revision_tracking": revision_tracking,
                  "escalation_schema_version": 3 if escalation_v3 else 2,
                  "project_intent_tracking": intent_tracking,
                  "project_intents": len(intent_state["intent_rows"]),
                  "manuscript_contracts": len(intent_state["manuscript_rows"]),
                  "global_thesis_audits": len(intent_state["audit_rows"]),
              },
              "issues": issues,
              "issue_count": len(issues),
          }
      
      
      def main(argv: Iterable[str] | None = None) -> int:
          parser = argparse.ArgumentParser(description="Validate thesis-control spine, contract, and drift-audit files.")
          parser.add_argument("base_dir", nargs="?", default=".", help="Project root containing thesis_control/")
          parser.add_argument("--strict", action="store_true", help="Require every applied contract to have a drift audit")
          parser.add_argument("--json", action="store_true", dest="emit_json", help="Emit JSON output")
          args = parser.parse_args(list(argv) if argv is not None else None)
      
          root = Path(args.base_dir).resolve()
          payload = validate_packet(root, strict=args.strict)
      
          if args.emit_json:
              print(json.dumps(payload, indent=2))
          else:
              print(f"Thesis-control root: {root}")
              if payload["issues"]:
                  for issue in payload["issues"]:
                      print(f"- {issue['location']}: {issue['kind']}: {issue['message']}")
              else:
                  print("- no thesis-control issues detected")
      
          return 1 if payload["issues"] else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • project_intent_control.py 23 KB
      #!/usr/bin/env python3
      """Validate the project-intent and manuscript-level thesis-control layer.
      
      This module is deliberately structural. It validates durable author-approved
      objects and their gates; it does not infer whether manuscript prose is
      semantically aligned with the recorded intent.
      """
      
      from __future__ import annotations
      
      import re
      from pathlib import Path
      from typing import Dict, List, Mapping, Optional, Sequence, Tuple
      
      from thesis_control_io import CsvShapeError, read_csv_table
      
      
      PROJECT_INTENT_COLUMNS = [
          "intent_id",
          "intent_version",
          "supersedes_intent_id",
          "primary_domain",
          "research_object",
          "core_research_question",
          "target_venue",
          "must_include_concepts",
          "excluded_reframes",
          "amendment_reason",
          "approval_evidence",
          "human_approved",
          "status",
      ]
      
      MANUSCRIPT_CONTRACT_COLUMNS = [
          "manuscript_id",
          "intent_id",
          "manuscript_version",
          "supersedes_manuscript_id",
          "title",
          "abstract_focus",
          "primary_domain",
          "research_object",
          "research_question",
          "contribution_scope",
          "structure_summary",
          "change_summary",
          "human_approved",
          "status",
      ]
      
      GLOBAL_THESIS_AUDIT_COLUMNS = [
          "global_audit_id",
          "intent_id",
          "manuscript_id",
          "manuscript_version",
          "title_alignment",
          "abstract_alignment",
          "primary_domain_alignment",
          "research_object_alignment",
          "research_question_alignment",
          "contribution_alignment",
          "structure_alignment",
          "detected_reframe",
          "reframe_summary",
          "human_review_required",
          "human_decision",
          "status",
      ]
      
      PROJECT_INTENT_FILES = {
          "project_intent.csv": PROJECT_INTENT_COLUMNS,
          "manuscript_contracts.csv": MANUSCRIPT_CONTRACT_COLUMNS,
          "global_thesis_audits.csv": GLOBAL_THESIS_AUDIT_COLUMNS,
      }
      
      INTENT_STATUSES = {"draft", "active", "superseded", "rejected"}
      MANUSCRIPT_STATUSES = {"draft", "active", "superseded", "rejected"}
      GLOBAL_AUDIT_STATUSES = {"passed", "needs_review", "failed"}
      ALIGNMENT_STATUSES = {"aligned", "drifted", "not_assessed"}
      GLOBAL_DECISIONS = {
          "accept",
          "revise_manuscript",
          "rollback",
          "amend_intent",
          "pending",
      }
      EMPTY_MARKERS = {"", "none", "n/a", "na", "no", "not applicable"}
      IDENTIFIER_RE = re.compile(r"^[A-Za-z0-9](?:[A-Za-z0-9._-]{0,118}[A-Za-z0-9])?$")
      ALIGNMENT_COLUMNS = [
          "title_alignment",
          "abstract_alignment",
          "primary_domain_alignment",
          "research_object_alignment",
          "research_question_alignment",
          "contribution_alignment",
          "structure_alignment",
      ]
      
      
      def is_empty(value: str) -> bool:
          return value.strip().lower() in EMPTY_MARKERS
      
      
      def parse_bool(value: str) -> bool | None:
          lowered = value.strip().lower()
          if lowered in {"true", "yes", "y", "1"}:
              return True
          if lowered in {"false", "no", "n", "0"}:
              return False
          return None
      
      
      def is_valid_identifier(value: str) -> bool:
          stripped = value.strip()
          return bool(IDENTIFIER_RE.fullmatch(stripped)) and ".." not in stripped
      
      
      def row_location(filename: str, index: int) -> str:
          return f"thesis_control/{filename}:row {index + 2}"
      
      
      def add_issue(issues: List[dict], kind: str, location: str, message: str) -> None:
          issues.append({"kind": kind, "location": location, "message": message})
      
      
      def read_control_table(
          root: Path,
          filename: str,
          required_columns: Sequence[str],
          table: Optional[Tuple[Sequence[str], Sequence[Mapping[str, str]]]] = None,
      ) -> Tuple[List[Dict[str, str]], List[dict]]:
          path = root / "thesis_control" / filename
          issues: List[dict] = []
          if table is None:
              if not path.is_file():
                  return [], [
                      {
                          "kind": "missing-file",
                          "location": str(path),
                          "message": f"missing {filename}",
                      }
                  ]
              try:
                  fieldnames, rows = read_csv_table(path)
              except CsvShapeError as exc:
                  return [], [
                      {"kind": exc.kind, "location": exc.location, "message": exc.message}
                  ]
              except OSError as exc:
                  return [], [
                      {
                          "kind": "csv-error",
                          "location": str(path),
                          "message": f"cannot read file: {exc}",
                      }
                  ]
          else:
              fieldnames = list(table[0])
              rows = [dict(row) for row in table[1]]
      
          for column in required_columns:
              if column not in fieldnames:
                  add_issue(
                      issues,
                      "missing-column",
                      str(path),
                      f"{filename} missing column: {column}",
                  )
          if not rows:
              add_issue(
                  issues,
                  "empty-file",
                  str(path),
                  f"{filename} has no data rows",
              )
          return rows, issues
      
      
      def parse_positive_version(
          issues: List[dict], value: str, kind: str, location: str, label: str
      ) -> Optional[int]:
          try:
              version = int(value.strip())
          except ValueError:
              version = 0
          if version < 1:
              add_issue(issues, kind, location, f"{label} must be a positive integer")
              return None
          return version
      
      
      def validate_project_intent_layer(
          root: Path,
          required: bool,
          table_overrides: Optional[
              Mapping[str, Tuple[Sequence[str], Sequence[Mapping[str, str]]]]
          ] = None,
      ) -> Tuple[dict, List[dict]]:
          """Return validated intent-layer state and located structural issues."""
      
          overrides = table_overrides or {}
          enabled = required or any(
              filename in overrides or (root / "thesis_control" / filename).exists()
              for filename in PROJECT_INTENT_FILES
          )
          empty_state = {
              "enabled": False,
              "intent_rows": [],
              "manuscript_rows": [],
              "audit_rows": [],
              "intents_by_id": {},
              "manuscripts_by_id": {},
              "audits_by_id": {},
              "active_intent_ids": set(),
              "active_manuscript_ids": set(),
              "ready_audit_ids": set(),
          }
          if not enabled:
              return empty_state, []
      
          issues: List[dict] = []
          tables: Dict[str, List[Dict[str, str]]] = {}
          for filename, columns in PROJECT_INTENT_FILES.items():
              rows, table_issues = read_control_table(
                  root,
                  filename,
                  columns,
                  table=overrides.get(filename),
              )
              tables[filename] = rows
              issues.extend(table_issues)
      
          intent_rows = tables["project_intent.csv"]
          manuscript_rows = tables["manuscript_contracts.csv"]
          audit_rows = tables["global_thesis_audits.csv"]
      
          intents_by_id: Dict[str, dict] = {}
          intent_versions: Dict[int, str] = {}
          active_intent_ids = set()
          for index, row in enumerate(intent_rows):
              location = row_location("project_intent.csv", index)
              intent_id = row.get("intent_id", "").strip()
              status = row.get("status", "").strip().lower()
              approved = parse_bool(row.get("human_approved", ""))
              version = parse_positive_version(
                  issues,
                  row.get("intent_version", ""),
                  "invalid-intent-version",
                  location,
                  "intent_version",
              )
              if not intent_id:
                  add_issue(issues, "missing-intent-id", location, "project intent has no intent_id")
              elif not is_valid_identifier(intent_id):
                  add_issue(issues, "invalid-intent-id", location, "intent_id must be a safe identifier")
              elif intent_id in intents_by_id:
                  add_issue(issues, "duplicate-intent-id", location, f"duplicate intent_id: {intent_id}")
              else:
                  intents_by_id[intent_id] = {
                      "row": row,
                      "status": status,
                      "approved": approved,
                      "version": version,
                      "location": location,
                  }
              if version is not None:
                  if version in intent_versions:
                      add_issue(
                          issues,
                          "duplicate-intent-version",
                          location,
                          f"intent_version {version} is already used by {intent_versions[version]}",
                      )
                  else:
                      intent_versions[version] = intent_id
              if status not in INTENT_STATUSES:
                  add_issue(issues, "invalid-intent-status", location, f"invalid project intent status: {status}")
              if approved is None:
                  add_issue(issues, "invalid-human-approved", location, "human_approved must be true or false")
              if status in {"active", "superseded"} and approved is not True:
                  add_issue(
                      issues,
                      "missing-intent-approval",
                      location,
                      "active/superseded project intent requires human_approved=true",
                  )
              if status == "active" and intent_id:
                  active_intent_ids.add(intent_id)
              for column in [
                  "primary_domain",
                  "research_object",
                  "core_research_question",
                  "target_venue",
                  "must_include_concepts",
                  "excluded_reframes",
                  "amendment_reason",
                  "approval_evidence",
              ]:
                  if is_empty(row.get(column, "")):
                      add_issue(
                          issues,
                          "empty-intent-field",
                          location,
                          f"project intent field is empty: {column}",
                      )
      
          if len(active_intent_ids) > 1:
              add_issue(
                  issues,
                  "multiple-active-intents",
                  "thesis_control/project_intent.csv",
                  "project intent history must have at most one active row",
              )
      
          for intent_id, record in intents_by_id.items():
              row = record["row"]
              version = record["version"]
              supersedes = row.get("supersedes_intent_id", "").strip()
              location = record["location"]
              if version == 1 and supersedes:
                  add_issue(
                      issues,
                      "invalid-intent-lineage",
                      location,
                      "intent_version 1 must not supersede another intent",
                  )
              if version is not None and version > 1:
                  if not supersedes:
                      add_issue(
                          issues,
                          "missing-intent-amendment",
                          location,
                          "later project intent versions must name supersedes_intent_id",
                      )
                  elif supersedes not in intents_by_id:
                      add_issue(
                          issues,
                          "unknown-superseded-intent",
                          location,
                          f"unknown supersedes_intent_id: {supersedes}",
                      )
                  else:
                      previous = intents_by_id[supersedes]
                      if previous["version"] != version - 1:
                          add_issue(
                              issues,
                              "nonsequential-intent-lineage",
                              location,
                              "project intent amendments must supersede the immediately previous version",
                          )
                      if record["status"] == "active" and previous["status"] != "superseded":
                          add_issue(
                              issues,
                              "unsuperseded-prior-intent",
                              location,
                              "an active amendment requires the previous intent status=superseded",
                          )
      
          manuscripts_by_id: Dict[str, dict] = {}
          manuscript_versions: Dict[int, str] = {}
          active_manuscript_ids = set()
          for index, row in enumerate(manuscript_rows):
              location = row_location("manuscript_contracts.csv", index)
              manuscript_id = row.get("manuscript_id", "").strip()
              intent_id = row.get("intent_id", "").strip()
              status = row.get("status", "").strip().lower()
              approved = parse_bool(row.get("human_approved", ""))
              version = parse_positive_version(
                  issues,
                  row.get("manuscript_version", ""),
                  "invalid-manuscript-version",
                  location,
                  "manuscript_version",
              )
              if not manuscript_id:
                  add_issue(issues, "missing-manuscript-id", location, "manuscript contract has no manuscript_id")
              elif not is_valid_identifier(manuscript_id):
                  add_issue(issues, "invalid-manuscript-id", location, "manuscript_id must be a safe identifier")
              elif manuscript_id in manuscripts_by_id:
                  add_issue(
                      issues,
                      "duplicate-manuscript-id",
                      location,
                      f"duplicate manuscript_id: {manuscript_id}",
                  )
              else:
                  manuscripts_by_id[manuscript_id] = {
                      "row": row,
                      "intent_id": intent_id,
                      "status": status,
                      "approved": approved,
                      "version": version,
                      "location": location,
                  }
              if version is not None:
                  if version in manuscript_versions:
                      add_issue(
                          issues,
                          "duplicate-manuscript-version",
                          location,
                          f"manuscript_version {version} is already used by {manuscript_versions[version]}",
                      )
                  else:
                      manuscript_versions[version] = manuscript_id
              if not intent_id:
                  add_issue(issues, "missing-intent-id", location, "manuscript contract has no intent_id")
              elif intent_id not in intents_by_id:
                  add_issue(issues, "unknown-intent-id", location, f"unknown intent_id: {intent_id}")
              if status not in MANUSCRIPT_STATUSES:
                  add_issue(issues, "invalid-manuscript-status", location, f"invalid manuscript status: {status}")
              if approved is None:
                  add_issue(issues, "invalid-human-approved", location, "human_approved must be true or false")
              if status in {"active", "superseded"} and approved is not True:
                  add_issue(
                      issues,
                      "missing-manuscript-approval",
                      location,
                      "active/superseded manuscript contract requires human_approved=true",
                  )
              if status == "active" and manuscript_id:
                  active_manuscript_ids.add(manuscript_id)
                  if intent_id not in active_intent_ids:
                      add_issue(
                          issues,
                          "inactive-intent-reference",
                          location,
                          "active manuscript contract must reference the active project intent",
                      )
              for column in [
                  "title",
                  "abstract_focus",
                  "primary_domain",
                  "research_object",
                  "research_question",
                  "contribution_scope",
                  "structure_summary",
                  "change_summary",
              ]:
                  if is_empty(row.get(column, "")):
                      add_issue(
                          issues,
                          "empty-manuscript-field",
                          location,
                          f"manuscript contract field is empty: {column}",
                      )
      
          if len(active_manuscript_ids) > 1:
              add_issue(
                  issues,
                  "multiple-active-manuscripts",
                  "thesis_control/manuscript_contracts.csv",
                  "manuscript contract history must have at most one active row",
              )
      
          for manuscript_id, record in manuscripts_by_id.items():
              row = record["row"]
              version = record["version"]
              supersedes = row.get("supersedes_manuscript_id", "").strip()
              location = record["location"]
              if version == 1 and supersedes:
                  add_issue(
                      issues,
                      "invalid-manuscript-lineage",
                      location,
                      "manuscript_version 1 must not supersede another manuscript contract",
                  )
              if version is not None and version > 1:
                  if not supersedes:
                      add_issue(
                          issues,
                          "missing-manuscript-amendment",
                          location,
                          "later manuscript versions must name supersedes_manuscript_id",
                      )
                  elif supersedes not in manuscripts_by_id:
                      add_issue(
                          issues,
                          "unknown-superseded-manuscript",
                          location,
                          f"unknown supersedes_manuscript_id: {supersedes}",
                      )
                  else:
                      previous = manuscripts_by_id[supersedes]
                      if previous["version"] != version - 1:
                          add_issue(
                              issues,
                              "nonsequential-manuscript-lineage",
                              location,
                              "manuscript amendments must supersede the immediately previous version",
                          )
                      if record["status"] == "active" and previous["status"] != "superseded":
                          add_issue(
                              issues,
                              "unsuperseded-prior-manuscript",
                              location,
                              "an active manuscript amendment requires the previous status=superseded",
                          )
      
          audits_by_id: Dict[str, dict] = {}
          ready_audit_ids = set()
          for index, row in enumerate(audit_rows):
              location = row_location("global_thesis_audits.csv", index)
              audit_id = row.get("global_audit_id", "").strip()
              intent_id = row.get("intent_id", "").strip()
              manuscript_id = row.get("manuscript_id", "").strip()
              status = row.get("status", "").strip().lower()
              decision = row.get("human_decision", "").strip().lower()
              detected_reframe = parse_bool(row.get("detected_reframe", ""))
              review_required = parse_bool(row.get("human_review_required", ""))
              version = parse_positive_version(
                  issues,
                  row.get("manuscript_version", ""),
                  "invalid-audit-manuscript-version",
                  location,
                  "manuscript_version",
              )
              if not audit_id:
                  add_issue(issues, "missing-global-audit-id", location, "global thesis audit has no global_audit_id")
              elif not is_valid_identifier(audit_id):
                  add_issue(issues, "invalid-global-audit-id", location, "global_audit_id must be a safe identifier")
              elif audit_id in audits_by_id:
                  add_issue(
                      issues,
                      "duplicate-global-audit-id",
                      location,
                      f"duplicate global_audit_id: {audit_id}",
                  )
              else:
                  audits_by_id[audit_id] = {
                      "row": row,
                      "intent_id": intent_id,
                      "manuscript_id": manuscript_id,
                      "status": status,
                      "location": location,
                  }
              if intent_id not in intents_by_id:
                  add_issue(issues, "unknown-intent-id", location, f"unknown intent_id: {intent_id}")
              if manuscript_id not in manuscripts_by_id:
                  add_issue(issues, "unknown-manuscript-id", location, f"unknown manuscript_id: {manuscript_id}")
              else:
                  manuscript = manuscripts_by_id[manuscript_id]
                  if manuscript["intent_id"] != intent_id:
                      add_issue(
                          issues,
                          "audit-intent-mismatch",
                          location,
                          "global thesis audit intent_id does not match its manuscript contract",
                      )
                  if version is not None and manuscript["version"] != version:
                      add_issue(
                          issues,
                          "stale-global-audit-version",
                          location,
                          "global thesis audit manuscript_version does not match its manuscript contract",
                      )
              alignments = [row.get(column, "").strip().lower() for column in ALIGNMENT_COLUMNS]
              for column, value in zip(ALIGNMENT_COLUMNS, alignments):
                  if value not in ALIGNMENT_STATUSES:
                      add_issue(
                          issues,
                          "invalid-global-alignment",
                          location,
                          f"invalid {column}: {value}",
                      )
              if detected_reframe is None:
                  add_issue(issues, "invalid-detected-reframe", location, "detected_reframe must be true or false")
              if review_required is None:
                  add_issue(
                      issues,
                      "invalid-human-review-required",
                      location,
                      "human_review_required must be true or false",
                  )
              if decision not in GLOBAL_DECISIONS:
                  add_issue(issues, "invalid-global-decision", location, f"invalid human_decision: {decision}")
              if status not in GLOBAL_AUDIT_STATUSES:
                  add_issue(issues, "invalid-global-audit-status", location, f"invalid global audit status: {status}")
              if is_empty(row.get("reframe_summary", "")):
                  add_issue(issues, "empty-reframe-summary", location, "global thesis audit needs reframe_summary")
      
              has_drift = detected_reframe is True or "drifted" in alignments
              not_assessed = "not_assessed" in alignments
              if has_drift and review_required is not True:
                  add_issue(
                      issues,
                      "missing-global-human-review",
                      location,
                      "global thesis drift or reframing requires human_review_required=true",
                  )
              if has_drift and (status == "passed" or decision == "accept"):
                  add_issue(
                      issues,
                      "unsafe-global-pass",
                      location,
                      "global thesis drift cannot pass or be accepted; revise, rollback, or amend intent",
                  )
              if not_assessed and status != "needs_review":
                  add_issue(
                      issues,
                      "unassessed-global-audit",
                      location,
                      "not_assessed alignment requires status=needs_review",
                  )
              if status == "needs_review" and decision != "pending":
                  add_issue(
                      issues,
                      "invalid-pending-global-decision",
                      location,
                      "status=needs_review requires human_decision=pending",
                  )
              if status == "passed":
                  passed_shape = (
                      all(value == "aligned" for value in alignments)
                      and detected_reframe is False
                      and decision == "accept"
                      and intent_id in active_intent_ids
                      and manuscript_id in active_manuscript_ids
                  )
                  if not passed_shape:
                      add_issue(
                          issues,
                          "invalid-global-pass",
                          location,
                          "passed audit requires complete alignment and active approved intent/manuscript contracts",
                      )
                  elif audit_id:
                      ready_audit_ids.add(audit_id)
      
          return {
              "enabled": True,
              "intent_rows": intent_rows,
              "manuscript_rows": manuscript_rows,
              "audit_rows": audit_rows,
              "intents_by_id": intents_by_id,
              "manuscripts_by_id": manuscripts_by_id,
              "audits_by_id": audits_by_id,
              "active_intent_ids": active_intent_ids,
              "active_manuscript_ids": active_manuscript_ids,
              "ready_audit_ids": ready_audit_ids,
          }, issues
      
    • scaffold_thesis_control.py 28.3 KB
      #!/usr/bin/env python3
      """Scaffold a thesis-control draft packet from a real manuscript unit.
      
      The scaffold is deliberately conservative. It creates a spine-card row and a
      draft edit-contract row, but it does not mark any prose change as approved or
      audited. The author still owns the scholarly judgement.
      """
      
      from __future__ import annotations
      
      import argparse
      import json
      import re
      import sys
      from pathlib import Path
      from typing import Dict, Iterable, List, Optional, Sequence, Tuple
      
      from check_thesis_control import validate_packet
      from project_intent_control import (
          GLOBAL_THESIS_AUDIT_COLUMNS,
          MANUSCRIPT_CONTRACT_COLUMNS,
          PROJECT_INTENT_COLUMNS,
      )
      from thesis_control_io import (
          atomic_write_batch,
          ensure_internal_paths,
          read_csv_table,
          render_csv_table,
      )
      
      
      SPINE_COLUMNS = [
          "unit_id",
          "manuscript_id",
          "path",
          "section_title",
          "spine_sentence",
          "scope_boundary",
          "core_claims",
          "do_not_change",
      ]
      
      CONTRACT_COLUMNS = [
          "contract_id",
          "unit_id",
          "global_audit_id",
          "revision_issue_id",
          "attempt_no",
          "change_scope",
          "allowed_changes",
          "forbidden_changes",
          "adjacent_context",
          "acceptance_checks",
          "human_approved",
          "status",
      ]
      
      AUDIT_COLUMNS = [
          "audit_id",
          "contract_id",
          "changed_claims",
          "changed_boundaries",
          "new_unsupported_claims",
          "missed_adjacent_updates",
          "drift_decision",
          "human_review_required",
          "status",
      ]
      
      ESCALATION_COLUMNS = [
          "escalation_id",
          "revision_issue_id",
          "escalation_kind",
          "trigger_contracts",
          "approved_after_attempt",
          "primary_category",
          "writing_scope",
          "valid_requirements",
          "missing_or_conflicting_information",
          "latest_author_approved_version",
          "recommended_next_action",
          "human_approved",
          "status",
      ]
      
      AUTHOR_REVIEW = "AUTHOR_REVIEW_REQUIRED"
      IDENTIFIER_RE = re.compile(r"^[A-Za-z0-9](?:[A-Za-z0-9._-]{0,118}[A-Za-z0-9])?$")
      
      
      def review_required(label: str) -> str:
          return f"{AUTHOR_REVIEW}: {label}"
      
      
      def validate_identifier(label: str, value: str) -> None:
          if not IDENTIFIER_RE.fullmatch(value) or ".." in value:
              raise ValueError(
                  f"{label} must be 1-120 ASCII letters, numbers, dots, underscores, or hyphens; "
                  "it must start and end with a letter or number and must not contain '..'"
              )
      
      
      def resolve_source(project_root: Path, source: str) -> Path:
          candidate = Path(source).expanduser()
          if not candidate.is_absolute():
              candidate = project_root / candidate
          return candidate.resolve()
      
      
      def relative_display(path: Path, base: Path) -> str:
          try:
              return path.resolve().relative_to(base.resolve()).as_posix()
          except ValueError:
              return str(path.resolve())
      
      
      def infer_unit_id(source: Path, start_line: Optional[int], end_line: Optional[int]) -> str:
          stem = re.sub(r"[^A-Za-z0-9]+", "-", source.stem).strip("-").lower() or "unit"
          if start_line is not None or end_line is not None:
              return f"{stem}-l{start_line or 1}-l{end_line or 'end'}"
          return stem
      
      
      def read_excerpt(path: Path, start_line: Optional[int], end_line: Optional[int]) -> str:
          if start_line is not None and start_line < 1:
              raise ValueError("--start-line must be >= 1")
          if end_line is not None and end_line < 1:
              raise ValueError("--end-line must be >= 1")
          if start_line is not None and end_line is not None and end_line < start_line:
              raise ValueError("--end-line must be greater than or equal to --start-line")
      
          lines = path.read_text(encoding="utf-8").splitlines()
          if not lines:
              raise ValueError(f"source file is empty: {path}")
          total_lines = len(lines)
          if start_line is not None and start_line > total_lines:
              raise ValueError(f"--start-line {start_line} is beyond end of file ({total_lines} lines)")
          if end_line is not None and end_line > total_lines:
              raise ValueError(f"--end-line {end_line} is beyond end of file ({total_lines} lines)")
      
          start = (start_line - 1) if start_line is not None else 0
          end = end_line if end_line is not None else len(lines)
          excerpt = "\n".join(lines[start:end]).rstrip()
          if not excerpt.strip():
              raise ValueError("selected source excerpt is empty")
          return excerpt + "\n"
      
      
      def infer_section_title(excerpt: str, source: Path) -> str:
          for line in excerpt.splitlines():
              stripped = line.strip()
              if stripped.startswith("#"):
                  return stripped.lstrip("#").strip() or source.stem
          return source.stem.replace("_", " ").replace("-", " ").strip() or source.name
      
      
      def read_owned_csv(path: Path, columns: Sequence[str]) -> List[Dict[str, str]]:
          """Read one scaffold-owned CSV without changing the project tree."""
      
          if not path.exists():
              return []
      
          fieldnames, rows = read_csv_table(path)
          missing = [column for column in columns if column not in fieldnames]
          if missing:
              raise ValueError(f"{path} missing column(s): {', '.join(missing)}")
          unsupported = [column for column in fieldnames if column not in columns]
          if unsupported:
              raise ValueError(f"{path} unsupported column(s): {', '.join(unsupported)}")
          return rows
      
      
      def upsert_row_candidate(
          path: Path,
          rows: Sequence[Dict[str, str]],
          key: str,
          row: Dict[str, str],
          force: bool,
      ) -> Tuple[List[Dict[str, str]], str]:
          matches = sum(1 for existing in rows if existing.get(key) == row[key])
          if matches > 1:
              raise ValueError(f"{path} contains duplicate {key}={row[key]}")
          if matches == 1 and not force:
              raise ValueError(f"{path} already contains {key}={row[key]}; pass --force to replace it")
      
          output: List[Dict[str, str]] = []
          for existing in rows:
              if existing.get(key) == row[key]:
                  output.append(row)
              else:
                  output.append(existing)
          if matches == 0:
              output.append(row)
          return output, "replaced" if matches == 1 else "added"
      
      
      def validate_revision_attempt(
          path: Path,
          rows: Sequence[Dict[str, str]],
          contract_id: str,
          revision_issue_id: str,
          attempt_no: int,
          force: bool,
      ) -> None:
          candidate_rows: List[Dict[str, str]] = []
          for row in rows:
              if row.get("contract_id", "").strip() == contract_id:
                  if not force:
                      raise ValueError(
                          f"{path} already contains contract_id={contract_id}; pass --force to replace it"
                      )
                  continue
              candidate_rows.append(row)
      
          candidate_rows.append(
              {
                  "contract_id": contract_id,
                  "revision_issue_id": revision_issue_id,
                  "attempt_no": str(attempt_no),
              }
          )
          attempts_by_issue: Dict[str, List[int]] = {}
          for row in candidate_rows:
              issue_id = row.get("revision_issue_id", "").strip()
              validate_identifier("revision_issue_id", issue_id)
              attempt_text = row.get("attempt_no", "").strip()
              try:
                  row_attempt = int(attempt_text)
              except ValueError as exc:
                  raise ValueError(f"attempt_no must be a positive integer: {attempt_text}") from exc
              if row_attempt < 1:
                  raise ValueError(f"attempt_no must be a positive integer: {attempt_text}")
              attempts_by_issue.setdefault(issue_id, []).append(row_attempt)
      
          for issue_id, attempts in attempts_by_issue.items():
              ordered = sorted(attempts)
              expected = list(range(1, len(ordered) + 1))
              if ordered != expected:
                  raise ValueError(
                      f"revision issue {issue_id} attempts must be unique and sequential from 1"
                  )
      
      
      def render_review_packet(
          unit_id: str,
          section_title: str,
          excerpt_path: str,
          intent_id: str,
          manuscript_id: str,
          global_audit_id: str,
          spine_row: Dict[str, str],
          contract_row: Dict[str, str],
      ) -> str:
          return f"""# Thesis-Control Review Packet: {unit_id}
      
      ## Unit
      
      - Source: `{excerpt_path}`
      - Section: {section_title}
      - Contract: `{contract_row['contract_id']}`
      - Project intent: `{intent_id}`
      - Manuscript contract: `{manuscript_id}`
      - Global thesis audit: `{global_audit_id}`
      - Revision issue: `{contract_row['revision_issue_id']}`
      - Attempt: {contract_row['attempt_no']}
      - Status: draft, not approved
      
      ## Spine Card Draft
      
      - Spine sentence: {spine_row['spine_sentence']}
      - Scope boundary: {spine_row['scope_boundary']}
      - Core claims: {spine_row['core_claims']}
      - Do not change: {spine_row['do_not_change']}
      
      ## Edit Contract Draft
      
      - Change scope: {contract_row['change_scope']}
      - Allowed changes: {contract_row['allowed_changes']}
      - Forbidden changes: {contract_row['forbidden_changes']}
      - Adjacent context: {contract_row['adjacent_context']}
      - Acceptance checks: {contract_row['acceptance_checks']}
      
      ## Author Gate
      
      - Approve the project intent and manuscript contract before approving any edit contract.
      - Resolve the global thesis audit as `passed` only when every alignment field is `aligned` and no reframe is detected.
      - If title, abstract, primary domain, research object, research question, contribution, or structure drifts, revise or roll back the manuscript, or create an explicitly approved new intent version. Never overwrite the earlier intent row.
      - Before editing prose, replace every `{AUTHOR_REVIEW}` field with a concrete judgement.
      - Keep `human_approved=false` until the author explicitly approves the contract.
      - After any applied edit, add a drift-audit row before accepting the prose.
      - Reuse the same revision issue id and increment the attempt number when a new contract retries the same unresolved problem.
      - Count an attempt as unsuccessful only after an applied contract receives a resolved `revise` or `rollback` audit with `status=failed`.
      - An `early_diagnostic` escalation may record one or two warning triggers, but it never closes or pre-authorises a failure cycle.
      - After three unsuccessful attempts, create a distinct `cycle_gate` that lists exactly those three trigger contracts, records the third attempt as its approval boundary, and receives explicit author approval before applying a later contract.
      """
      
      
      def scaffold(args: argparse.Namespace) -> dict:
          project_root = Path(args.project_root).expanduser().resolve()
          output_dir = Path(args.output_dir).expanduser().resolve() if args.output_dir else project_root
          source = resolve_source(project_root, args.source)
          if not source.is_file():
              raise FileNotFoundError(f"source file not found: {source}")
      
          excerpt = read_excerpt(source, args.start_line, args.end_line)
          unit_id = args.unit_id or infer_unit_id(source, args.start_line, args.end_line)
          section_title = args.section_title or infer_section_title(excerpt, source)
          attempt_no = args.attempt_no
          contract_id = args.contract_id or f"ec-{unit_id}-{attempt_no:03d}"
          revision_issue_id = args.revision_issue_id or f"ri-{contract_id}"
          validate_identifier("unit_id", unit_id)
          validate_identifier("contract_id", contract_id)
          validate_identifier("revision_issue_id", revision_issue_id)
          if attempt_no < 1:
              raise ValueError("--attempt-no must be >= 1")
      
          control_dir = output_dir / "thesis_control"
          intent_path = control_dir / "project_intent.csv"
          manuscript_path = control_dir / "manuscript_contracts.csv"
          global_audit_path = control_dir / "global_thesis_audits.csv"
          spine_path = control_dir / "spine_cards.csv"
          contract_path = control_dir / "edit_contracts.csv"
          audit_path = control_dir / "drift_audits.csv"
          escalation_path = control_dir / "revision_escalations.csv"
          packet_path = control_dir / f"{unit_id}_review_packet.md"
      
          excerpt_file: Optional[Path] = None
          if args.copy_source:
              excerpt_dir = output_dir / "source_excerpts"
              excerpt_file = excerpt_dir / f"{unit_id}.md"
          else:
              try:
                  source.resolve().relative_to(output_dir)
              except ValueError as exc:
                  raise ValueError("source must be inside output-dir when --copy-source is not used") from exc
              source_display = relative_display(source, output_dir)
      
          output_targets = [
              intent_path,
              manuscript_path,
              global_audit_path,
              spine_path,
              contract_path,
              audit_path,
              escalation_path,
              packet_path,
          ]
          if excerpt_file is not None:
              output_targets.append(excerpt_file)
          ensure_internal_paths(output_dir, output_targets)
      
          if excerpt_file is not None:
              if excerpt_file.exists() and not args.force:
                  raise ValueError(f"{excerpt_file} already exists; pass --force to replace it")
              source_display = relative_display(excerpt_file, output_dir)
      
          if packet_path.exists() and not args.force:
              raise ValueError(f"{packet_path} already exists; pass --force to replace it")
      
          intent_rows = read_owned_csv(intent_path, PROJECT_INTENT_COLUMNS)
          manuscript_rows = read_owned_csv(manuscript_path, MANUSCRIPT_CONTRACT_COLUMNS)
          global_audit_rows = read_owned_csv(global_audit_path, GLOBAL_THESIS_AUDIT_COLUMNS)
          intent_needs_write = not intent_rows
          manuscript_needs_write = not manuscript_rows
          global_audit_needs_write = not global_audit_rows
          spine_rows = read_owned_csv(spine_path, SPINE_COLUMNS)
          contract_rows = read_owned_csv(contract_path, CONTRACT_COLUMNS)
          audit_rows = read_owned_csv(audit_path, AUDIT_COLUMNS)
          escalation_rows = read_owned_csv(escalation_path, ESCALATION_COLUMNS)
          validate_revision_attempt(
              contract_path,
              contract_rows,
              contract_id,
              revision_issue_id,
              attempt_no,
              args.force,
          )
      
          if not intent_rows:
              intent_id = args.intent_id or "pi-project-001"
              validate_identifier("intent_id", intent_id)
              intent_rows = [
                  {
                      "intent_id": intent_id,
                      "intent_version": "1",
                      "supersedes_intent_id": "",
                      "primary_domain": review_required("Name the manuscript's primary scholarly domain."),
                      "research_object": review_required("Define the object being studied or reviewed."),
                      "core_research_question": review_required("State the author-approved project-level research question."),
                      "target_venue": review_required("Name the intended venue or audience."),
                      "must_include_concepts": review_required("List concepts that must remain visible in the title or abstract."),
                      "excluded_reframes": review_required("List reframes that require a new approved intent version."),
                      "amendment_reason": review_required("Record that this is the initial intent contract."),
                      "approval_evidence": review_required("Record how and when the author approved this intent."),
                      "human_approved": "false",
                      "status": "draft",
                  }
              ]
          else:
              intent_candidates = [
                  row
                  for row in intent_rows
                  if args.intent_id
                  and row.get("intent_id", "").strip() == args.intent_id
              ]
              if not args.intent_id:
                  active = [row for row in intent_rows if row.get("status", "").strip().lower() == "active"]
                  intent_candidates = active or [
                      row for row in intent_rows if row.get("status", "").strip().lower() == "draft"
                  ]
              if len(intent_candidates) != 1:
                  raise ValueError("select exactly one current project intent with --intent-id")
              intent_id = intent_candidates[0].get("intent_id", "").strip()
              validate_identifier("intent_id", intent_id)
      
          if not manuscript_rows:
              manuscript_id = args.manuscript_id or "mc-project-001"
              validate_identifier("manuscript_id", manuscript_id)
              manuscript_rows = [
                  {
                      "manuscript_id": manuscript_id,
                      "intent_id": intent_id,
                      "manuscript_version": "1",
                      "supersedes_manuscript_id": "",
                      "title": review_required("Record the current manuscript title."),
                      "abstract_focus": review_required("Summarise the current abstract's primary focus."),
                      "primary_domain": review_required("Record the domain currently treated as primary."),
                      "research_object": review_required("Record the manuscript's current research object."),
                      "research_question": review_required("Record the manuscript's current research question."),
                      "contribution_scope": review_required("Record the current contribution boundary."),
                      "structure_summary": review_required("Summarise the manuscript's current section logic."),
                      "change_summary": review_required("Record that this is the initial manuscript contract."),
                      "human_approved": "false",
                      "status": "draft",
                  }
              ]
          else:
              manuscript_candidates = [
                  row
                  for row in manuscript_rows
                  if args.manuscript_id
                  and row.get("manuscript_id", "").strip() == args.manuscript_id
              ]
              if not args.manuscript_id:
                  active = [
                      row
                      for row in manuscript_rows
                      if row.get("status", "").strip().lower() == "active"
                      and row.get("intent_id", "").strip() == intent_id
                  ]
                  manuscript_candidates = active or [
                      row
                      for row in manuscript_rows
                      if row.get("status", "").strip().lower() == "draft"
                      and row.get("intent_id", "").strip() == intent_id
                  ]
              if len(manuscript_candidates) != 1:
                  raise ValueError("select exactly one current manuscript contract with --manuscript-id")
              manuscript_id = manuscript_candidates[0].get("manuscript_id", "").strip()
              validate_identifier("manuscript_id", manuscript_id)
      
          manuscript_version = next(
              row.get("manuscript_version", "").strip()
              for row in manuscript_rows
              if row.get("manuscript_id", "").strip() == manuscript_id
          )
          if not global_audit_rows:
              global_audit_id = args.global_audit_id or "ga-project-001"
              validate_identifier("global_audit_id", global_audit_id)
              global_audit_rows = [
                  {
                      "global_audit_id": global_audit_id,
                      "intent_id": intent_id,
                      "manuscript_id": manuscript_id,
                      "manuscript_version": manuscript_version,
                      "title_alignment": "not_assessed",
                      "abstract_alignment": "not_assessed",
                      "primary_domain_alignment": "not_assessed",
                      "research_object_alignment": "not_assessed",
                      "research_question_alignment": "not_assessed",
                      "contribution_alignment": "not_assessed",
                      "structure_alignment": "not_assessed",
                      "detected_reframe": "false",
                      "reframe_summary": review_required("Compare the manuscript contract with the approved project intent."),
                      "human_review_required": "true",
                      "human_decision": "pending",
                      "status": "needs_review",
                  }
              ]
          else:
              audit_candidates = [
                  row
                  for row in global_audit_rows
                  if args.global_audit_id
                  and row.get("global_audit_id", "").strip() == args.global_audit_id
              ]
              if not args.global_audit_id:
                  audit_candidates = [
                      row
                      for row in global_audit_rows
                      if row.get("intent_id", "").strip() == intent_id
                      and row.get("manuscript_id", "").strip() == manuscript_id
                      and row.get("manuscript_version", "").strip() == manuscript_version
                  ]
              if len(audit_candidates) != 1:
                  raise ValueError("select exactly one current global thesis audit with --global-audit-id")
              global_audit_id = audit_candidates[0].get("global_audit_id", "").strip()
              validate_identifier("global_audit_id", global_audit_id)
      
          spine_row = {
              "unit_id": unit_id,
              "manuscript_id": manuscript_id,
              "path": source_display,
              "section_title": section_title,
              "spine_sentence": args.spine_sentence or review_required("State the one-sentence argument spine for this unit."),
              "scope_boundary": args.scope_boundary or review_required("State what this unit is allowed to claim and what belongs elsewhere."),
              "core_claims": args.core_claims or review_required("List the claims that must survive the edit."),
              "do_not_change": args.do_not_change or review_required("List caveats, citations, scope limits, and terms that must not be changed."),
          }
          contract_row = {
              "contract_id": contract_id,
              "unit_id": unit_id,
              "global_audit_id": global_audit_id,
              "revision_issue_id": revision_issue_id,
              "attempt_no": str(attempt_no),
              "change_scope": args.change_scope or review_required("Specify exact paragraphs, lines, or local issue before editing."),
              "allowed_changes": args.allowed_changes or review_required("Specify what the edit may change."),
              "forbidden_changes": args.forbidden_changes or review_required("Specify claims, evidence, caveats, and boundaries the edit must preserve."),
              "adjacent_context": args.adjacent_context or review_required("Name neighbouring paragraphs or sections that must be checked."),
              "acceptance_checks": args.acceptance_checks or review_required("Define concrete checks for accepting, revising, or rolling back the edit."),
              "human_approved": "false",
              "status": "draft",
          }
      
          spine_rows, spine_action = upsert_row_candidate(
              spine_path, spine_rows, "unit_id", spine_row, args.force
          )
          contract_rows, contract_action = upsert_row_candidate(
              contract_path, contract_rows, "contract_id", contract_row, args.force
          )
          review_packet = render_review_packet(
              unit_id,
              section_title,
              source_display,
              intent_id,
              manuscript_id,
              global_audit_id,
              spine_row,
              contract_row,
          )
      
          validation = validate_packet(
              output_dir,
              strict=True,
              table_overrides={
                  "project_intent.csv": (PROJECT_INTENT_COLUMNS, intent_rows),
                  "manuscript_contracts.csv": (MANUSCRIPT_CONTRACT_COLUMNS, manuscript_rows),
                  "global_thesis_audits.csv": (GLOBAL_THESIS_AUDIT_COLUMNS, global_audit_rows),
                  "spine_cards.csv": (SPINE_COLUMNS, spine_rows),
                  "edit_contracts.csv": (CONTRACT_COLUMNS, contract_rows),
                  "drift_audits.csv": (AUDIT_COLUMNS, audit_rows),
                  "revision_escalations.csv": (ESCALATION_COLUMNS, escalation_rows),
              },
              prospective_files=[excerpt_file] if excerpt_file is not None else None,
          )
          if validation["issues"]:
              issue = validation["issues"][0]
              raise ValueError(
                  "candidate packet is not strict-valid: "
                  f"{issue['kind']} at {issue['location']}: {issue['message']}"
              )
      
          contents = {
              spine_path: render_csv_table(SPINE_COLUMNS, spine_rows),
              contract_path: render_csv_table(CONTRACT_COLUMNS, contract_rows),
              packet_path: review_packet.encode("utf-8"),
          }
          if intent_needs_write:
              contents[intent_path] = render_csv_table(PROJECT_INTENT_COLUMNS, intent_rows)
          if manuscript_needs_write:
              contents[manuscript_path] = render_csv_table(
                  MANUSCRIPT_CONTRACT_COLUMNS, manuscript_rows
              )
          if global_audit_needs_write:
              contents[global_audit_path] = render_csv_table(
                  GLOBAL_THESIS_AUDIT_COLUMNS, global_audit_rows
              )
          if not audit_path.exists():
              contents[audit_path] = render_csv_table(AUDIT_COLUMNS, audit_rows)
          if not escalation_path.exists():
              contents[escalation_path] = render_csv_table(ESCALATION_COLUMNS, escalation_rows)
          if excerpt_file is not None:
              contents[excerpt_file] = excerpt.encode("utf-8")
          atomic_write_batch(contents)
      
          return {
              "schema_version": 4,
              "output_dir": str(output_dir),
              "unit_id": unit_id,
              "contract_id": contract_id,
              "revision_issue_id": revision_issue_id,
              "attempt_no": attempt_no,
              "source": str(source),
              "source_recorded_as": source_display,
              "spine_cards": str(spine_path),
              "edit_contracts": str(contract_path),
              "drift_audits": str(audit_path),
              "revision_escalations": str(escalation_path),
              "project_intent": str(intent_path),
              "manuscript_contracts": str(manuscript_path),
              "global_thesis_audits": str(global_audit_path),
              "review_packet": str(packet_path),
              "actions": {
                  "spine_card": spine_action,
                  "edit_contract": contract_action,
                  "drift_audits": "ensured-header",
                  "revision_escalations": "ensured-header",
                  "project_intent": "ensured-current-contract",
                  "manuscript_contracts": "ensured-current-contract",
                  "global_thesis_audits": "ensured-current-audit",
              },
          }
      
      
      def build_parser() -> argparse.ArgumentParser:
          parser = argparse.ArgumentParser(
              description="Create a draft thesis-control packet for one manuscript unit."
          )
          parser.add_argument("project_root", nargs="?", default=".", help="Project root used to resolve --source")
          parser.add_argument("--source", required=True, help="Markdown source file, absolute or relative to project_root")
          parser.add_argument("--output-dir", help="Directory where thesis_control/ should be written; defaults to project_root")
          parser.add_argument("--unit-id", help="Stable unit id; defaults to source stem plus line range")
          parser.add_argument(
              "--contract-id",
              help="Stable contract id; defaults to ec-<unit-id>-<attempt-no padded to three digits>",
          )
          parser.add_argument(
              "--revision-issue-id",
              help="Stable issue id shared by contract versions for the same unresolved problem",
          )
          parser.add_argument("--intent-id", help="Project intent id to bind this section to")
          parser.add_argument("--manuscript-id", help="Manuscript contract id to bind this section to")
          parser.add_argument("--global-audit-id", help="Global thesis audit id that must gate this edit")
          parser.add_argument("--attempt-no", type=int, default=1, help="Positive attempt number within the revision issue")
          parser.add_argument("--section-title", help="Section title; defaults to first Markdown heading in the excerpt")
          parser.add_argument("--start-line", type=int, help="1-based start line for a source excerpt")
          parser.add_argument("--end-line", type=int, help="1-based end line for a source excerpt")
          parser.add_argument("--copy-source", action="store_true", help="Copy the selected excerpt into output-dir/source_excerpts/")
          parser.add_argument("--force", action="store_true", help="Replace existing rows or review packet with the same ids")
          parser.add_argument("--json", action="store_true", dest="emit_json", help="Emit JSON output")
      
          parser.add_argument("--spine-sentence", help="Concrete spine sentence for the unit")
          parser.add_argument("--scope-boundary", help="Concrete scope boundary for the unit")
          parser.add_argument("--core-claims", help="Concrete core claims that must survive editing")
          parser.add_argument("--do-not-change", help="Concrete do-not-change list")
          parser.add_argument("--change-scope", help="Concrete edit scope")
          parser.add_argument("--allowed-changes", help="Concrete allowed changes")
          parser.add_argument("--forbidden-changes", help="Concrete forbidden changes")
          parser.add_argument("--adjacent-context", help="Concrete adjacent context to inspect")
          parser.add_argument("--acceptance-checks", help="Concrete checks for accepting the edit")
          return parser
      
      
      def main(argv: Optional[Iterable[str]] = None) -> int:
          parser = build_parser()
          args = parser.parse_args(list(argv) if argv is not None else None)
      
          try:
              payload = scaffold(args)
          except (OSError, ValueError) as exc:
              print(f"error: {exc}", file=sys.stderr)
              return 1
      
          if args.emit_json:
              print(json.dumps(payload, indent=2))
          else:
              print(f"Thesis-control draft created: {payload['output_dir']}")
              print(f"- unit_id: {payload['unit_id']}")
              print(f"- contract_id: {payload['contract_id']}")
              print(f"- revision_issue_id: {payload['revision_issue_id']}")
              print(f"- attempt_no: {payload['attempt_no']}")
              print(f"- review packet: {payload['review_packet']}")
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • thesis_control_io.py 7.3 KB
      #!/usr/bin/env python3
      """Shared strict CSV and atomic file I/O for thesis-control scripts."""
      
      from __future__ import annotations
      
      import csv
      import io
      import os
      import stat
      import tempfile
      from collections import Counter
      from pathlib import Path
      from typing import Dict, List, Mapping, Sequence, Tuple
      
      
      class CsvShapeError(ValueError):
          """A structural CSV error with a stable issue kind and location."""
      
          def __init__(self, kind: str, location: str, message: str) -> None:
              super().__init__(message)
              self.kind = kind
              self.location = location
              self.message = message
      
      
      def read_csv_table(path: Path) -> Tuple[List[str], List[Dict[str, str]]]:
          """Read a CSV only when its header is unique and every row has full width."""
      
          try:
              with path.open(newline="", encoding="utf-8") as handle:
                  reader = csv.reader(handle, strict=True)
                  try:
                      fieldnames = next(reader)
                  except StopIteration as exc:
                      raise CsvShapeError("missing-header", str(path), "missing header row") from exc
      
                  if not fieldnames:
                      raise CsvShapeError("missing-header", str(path), "missing header row")
      
                  empty_columns = [index + 1 for index, name in enumerate(fieldnames) if not name.strip()]
                  if empty_columns:
                      positions = ", ".join(str(position) for position in empty_columns)
                      raise CsvShapeError(
                          "empty-column",
                          str(path),
                          f"header contains empty column name(s) at position(s): {positions}",
                      )
      
                  duplicates = sorted(name for name, count in Counter(fieldnames).items() if count > 1)
                  if duplicates:
                      raise CsvShapeError(
                          "duplicate-column",
                          str(path),
                          f"header contains duplicate column(s): {', '.join(duplicates)}",
                      )
      
                  rows: List[Dict[str, str]] = []
                  for row_number, values in enumerate(reader, start=2):
                      if len(values) != len(fieldnames):
                          raise CsvShapeError(
                              "row-width-mismatch",
                              f"{path}:row {row_number}",
                              f"row has {len(values)} cell(s); expected {len(fieldnames)}",
                          )
                      rows.append(dict(zip(fieldnames, values)))
          except csv.Error as exc:
              raise CsvShapeError("csv-parse-error", str(path), f"CSV parse error: {exc}") from exc
          except UnicodeError as exc:
              raise CsvShapeError(
                  "csv-decode-error",
                  str(path),
                  f"CSV is not valid UTF-8: {exc}",
              ) from exc
      
          return fieldnames, rows
      
      
      def ensure_internal_paths(root: Path, paths: Sequence[Path]) -> None:
          """Reject targets that escape the root or traverse an internal symlink."""
      
          root = root.resolve()
          for raw_path in paths:
              target = Path(os.path.abspath(str(raw_path)))
              try:
                  relative = target.relative_to(root)
              except ValueError as exc:
                  raise ValueError(f"output path escapes the packet root: {target}") from exc
      
              cursor = root
              for part in relative.parts:
                  cursor = cursor / part
                  if cursor.is_symlink():
                      raise ValueError(f"refusing internal symlink path: {cursor}")
      
              try:
                  target.resolve().relative_to(root)
              except (OSError, RuntimeError, ValueError) as exc:
                  raise ValueError(f"output path escapes the packet root: {target}") from exc
      
      
      def render_csv_table(fieldnames: Sequence[str], rows: Sequence[Mapping[str, str]]) -> bytes:
          """Render a validated CSV table with stable LF line endings."""
      
          buffer = io.StringIO(newline="")
          writer = csv.DictWriter(
              buffer,
              fieldnames=list(fieldnames),
              extrasaction="raise",
              lineterminator="\n",
          )
          writer.writeheader()
          writer.writerows(rows)
          return buffer.getvalue().encode("utf-8")
      
      
      def _stage_bytes(path: Path, content: bytes, mode: int) -> Path:
          descriptor, temporary_name = tempfile.mkstemp(
              prefix=f".{path.name}.",
              suffix=".tmp",
              dir=str(path.parent),
          )
          temporary = Path(temporary_name)
          try:
              with os.fdopen(descriptor, "wb") as handle:
                  handle.write(content)
                  handle.flush()
                  os.fsync(handle.fileno())
              os.chmod(temporary, mode)
          except Exception:
              temporary.unlink(missing_ok=True)
              raise
          return temporary
      
      
      def atomic_write_batch(contents: Mapping[Path, bytes]) -> None:
          """Stage and replace a set of files, restoring prior bytes on failure."""
      
          if not contents:
              return
      
          targets = sorted((Path(path), content) for path, content in contents.items())
          originals: Dict[Path, Tuple[bytes, int]] = {}
          created_directories: List[Path] = []
          staged: Dict[Path, Path] = {}
          replaced: List[Path] = []
      
          for target, _ in targets:
              if target.is_symlink():
                  raise ValueError(f"refusing to replace symlink target: {target}")
              if target.exists() and not target.is_file():
                  raise ValueError(f"output target is not a regular file: {target}")
              if target.exists():
                  originals[target] = (
                      target.read_bytes(),
                      stat.S_IMODE(target.stat().st_mode),
                  )
      
          try:
              known_directories = set()
              for target, _ in targets:
                  missing: List[Path] = []
                  parent = target.parent
                  while not parent.exists():
                      missing.append(parent)
                      parent = parent.parent
                  if not parent.is_dir():
                      raise ValueError(f"output parent is not a directory: {parent}")
                  for directory in reversed(missing):
                      if directory in known_directories:
                          continue
                      directory.mkdir()
                      created_directories.append(directory)
                      known_directories.add(directory)
      
              for target, content in targets:
                  mode = originals.get(target, (b"", 0o644))[1]
                  staged[target] = _stage_bytes(target, content, mode)
      
              for target, _ in targets:
                  os.replace(staged[target], target)
                  staged.pop(target, None)
                  replaced.append(target)
          except Exception as exc:
              rollback_errors = []
              for target in reversed(replaced):
                  try:
                      if target in originals:
                          original_bytes, original_mode = originals[target]
                          restore = _stage_bytes(target, original_bytes, original_mode)
                          os.replace(restore, target)
                      else:
                          target.unlink(missing_ok=True)
                  except Exception as rollback_exc:
                      rollback_errors.append(f"{target}: {rollback_exc}")
              for temporary in staged.values():
                  temporary.unlink(missing_ok=True)
              for directory in reversed(created_directories):
                  try:
                      directory.rmdir()
                  except OSError:
                      pass
              if rollback_errors:
                  details = "; ".join(rollback_errors)
                  raise OSError(f"{exc}; rollback failed: {details}") from exc
              raise
          finally:
              for temporary in staged.values():
                  temporary.unlink(missing_ok=True)
      
    • upgrade_thesis_control_project_intent.py 10.7 KB
      #!/usr/bin/env python3
      """Upgrade a schema-v3 thesis-control packet to a blocked schema-v4 draft.
      
      The helper never infers or approves scholarly intent. It adds the project-level
      link columns and creates AUTHOR_REVIEW_REQUIRED draft objects atomically. Any
      previously approved or applied edit remains blocked until the author completes
      and approves the new layer.
      """
      
      from __future__ import annotations
      
      import argparse
      import json
      import sys
      from pathlib import Path
      from typing import Dict, Iterable, List, Sequence, Tuple
      
      from check_thesis_control import validate_packet
      from project_intent_control import (
          GLOBAL_THESIS_AUDIT_COLUMNS,
          MANUSCRIPT_CONTRACT_COLUMNS,
          PROJECT_INTENT_COLUMNS,
      )
      from thesis_control_io import (
          atomic_write_batch,
          ensure_internal_paths,
          read_csv_table,
          render_csv_table,
      )
      
      
      SPINE_BASE_COLUMNS = [
          "unit_id",
          "path",
          "section_title",
          "spine_sentence",
          "scope_boundary",
          "core_claims",
          "do_not_change",
      ]
      CONTRACT_BASE_COLUMNS = [
          "contract_id",
          "unit_id",
          "revision_issue_id",
          "attempt_no",
          "change_scope",
          "allowed_changes",
          "forbidden_changes",
          "adjacent_context",
          "acceptance_checks",
          "human_approved",
          "status",
      ]
      AUDIT_COLUMNS = [
          "audit_id",
          "contract_id",
          "changed_claims",
          "changed_boundaries",
          "new_unsupported_claims",
          "missed_adjacent_updates",
          "drift_decision",
          "human_review_required",
          "status",
      ]
      ESCALATION_COLUMNS = [
          "escalation_id",
          "revision_issue_id",
          "escalation_kind",
          "trigger_contracts",
          "approved_after_attempt",
          "primary_category",
          "writing_scope",
          "valid_requirements",
          "missing_or_conflicting_information",
          "latest_author_approved_version",
          "recommended_next_action",
          "human_approved",
          "status",
      ]
      AUTHOR_REVIEW = "AUTHOR_REVIEW_REQUIRED"
      ALLOWED_BLOCKING_ISSUES = {
          "global-thesis-gate-required",
          "active-project-intent-required",
          "active-manuscript-contract-required",
      }
      
      
      def review_required(label: str) -> str:
          return f"{AUTHOR_REVIEW}: {label}"
      
      
      def read_required(path: Path, required: Sequence[str]) -> Tuple[List[str], List[Dict[str, str]]]:
          if not path.is_file():
              raise FileNotFoundError(f"missing required control file: {path}")
          fieldnames, rows = read_csv_table(path)
          missing = [column for column in required if column not in fieldnames]
          if missing:
              raise ValueError(f"{path} missing column(s): {', '.join(missing)}")
          return fieldnames, rows
      
      
      def insert_after(columns: Sequence[str], anchor: str, column: str) -> List[str]:
          output = list(columns)
          output.insert(output.index(anchor) + 1, column)
          return output
      
      
      def upgrade(project_root: Path) -> dict:
          project_root = project_root.expanduser().resolve()
          control_dir = project_root / "thesis_control"
          intent_path = control_dir / "project_intent.csv"
          manuscript_path = control_dir / "manuscript_contracts.csv"
          global_audit_path = control_dir / "global_thesis_audits.csv"
          spine_path = control_dir / "spine_cards.csv"
          contract_path = control_dir / "edit_contracts.csv"
          audit_path = control_dir / "drift_audits.csv"
          escalation_path = control_dir / "revision_escalations.csv"
          targets = [
              intent_path,
              manuscript_path,
              global_audit_path,
              spine_path,
              contract_path,
              audit_path,
              escalation_path,
          ]
          ensure_internal_paths(project_root, targets)
      
          spine_columns, spine_rows = read_required(spine_path, SPINE_BASE_COLUMNS)
          contract_columns, contract_rows = read_required(contract_path, CONTRACT_BASE_COLUMNS)
          audit_columns, audit_rows = read_required(audit_path, AUDIT_COLUMNS)
          escalation_columns, escalation_rows = read_required(escalation_path, ESCALATION_COLUMNS)
      
          new_files = [intent_path, manuscript_path, global_audit_path]
          file_presence = [path.exists() for path in new_files]
          has_spine_link = "manuscript_id" in spine_columns
          has_contract_link = "global_audit_id" in contract_columns
          if all(file_presence) and has_spine_link and has_contract_link:
              return {
                  "schema_version": 4,
                  "status": "unchanged",
                  "author_action_required": True,
                  "message": "project-intent layer already exists; validate and resolve its current author gates",
              }
          if any(file_presence) or has_spine_link or has_contract_link:
              raise ValueError(
                  "partial project-intent schema detected; no files were changed and an author decision is required"
              )
      
          manuscript_id = "mc-project-001"
          global_audit_id = "ga-project-001"
          intent_id = "pi-project-001"
          output_spine_columns = insert_after(spine_columns, "unit_id", "manuscript_id")
          output_contract_columns = insert_after(contract_columns, "unit_id", "global_audit_id")
          output_spine_rows = []
          for row in spine_rows:
              output = dict(row)
              output["manuscript_id"] = manuscript_id
              output_spine_rows.append(output)
          output_contract_rows = []
          for row in contract_rows:
              output = dict(row)
              output["global_audit_id"] = global_audit_id
              output_contract_rows.append(output)
      
          intent_rows = [
              {
                  "intent_id": intent_id,
                  "intent_version": "1",
                  "supersedes_intent_id": "",
                  "primary_domain": review_required("Name the manuscript's primary scholarly domain."),
                  "research_object": review_required("Define the project-level research object."),
                  "core_research_question": review_required("State the author-approved research question."),
                  "target_venue": review_required("Name the target venue or audience."),
                  "must_include_concepts": review_required("List concepts that must remain visible."),
                  "excluded_reframes": review_required("List reframes that need a new approved intent version."),
                  "amendment_reason": review_required("Record that this is the initial intent contract."),
                  "approval_evidence": review_required("Record explicit author approval."),
                  "human_approved": "false",
                  "status": "draft",
              }
          ]
          manuscript_rows = [
              {
                  "manuscript_id": manuscript_id,
                  "intent_id": intent_id,
                  "manuscript_version": "1",
                  "supersedes_manuscript_id": "",
                  "title": review_required("Record the current manuscript title."),
                  "abstract_focus": review_required("Summarise the current abstract focus."),
                  "primary_domain": review_required("Record the domain currently treated as primary."),
                  "research_object": review_required("Record the current research object."),
                  "research_question": review_required("Record the current research question."),
                  "contribution_scope": review_required("Record the current contribution boundary."),
                  "structure_summary": review_required("Summarise the current manuscript structure."),
                  "change_summary": review_required("Record that this is the initial manuscript contract."),
                  "human_approved": "false",
                  "status": "draft",
              }
          ]
          global_audit_rows = [
              {
                  "global_audit_id": global_audit_id,
                  "intent_id": intent_id,
                  "manuscript_id": manuscript_id,
                  "manuscript_version": "1",
                  "title_alignment": "not_assessed",
                  "abstract_alignment": "not_assessed",
                  "primary_domain_alignment": "not_assessed",
                  "research_object_alignment": "not_assessed",
                  "research_question_alignment": "not_assessed",
                  "contribution_alignment": "not_assessed",
                  "structure_alignment": "not_assessed",
                  "detected_reframe": "false",
                  "reframe_summary": review_required("Compare the current manuscript with the project intent."),
                  "human_review_required": "true",
                  "human_decision": "pending",
                  "status": "needs_review",
              }
          ]
      
          candidate = validate_packet(
              project_root,
              strict=False,
              table_overrides={
                  "project_intent.csv": (PROJECT_INTENT_COLUMNS, intent_rows),
                  "manuscript_contracts.csv": (MANUSCRIPT_CONTRACT_COLUMNS, manuscript_rows),
                  "global_thesis_audits.csv": (GLOBAL_THESIS_AUDIT_COLUMNS, global_audit_rows),
                  "spine_cards.csv": (output_spine_columns, output_spine_rows),
                  "edit_contracts.csv": (output_contract_columns, output_contract_rows),
                  "drift_audits.csv": (audit_columns, audit_rows),
                  "revision_escalations.csv": (escalation_columns, escalation_rows),
              },
          )
          unexpected = [
              issue for issue in candidate["issues"] if issue["kind"] not in ALLOWED_BLOCKING_ISSUES
          ]
          if unexpected:
              issue = unexpected[0]
              raise ValueError(
                  "candidate packet has a pre-existing or unrelated issue: "
                  f"{issue['kind']} at {issue['location']}: {issue['message']}"
              )
      
          atomic_write_batch(
              {
                  intent_path: render_csv_table(PROJECT_INTENT_COLUMNS, intent_rows),
                  manuscript_path: render_csv_table(MANUSCRIPT_CONTRACT_COLUMNS, manuscript_rows),
                  global_audit_path: render_csv_table(GLOBAL_THESIS_AUDIT_COLUMNS, global_audit_rows),
                  spine_path: render_csv_table(output_spine_columns, output_spine_rows),
                  contract_path: render_csv_table(output_contract_columns, output_contract_rows),
              }
          )
          return {
              "schema_version": 4,
              "status": "upgraded_blocked",
              "author_action_required": True,
              "project_intent": str(intent_path),
              "manuscript_contracts": str(manuscript_path),
              "global_thesis_audits": str(global_audit_path),
              "blocking_issue_count": len(candidate["issues"]),
              "blocking_issue_kinds": sorted({issue["kind"] for issue in candidate["issues"]}),
          }
      
      
      def main(argv: Iterable[str] | None = None) -> int:
          parser = argparse.ArgumentParser(
              description="Upgrade a schema-v3 thesis-control packet to a blocked schema-v4 project-intent draft."
          )
          parser.add_argument("project_root", nargs="?", default=".")
          parser.add_argument("--json", action="store_true", dest="emit_json")
          args = parser.parse_args(list(argv) if argv is not None else None)
          try:
              payload = upgrade(Path(args.project_root).expanduser().resolve())
          except (OSError, ValueError) as exc:
              print(f"error: {exc}", file=sys.stderr)
              return 1
          if args.emit_json:
              print(json.dumps(payload, indent=2))
          else:
              print(f"Project-intent upgrade: {payload['status']}")
              print("- author action required before approved or applied edits can pass strict validation")
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • upgrade_thesis_control_revision_tracking.py 14.2 KB
      #!/usr/bin/env python3
      """Upgrade a complete thesis-control packet to revision-tracking schema v3."""
      
      from __future__ import annotations
      
      import argparse
      import hashlib
      import json
      import sys
      from pathlib import Path
      from typing import Dict, Iterable, List, Optional, Sequence, Tuple
      
      from check_thesis_control import is_valid_identifier, validate_packet
      from thesis_control_io import (
          atomic_write_batch,
          ensure_internal_paths,
          read_csv_table,
          render_csv_table,
      )
      
      
      REVISION_COLUMNS = ["revision_issue_id", "attempt_no"]
      ESCALATION_V2_COLUMNS = [
          "escalation_id",
          "revision_issue_id",
          "trigger_contracts",
          "primary_category",
          "writing_scope",
          "valid_requirements",
          "missing_or_conflicting_information",
          "latest_author_approved_version",
          "recommended_next_action",
          "human_approved",
          "status",
      ]
      ESCALATION_V3_COLUMNS = [
          "escalation_id",
          "revision_issue_id",
          "escalation_kind",
          "trigger_contracts",
          "approved_after_attempt",
          "primary_category",
          "writing_scope",
          "valid_requirements",
          "missing_or_conflicting_information",
          "latest_author_approved_version",
          "recommended_next_action",
          "human_approved",
          "status",
      ]
      
      
      def default_revision_issue_id(contract_id: str) -> str:
          candidate = f"ri-{contract_id}"
          if len(candidate) <= 120:
              return candidate
          digest = hashlib.sha256(contract_id.encode("utf-8")).hexdigest()[:24]
          return f"ri-{digest}"
      
      
      def read_required_csv(path: Path, label: str) -> Tuple[List[str], List[Dict[str, str]]]:
          if not path.is_file():
              raise FileNotFoundError(f"missing {label}: {path}")
          return read_csv_table(path)
      
      
      def require_columns(path: Path, fieldnames: Sequence[str], required: Sequence[str]) -> None:
          missing = [column for column in required if column not in fieldnames]
          if missing:
              raise ValueError(f"{path} missing column(s): {', '.join(missing)}")
      
      
      def prepare_contracts(
          path: Path,
          fieldnames: Sequence[str],
          input_rows: Sequence[Dict[str, str]],
      ) -> Tuple[List[str], List[Dict[str, str]], int, bool]:
          require_columns(path, fieldnames, ["contract_id"])
          has_issue = "revision_issue_id" in fieldnames
          has_attempt = "attempt_no" in fieldnames
          if has_issue != has_attempt:
              raise ValueError(
                  f"{path} has a partial revision schema; an author decision is required "
                  "to establish historical issue grouping and attempt order"
              )
      
          rows = [dict(row) for row in input_rows]
          output_columns = list(fieldnames)
          legacy = not has_issue
          if legacy:
              insertion_point = (
                  output_columns.index("unit_id") + 1 if "unit_id" in output_columns else 1
              )
              for column in reversed(REVISION_COLUMNS):
                  output_columns.insert(insertion_point, column)
      
          contract_ids = set()
          attempts_by_issue: Dict[str, List[int]] = {}
          for row in rows:
              contract_id = row.get("contract_id", "").strip()
              if not contract_id:
                  raise ValueError("cannot upgrade a contract row without contract_id")
              if contract_id in contract_ids:
                  raise ValueError(f"duplicate contract_id requires an author decision: {contract_id}")
              contract_ids.add(contract_id)
      
              if legacy:
                  row["revision_issue_id"] = default_revision_issue_id(contract_id)
                  row["attempt_no"] = "1"
              else:
                  revision_issue_id = row.get("revision_issue_id", "").strip()
                  attempt_text = row.get("attempt_no", "").strip()
                  if not revision_issue_id or not attempt_text:
                      raise ValueError(
                          f"contract {contract_id} has partial revision values; an author decision "
                          "is required to establish historical issue grouping and attempt order"
                      )
      
              revision_issue_id = row.get("revision_issue_id", "").strip()
              if not is_valid_identifier(revision_issue_id):
                  raise ValueError(
                      f"contract {contract_id} has an unsafe revision_issue_id: {revision_issue_id}"
                  )
              attempt_text = row.get("attempt_no", "").strip()
              try:
                  attempt_no = int(attempt_text)
              except ValueError as exc:
                  raise ValueError(
                      f"contract {contract_id} attempt_no must be a positive integer: {attempt_text}"
                  ) from exc
              if attempt_no < 1:
                  raise ValueError(
                      f"contract {contract_id} attempt_no must be a positive integer: {attempt_text}"
                  )
              attempts_by_issue.setdefault(revision_issue_id, []).append(attempt_no)
      
          for revision_issue_id, attempts in attempts_by_issue.items():
              ordered = sorted(attempts)
              expected = list(range(1, len(ordered) + 1))
              if ordered != expected:
                  raise ValueError(
                      f"revision issue {revision_issue_id} attempts are missing, duplicate, or "
                      "non-sequential; an author decision is required"
                  )
      
          return output_columns, rows, len(rows) if legacy else 0, legacy
      
      
      def completed_failure_groups(
          contract_rows: Sequence[Dict[str, str]],
          audit_rows: Sequence[Dict[str, str]],
      ) -> Dict[str, List[Tuple[set, int, List[str]]]]:
          contracts: Dict[str, Tuple[str, int, str]] = {}
          for row in contract_rows:
              contract_id = row.get("contract_id", "").strip()
              contracts[contract_id] = (
                  row.get("revision_issue_id", "").strip(),
                  int(row.get("attempt_no", "").strip()),
                  row.get("status", "").strip().lower(),
              )
      
          unsuccessful = set()
          for row in audit_rows:
              contract_id = row.get("contract_id", "").strip()
              decision = row.get("drift_decision", "").strip().lower()
              status = row.get("status", "").strip().lower()
              contract = contracts.get(contract_id)
              if (
                  contract is not None
                  and contract[2] == "applied"
                  and decision in {"revise", "rollback"}
                  and status == "failed"
              ):
                  unsuccessful.add(contract_id)
      
          ordered_by_issue: Dict[str, List[Tuple[int, str]]] = {}
          for contract_id in unsuccessful:
              revision_issue_id, attempt_no, _ = contracts[contract_id]
              ordered_by_issue.setdefault(revision_issue_id, []).append((attempt_no, contract_id))
      
          groups: Dict[str, List[Tuple[set, int, List[str]]]] = {}
          for revision_issue_id, attempts in ordered_by_issue.items():
              ordered = sorted(attempts)
              for start in range(0, (len(ordered) // 3) * 3, 3):
                  group = ordered[start : start + 3]
                  groups.setdefault(revision_issue_id, []).append(
                      (
                          {contract_id for _, contract_id in group},
                          max(attempt for attempt, _ in group),
                          [contract_id for _, contract_id in group],
                      )
                  )
          return groups
      
      
      def classify_v2_escalation(
          path: Path,
          row: Dict[str, str],
          contracts: Dict[str, Tuple[str, int]],
          failure_groups: Dict[str, List[Tuple[set, int, List[str]]]],
      ) -> None:
          escalation_id = row.get("escalation_id", "").strip() or "<missing>"
          revision_issue_id = row.get("revision_issue_id", "").strip()
          triggers = [
              value.strip()
              for value in row.get("trigger_contracts", "").split(";")
              if value.strip()
          ]
          trigger_set = set(triggers)
          prefix = f"{path} escalation {escalation_id} has ambiguous v2 data:"
      
          if not triggers or len(triggers) != len(trigger_set):
              raise ValueError(f"{prefix} trigger contracts must be non-empty and unique")
          if len(triggers) > 3:
              raise ValueError(f"{prefix} oversized trigger sets cannot be classified")
          if revision_issue_id not in {value[0] for value in contracts.values()}:
              raise ValueError(f"{prefix} unknown revision_issue_id {revision_issue_id}")
          for trigger in triggers:
              contract = contracts.get(trigger)
              if contract is None:
                  raise ValueError(f"{prefix} unknown trigger contract {trigger}")
              if contract[0] != revision_issue_id:
                  raise ValueError(f"{prefix} trigger contracts cross revision issues")
      
          if len(triggers) <= 2:
              row["escalation_kind"] = "early_diagnostic"
              row["approved_after_attempt"] = ""
              return
      
          matching_group = next(
              (
                  (group, boundary, ordered)
                  for group, boundary, ordered in failure_groups.get(revision_issue_id, [])
                  if group == trigger_set
              ),
              None,
          )
          if matching_group is None:
              raise ValueError(
                  f"{prefix} three triggers do not match one completed unsuccessful group"
              )
          row["escalation_kind"] = "cycle_gate"
          row["approved_after_attempt"] = str(matching_group[1])
          row["trigger_contracts"] = ";".join(matching_group[2])
      
      
      def prepare_escalations(
          path: Path,
          contract_rows: Sequence[Dict[str, str]],
          audit_rows: Sequence[Dict[str, str]],
      ) -> Tuple[List[str], List[Dict[str, str]], str, bool]:
          if not path.exists():
              return list(ESCALATION_V3_COLUMNS), [], "created", True
          if not path.is_file():
              raise FileNotFoundError(f"revision escalation target is not a file: {path}")
      
          fieldnames, input_rows = read_csv_table(path)
          has_kind = "escalation_kind" in fieldnames
          has_boundary = "approved_after_attempt" in fieldnames
          if has_kind != has_boundary:
              raise ValueError(
                  f"{path} has a partial escalation schema; an author decision is required"
              )
      
          if has_kind:
              require_columns(path, fieldnames, ESCALATION_V3_COLUMNS)
              return list(fieldnames), [dict(row) for row in input_rows], "unchanged", False
      
          require_columns(path, fieldnames, ESCALATION_V2_COLUMNS)
          output_columns = list(fieldnames)
          trigger_index = output_columns.index("trigger_contracts")
          output_columns.insert(trigger_index, "escalation_kind")
          trigger_index = output_columns.index("trigger_contracts")
          output_columns.insert(trigger_index + 1, "approved_after_attempt")
      
          rows = [dict(row) for row in input_rows]
          contracts = {
              row.get("contract_id", "").strip(): (
                  row.get("revision_issue_id", "").strip(),
                  int(row.get("attempt_no", "").strip()),
              )
              for row in contract_rows
          }
          failure_groups = completed_failure_groups(contract_rows, audit_rows)
          for row in rows:
              classify_v2_escalation(path, row, contracts, failure_groups)
          return output_columns, rows, "upgraded", True
      
      
      def validate_candidate(
          project_root: Path,
          spine_table: Tuple[Sequence[str], Sequence[Dict[str, str]]],
          contract_table: Tuple[Sequence[str], Sequence[Dict[str, str]]],
          audit_table: Tuple[Sequence[str], Sequence[Dict[str, str]]],
          escalation_table: Tuple[Sequence[str], Sequence[Dict[str, str]]],
      ) -> None:
          payload = validate_packet(
              project_root,
              strict=True,
              require_project_intent=False,
              table_overrides={
                  "spine_cards.csv": spine_table,
                  "edit_contracts.csv": contract_table,
                  "drift_audits.csv": audit_table,
                  "revision_escalations.csv": escalation_table,
              },
          )
          if payload["issues"]:
              issue = payload["issues"][0]
              raise ValueError(
                  "candidate packet is not strict-valid: "
                  f"{issue['kind']} at {issue['location']}: {issue['message']}"
              )
      
      
      def upgrade(project_root: Path) -> dict:
          control_dir = project_root / "thesis_control"
          spine_path = control_dir / "spine_cards.csv"
          contract_path = control_dir / "edit_contracts.csv"
          audit_path = control_dir / "drift_audits.csv"
          escalation_path = control_dir / "revision_escalations.csv"
      
          ensure_internal_paths(
              project_root,
              [spine_path, contract_path, audit_path, escalation_path],
          )
      
          contract_fields, contract_input_rows = read_required_csv(
              contract_path, "edit contracts"
          )
          contract_columns, contract_rows, upgraded, contract_changed = prepare_contracts(
              contract_path, contract_fields, contract_input_rows
          )
      
          spine_fields, spine_rows = read_required_csv(spine_path, "spine cards")
          audit_fields, audit_rows = read_required_csv(audit_path, "drift audits")
          (
              escalation_columns,
              escalation_rows,
              escalation_action,
              escalation_changed,
          ) = prepare_escalations(escalation_path, contract_rows, audit_rows)
      
          validate_candidate(
              project_root,
              (spine_fields, spine_rows),
              (contract_columns, contract_rows),
              (audit_fields, audit_rows),
              (escalation_columns, escalation_rows),
          )
      
          contents = {}
          if contract_changed:
              contents[contract_path] = render_csv_table(contract_columns, contract_rows)
          if escalation_changed:
              contents[escalation_path] = render_csv_table(escalation_columns, escalation_rows)
          atomic_write_batch(contents)
      
          return {
              "schema_version": 3,
              "project_root": str(project_root),
              "edit_contracts": str(contract_path),
              "revision_escalations": str(escalation_path),
              "contracts_upgraded": upgraded,
              "escalation_file": escalation_action,
          }
      
      
      def build_parser() -> argparse.ArgumentParser:
          parser = argparse.ArgumentParser(
              description="Upgrade a complete thesis-control packet to strict revision schema v3."
          )
          parser.add_argument(
              "project_root", nargs="?", default=".", help="Packet root containing thesis_control/"
          )
          parser.add_argument("--json", action="store_true", dest="emit_json", help="Emit JSON output")
          return parser
      
      
      def main(argv: Optional[Iterable[str]] = None) -> int:
          parser = build_parser()
          args = parser.parse_args(list(argv) if argv is not None else None)
          try:
              payload = upgrade(Path(args.project_root).expanduser().resolve())
          except (OSError, ValueError) as exc:
              print(f"error: {exc}", file=sys.stderr)
              return 1
      
          if args.emit_json:
              print(json.dumps(payload, indent=2))
          else:
              print(f"Thesis-control revision tracking upgraded: {payload['project_root']}")
              print(f"- contracts upgraded: {payload['contracts_upgraded']}")
              print(f"- escalation file: {payload['escalation_file']}")
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • SKILL.md 19.8 KB
    ---
    name: thesis-control
    description: Use when AI-assisted thesis or manuscript edits risk claim drift, scope creep, loss of intended use, experiment-role promotion, or repeated revisions that fail to converge; provides author-intent control, lightweight or strict contracts, drift audits, revision escalation, and human gates.
    allowed-tools: Read, Glob, Grep, Edit, Write, Bash
    ---
    
    # /thesis-control - Thesis Drift Control
    
    ## Purpose
    
    Prevent AI-assisted writing from becoming fluent but distorted. Use this before and after substantive thesis or manuscript edits when the risk is not spelling or style, but loss of author control: project-level reframing, deletion of the real-world task or intended use, a primary domain becoming a secondary example, a changed research object or question, auxiliary analyses becoming primary, widened claims, blurred section purpose, missing caveats, unsynchronised adjacent paragraphs, or local edits that weaken the paper spine.
    
    ## Trigger Words
    
    This skill activates on: `thesis control`, `drift audit`, `edit contract`, `spine card`, `claim drift`, `author control`, `loss of control`, `scope creep`, `rewrite risk`, `/thesis-control`.
    
    ## Core Rule
    
    Do not edit thesis prose until the project intent, current manuscript contract,
    global thesis audit, section spine, and intended local change form one explicit
    and traceable contract chain.
    
    The contract must answer:
    
    ```text
    This edit is allowed to change [specific local issue] in [specific unit], while preserving [spine sentence], [scope boundary], [core claims], and [do-not-change items].
    ```
    
    If this sentence cannot be written, stop and diagnose the section instead of rewriting it.
    
    The control hierarchy is:
    
    ```text
    Author-approved Project Intent
    → Author-approved Manuscript Contract
    → Passed Global Thesis Audit
    → Section Spine Card
    → Edit Contract
    → Post-edit Drift Audit
    ```
    
    A lower layer cannot amend a higher one. If the title, abstract, primary
    domain, research object, research question, contribution scope, or manuscript
    structure no longer matches the approved intent, stop. Revise or roll back the
    manuscript, or create a new explicitly approved intent version that preserves
    the earlier row as history.
    
    ## Control Files
    
    Choose one control profile and name it as canonical.
    
    For a lightweight single-manuscript workflow, read
    `references/author_control_lightweight.md` and use:
    
    - `00_AUTHOR_INTENT.md`
    - `01_EVIDENCE_AND_CLAIMS.md`
    - `02_REVISION_LOG.md`
    
    Create and check the bundled templates with:
    
    ```bash
    python scripts/scaffold-author-control.py <project_root>
    python scripts/check-author-control.py <project_root> --strict
    ```
    
    The lightweight checker validates structure, approval state, and unresolved
    placeholders. It does not infer semantic alignment. Do not use the lightweight
    profile to bypass a gate that already requires the durable packet.
    
    Use a `thesis_control/` directory when the project needs durable tracking:
    
    - `project_intent.csv`
    - `manuscript_contracts.csv`
    - `global_thesis_audits.csv`
    - `spine_cards.csv`
    - `edit_contracts.csv`
    - `drift_audits.csv`
    - `revision_escalations.csv`
    
    Run the optional validator when Python is available:
    
    ```bash
    python {skill_dir}/scripts/check_thesis_control.py <project_root> --strict
    ```
    
    Strict validation requires the project-intent layer. It blocks approved or
    applied edit contracts unless they reference a passed global thesis audit for
    the active author-approved intent and manuscript contract. Non-strict mode can
    still inspect legacy packets that do not yet have this layer.
    
    The validator checks packet structure and recorded gate consistency. It does
    not infer semantic alignment or judge scholarly truth. The author or reviewer
    must compare the manuscript with the intent and record each alignment field
    honestly; the validator then prevents an unresolved or drifted audit from being
    used as authorisation.
    
    Strict validation requires revision-tracking schema v3. Upgrade a complete
    legacy packet without guessing historical revision families:
    
    ```bash
    python {skill_dir}/scripts/upgrade_thesis_control_revision_tracking.py <project_root>
    ```
    
    To create a draft packet from a real Markdown unit before editing prose:
    
    ```bash
    python {skill_dir}/scripts/scaffold_thesis_control.py <project_root> \
      --source chapters/ch1_introduction.md \
      --start-line 71 \
      --end-line 104 \
      --revision-issue-id ri-ch1-gap-clarity \
      --attempt-no 1 \
      --copy-source
    ```
    
    The scaffold writes schema v4 draft project-intent and manuscript contracts, a
    pending global thesis audit, `human_approved=false`, `status=draft`, and
    `AUTHOR_REVIEW_REQUIRED` fields. Replace those fields with concrete author
    judgement before applying a substantive edit. A scaffolded packet may be
    structurally valid while remaining non-executable. Its default contract id
    includes the attempt number, so attempts 1 and 2 become `ec-<unit>-001` and
    `ec-<unit>-002`. Reuse an explicit `revision_issue_id` for retries.
    
    The migration helper stops without writing when revision metadata is partial or
    when a legacy escalation cannot be classified from current contracts and
    resolved audits. It preserves named extension columns and converts one- or
    two-trigger legacy rows to `early_diagnostic`; a three-trigger row becomes a
    `cycle_gate` only when it already matches one completed failure group.
    
    Upgrade a complete schema-v3 packet into a deliberately blocked schema-v4
    draft without guessing author intent:
    
    ```bash
    python {skill_dir}/scripts/upgrade_thesis_control_project_intent.py \
      <project_root> --json
    ```
    
    The helper adds `manuscript_id` and `global_audit_id` links, preserves named
    extension columns, and creates `AUTHOR_REVIEW_REQUIRED` draft intent,
    manuscript, and global-audit rows through one atomic batch. Previously approved
    or applied edits remain blocked. Replace the draft fields with real author
    judgement, approve the active intent and manuscript contract, and resolve the
    global audit before strict validation can pass. A partial project-intent schema
    stops without mutation.
    
    ## Workflow
    
    ### 0. Establish The Project Intent And Manuscript Contract
    
    Before section-level planning, record:
    
    - the real-world problem and intended user or beneficiary
    - the intended application and the present method or software task
    - the primary scholarly domain
    - the research object
    - the core research question
    - the primary experiment that directly answers that question
    - supporting, robustness, exploratory, failed-development, and out-of-scope analyses
    - the strongest evidence-licensed headline claim
    - the current validation and evidence boundaries
    - the target venue or audience
    - concepts that must remain visible in the title or abstract
    - reframes that require fresh author approval
    - the current title, abstract focus, contribution scope, and structure
    - concrete approval evidence and the active version ids
    
    Keep one active author-approved project intent and one active author-approved
    manuscript contract. A later intent version must identify the immediately
    previous version in `supersedes_intent_id`, record the amendment reason, and
    leave the earlier version as `superseded`. Do not overwrite the original row.
    
    Run a global thesis audit whenever the title, abstract, primary domain,
    research object, research question, contribution scope, or overall structure
    changes. Record each dimension as `aligned`, `drifted`, or `not_assessed`.
    Only a fully aligned audit with `detected_reframe=false` can have
    `status=passed` and `human_decision=accept`.
    
    If any dimension is `drifted`, set `human_review_required=true` and use
    `needs_review` or `failed`. The author must choose to revise the manuscript,
    roll back, or approve a versioned intent amendment. Merely accepting the audit
    cannot authorise the reframe.
    
    Keep intended use and current validation separate. Narrow evidence may narrow
    the empirical or headline claim, but it must not silently delete a legitimate
    application problem or recast the evidence boundary as the paper's research
    object. Record future application as intended use and untested hardware,
    clinical, causal, deployment, or transfer outcomes as unvalidated boundaries.
    
    Before admitting a completed analysis into the paper, record whether it
    directly answers the core question, whether the main conclusion survives its
    removal, its one-sentence argumentative function, its role, destination, and
    author decision. Completion alone does not make an analysis a main
    contribution. Use the role and placement defaults in
    `references/author_control_lightweight.md`.
    
    ### 1. Establish Or Read The Spine Card
    
    Before editing a chapter, section, or paragraph cluster, identify:
    
    - unit id
    - source path
    - section title
    - spine sentence
    - scope boundary
    - core claims
    - do-not-change items
    - the active manuscript contract id
    
    The spine sentence should be narrow:
    
    ```text
    This unit argues that [specific claim] by showing [specific basis], so the chapter can [specific function].
    ```
    
    If the current text does not support a clear spine sentence, produce a diagnosis and ask for author direction before changing prose.
    
    ### 2. Create The Edit Contract
    
    For every substantive edit, state:
    
    - target unit and file range
    - change scope: `local_patch`, `section_restructure`, or `full_reframe`
    - allowed changes
    - forbidden changes
    - evidence baseline: the ref, artifact, data, configuration, table, or frozen numbers inherited
    - argument baseline: the author-approved intent and manuscript version inherited
    - adjacent context that must be checked
    - acceptance checks
    - whether human approval is required before editing
    - the passed global thesis audit id that covers the spine card's manuscript contract
    
    When using the scaffold helper, treat its output as a draft control packet, not
    as approval. A generated contract becomes actionable only after the author has
    replaced the `AUTHOR_REVIEW_REQUIRED` fields and explicitly approved the scope.
    
    Always require human approval for:
    
    - changing the section spine
    - adding or broadening claims
    - deleting caveats or limitations
    - moving evidence between sections
    - rewriting more than one paragraph
    - merging or splitting sections
    - changing the title, abstract, primary domain, research object, research question, contribution scope, or manuscript structure
    - changing the real-world problem, intended use, primary experiment, analysis prominence, headline claim, or evidence boundary
    
    Treat any change to the title, abstract thesis, research object, core question,
    primary experiment, contribution order, evidence chain, application purpose,
    or paper-wide structure as a `full_reframe`, even when the request calls it
    polishing. Before editing, show an old-versus-proposed spine comparison and
    obtain explicit author approval.
    
    ### 3. Apply Only Approved Changes
    
    After approval, edit only the approved scope. Do not apply the edit if the
    linked global thesis audit is pending, failed, stale, drifted, or attached to a
    different manuscript contract.
    
    Keep mechanical fixes separate from argument changes. Do not bundle style, structure, evidence, and claim changes into one patch unless the contract explicitly allows it.
    
    ### 4. Run The Drift Audit
    
    After editing, compare the new prose against the contract and report:
    
    - changed claims
    - changed boundaries or caveats
    - new unsupported claims
    - deleted evidence anchors
    - missed adjacent updates
    - section-spine change
    - research-object or core-question change
    - deleted, generalised, or demoted real-world task or intended use
    - promoted auxiliary analysis
    - evidence boundary rewritten as the paper topic
    - loss of application meaning caused by over-cautious wording
    - title, abstract, Introduction, Results, and Conclusion alignment
    - decision: accept, partial accept, revise, or rollback
    
    If any claim, boundary, or caveat changed, the result needs human review even if the prose is smoother.
    
    Use audit `status=needs_review` only while the author's post-edit decision is
    pending. Strict validation blocks an applied contract in that state. After the
    author decides, record `status=passed` for `accept` or `partial_accept`, and
    `status=failed` for `revise` or `rollback`. Do not treat a pending audit as a
    completed unsuccessful attempt.
    
    ### 5. Record Human Gate Outcome
    
    The author decides whether to accept, partially accept, revise, or rollback. Do not mark a high-risk edit as accepted without explicit human approval.
    
    ### 6. Run Post-Spine Readability Gates When Relevant
    
    Only after the research spine is stable, use `/logic-review` to audit repeated
    argument functions and `/self-review` to prepare the unfamiliar-reader packet.
    Do not solve repetitive AI prose by generating synonyms. Remove duplicated
    problem, gap, evidence, interpretation, or boundary functions while preserving
    essential local qualifiers. A model simulation cannot pass a human unfamiliar-
    reader gate; record it as advisory or `not_run` until an actual reader responds.
    
    ## Revision Escalation Rule
    
    Treat three unsuccessful attempts on the same revision issue as an operational escalation threshold, not as evidence that every task fails after three turns. Use `revision_issue_id` to keep successive contract versions attached to that issue. Only count an attempt when its drift decision is `revise` or `rollback` and its audit status is `failed`. Only applied contracts count as unsuccessful attempts. Record author rejection as one of those decisions. Multiple failed audits of one contract still count as one attempt; contradictory passed and failed resolved audits are invalid. Clarifying discussion, pending human reviews, and unexecuted proposals do not count.
    
    After three unsuccessful attempts, stop. Do not apply a fourth prose patch. Record a row in `revision_escalations.csv`; a later contract may become `approved` or `applied` only after the matching escalation has `human_approved=true` and `status=approved`.
    
    An approved escalation closes only that group of three unsuccessful contracts. If three later contracts also receive `revise` or `rollback`, require a new escalation before another contract can proceed.
    
    Only a `cycle_gate` whose three triggers exactly match one completed group of
    unsuccessful contracts, in attempt order, may close that group. Set
    `approved_after_attempt` to the final attempt number in that group. The gate is
    effective only with `human_approved=true` and `status=approved`. One gate cannot
    close more than one group. Do not repeat a trigger contract within a row or
    create multiple rows for the same issue and trigger set.
    
    Only an escalation whose trigger set exactly matches one completed group of three unsuccessful contracts may close that group. One escalation cannot close more than one group.
    
    Record an earlier warning as `early_diagnostic` with one or two unique triggers
    and an empty `approved_after_attempt`. It may be author-approved as a diagnosis,
    but it never closes or pre-authorises a later completed group.
    
    An earlier escalation with fewer than three trigger contracts does not close or pre-authorise a later completed group.
    
    Escalate earlier than three attempts when any of these signals is already visible:
    
    - the section spine cannot be stated consistently
    - the requested claim lacks supporting evidence
    - a revision changes a claim, caveat, or scope boundary outside the contract
    - the latest author-approved version cannot be identified
    - old assumptions, duplicated explanations, or conflicting requirements indicate version contamination
    
    ### Required Escalation Check
    
    Before editing again:
    
    1. Consolidate the currently valid requirements into one brief.
    2. Compare that brief with the spine card, evidence boundaries, current contract, and latest author-approved version.
    3. Classify the failure as one primary category:
       - **underspecified or conflicting intent** — the target, audience, venue, constraint, or acceptance condition is missing or inconsistent, or the feedback is evaluative but not operational, such as “weak”, “unclear”, or “still not right” without a concrete change target
       - **local execution failure** — the contract is clear, but the edit did not implement it correctly
       - **structural mismatch** — the problem affects the section purpose, research question, gap, contribution, evidence chain, or manuscript structure
       - **evidence gap** — the requested claim is not supported by the available sources, data, experiments, or files
       - **version contamination** — accumulated patches mix incompatible assumptions, duplicate reasoning, or obscure which prose the author approved
    4. Classify the writing scope:
       - **local patch** — wording or presentation changes that preserve the spine, claims, evidence, and adjacent-section relationships
       - **section-level restructure** — changes confined to one section without changing the research question, contribution, or evidence chain
       - **full reframing** — changes to the title, abstract, research question, gap, contribution, methods-results alignment, evidence chain, or discussion framing
    5. Recommend the smallest valid next action and wait for author approval.
    
    Use these default actions:
    
    - For a local execution failure, create a corrected local contract.
    - For underspecified or conflicting intent, ask for the missing decision before editing.
    - For a structural mismatch, propose a section-level restructure or full reframing plan before editing.
    - For an evidence gap, narrow, qualify, or remove the unsupported claim unless the author supplies more evidence.
    - For version contamination, restore or copy the latest author-approved version, then apply a consolidated contract. Create a separate branch or manuscript version only when the approved scope requires structural work.
    
    For full reframing, hand off a brief that states the target venue, old and
    proposed real-world problem, intended use, research object, research question,
    primary experiment, contribution order, headline claim, evidence boundary,
    evidence baseline, argument baseline, available evidence, claims that must not
    be made, and proposed new structure. Do not rewrite the manuscript until the
    author approves that brief.
    
    Return the escalation check in this form:
    
    ```text
    ## Revision Escalation Check
    
    Revision issue:
    Contract:
    Unsuccessful attempts:
    Trigger contracts:
    Primary category:
    Writing scope:
    Why the revisions did not converge:
    Valid requirements:
    Missing or conflicting information:
    Latest author-approved version:
    Recommended next action:
    Author decision required:
    ```
    
    ## Output Patterns
    
    ### Audit Only
    
    Return:
    
    - current spine diagnosis
    - likely drift risks
    - control gaps
    - recommended edit contracts
    - blocked items needing author decision
    
    ### Pre-Edit Contract
    
    Return:
    
    ```text
    ## Edit Contract
    
    Unit:
    Spine sentence:
    Scope: local_patch / section_restructure / full_reframe
    Evidence baseline:
    Argument baseline:
    Allowed changes:
    Forbidden changes:
    Adjacent context to check:
    Acceptance checks:
    Human approval required:
    Proceed only after approval:
    ```
    
    ### Post-Edit Drift Audit
    
    Return:
    
    ```text
    ## Drift Audit
    
    Contract:
    Changed claims:
    Changed boundaries:
    New unsupported claims:
    Missed adjacent updates:
    Research object or question changed:
    Intended use deleted or demoted:
    Auxiliary analysis promoted:
    Cross-section paper identity aligned:
    Decision:
    Human review required:
    Recommended next action:
    ```
    
    ## Stop Conditions
    
    Stop and ask for author direction if:
    
    - the section spine cannot be stated clearly
    - the requested edit would broaden a claim without evidence
    - a local edit requires adjacent updates outside the approved scope
    - the user asks for a full-chapter rewrite without a spine map
    - previous AI edits cannot be distinguished from author-approved text
    - the edit would remove caveats, limitations, or uncertainty language without explicit approval
    - the active project intent or manuscript contract cannot be identified
    - the real-world task, intended use, primary experiment, evidence baseline, or argument baseline cannot be identified
    - a global thesis audit is missing, unresolved, stale, or records project-level drift
    - a proposed local contract would preserve a section spine that conflicts with the author-approved project intent
    - a full reframe lacks an approved old-versus-proposed spine comparison
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related