Claude Skill

scholarly-citation

Verifiable-attribution and citation-discipline craft — atomic claim decomposition, evidence grading (primary/analysis/forecast tiers), claimant attribution, citation typing (cites/references/contradicts/defines), source anchoring (Xanadu). Use when writing or reviewing a referenc

LLM Mart · 0 points · 8 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download alfadur7-llm-wiki-newsroom-.claude_skills_scholarly-citation-d518712.zip · 14 KB
Part of alfadur7/llm-wiki-newsroom — 6 skills

Install

skills CLI npx skills add https://github.com/alfadur7/llm-wiki-newsroom/tree/main/.claude/skills/scholarly-citation
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install alfadur7-llm-wiki-newsroom@llmmart
Git git clone https://github.com/alfadur7/llm-wiki-newsroom.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole alfadur7/llm-wiki-newsroom collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

scholarly-citation

Verifiable-attribution craft drawn from scholarly/scientific citation traditions (citation classification, typing, anchoring). It ensures every claim in this wiki is atomized, evidence-graded, attributed to a speaker, typed by citation meaning, and anchored to a source location — and therefore traceable. Unlike the three borrowed traditions (narrative, exec summary, encyclopedic neutrality), this craft synthesizes external sources into an atomic-claim schema and is the project's signature. criteria.json is the SoT for each criterion's definition, comparator, and source; wiki-global state (page existence·hub type·section titles) is injected by the orchestrator (tools/_lint/source.py). The schema literals (## Key Claims, [fact], cites: etc.) are the wiki's actual schema markers and section headers.

Claim atomization & grading (cit.grade-marker · cit.atomic · cit.composite-split)

One claim, one line, one utterance. Each line of ## Key Claims (Key Claims) marks an evidence grade — [fact] (primary source), [analysis] (secondary analysis), [forecast] (forecast). Putting two subjects' claims on one line (joined by "and"/"&" or a comma series spanning two hubs; ko: the conjunctions 와/및), or chaining verb phrases on one line ("…did, and …did"; ko: ~했고, ~했다), breaks atomicity.

Claimant attribution (cit.claimant-link · cit.claimant-valid · cit.speaker-link)

Every claim is bound to who said it. The slot right after the grade marker names the speaker — a [[wikilink]] when they have a page, their plain-text real name and role when they fall below the page-creation threshold. What WP:ASF requires is the named source; the wikilink is one means to it, so demanding a link from a speaker who cannot have a page only pushes the author to link some other entity. Anonymous or collective subjects ("the government"·"industry sources"; ko 정부는·업계에서는) stay forbidden in either form ("avoid stating opinions as facts"; attribute them to a named source), as does naming in plain text a page that does exist. A blockquote names its speaker (APA author-date principle — identify the cited author), wikilinked when that speaker has a page. A speaker with no page stays plain text with their role spelled out. Any target that is linked must exist as a real page.

Citation typing (cit.cite-type · cit.cite-distribution · cit.cite-type-hub)

Each line of ## Connections (Connections) types the citation's meaning — cites: (direct grounds)·references: (context)·contradicts: (dispute)·defines: (definition). This extends scite Smart Citations' Supporting / Mentioning / Contrasting classes (with defines: added). If every line collapses to the default-safe references:, that signals avoidance of meaning classification. defines: points only to a concept hub.

Source anchoring (cit.anchor · cit.anchor-valid)

An external-evidence citation anchors to the source's section — [[slug#section]] (ko: [[slug#섹션]]). This applies Xanadu citation anchoring (visible links to the precise source location, Ted Nelson) at the source level; a gradual migration, so 0% still passes (advisory), but the anchor target's slug and section must exist.

Synthesis citation integrity (cit.cite-consistency · cit.grounding · cit.anchor-evidence)

A piece synthesizing many sources follows Select→Read→Cite discipline — the declared source set (frontmatter) is the top set, and both the sources read and the sources cited in the body are subsets of it.

  • cit.cite-consistency — every source cited in the body is registered in the declared source list (body ⊆ declaration). Citing an undeclared source is an editorial slip (omission).
  • cit.grounding — a source offered as representative evidence must actually be read and used. When at least one of the source's direct-quote / key-claim items appears in the body as a substring, that is objective evidence of having read it (writing from a one-line summary by memory/guess risks divergence from the source). The deterministic lint blocks only the "never read at all" extreme; full read-through rests on editor self-discipline.
  • cit.anchor-evidence — when a direct quote is placed in representative evidence, mark the citation location with a source section anchor ([[slug#section]]; ko [[slug#섹션]]) — the same Xanadu principle applied to synthesis evidence. A gradual migration, so advisory.

Schema meta-use (cit.grade-meta · cit.cite-type-meta)

The synthesizing piece's narrative consciously reflects the evidence-grade ([fact]·[analysis]·[forecast]) and citation-type (cites:·references:·contradicts:·defines:) schema attached to sources — noting which grades the evidence base rests on, the dominant claimant, strong coupling (cites) vs context (references) — to deepen attribution. Confine attribution strength to factual statements; do not let it spill into verdict conclusions. Deterministic measurement (count_grade_meta·count_cite_type_meta) counts grade/cite-type meta expressions as advisory; the lints judge the cite-type count only under WIKI_LANG=ko.

Sources

Each URL points to the relevant page as of the last verification.

Files (llm-wiki-newsroom)
  • checks.py 24.5 KB
    """scholarly-citation craft skill — deterministic checks.
    
    Verifiable-attribution craft — claim atomization, evidence grading, claimant
    attribution, citation typing, anchoring. It owns the measurement of the source
    page's cit.* criteria (legacy G1–G5·C1–C3·A1–A3). A craft synthesizing external
    sources (Toulmin·scite/Elicit·WP:ASF·scite Smart Citations·Xanadu·Hyper-G·APA)
    into an atomic-claim schema.
    
    content-type-agnostic: wiki-global state (page_index·section-title lookup) is
    injected by the orchestrator (tools/_lint/source.py) — this module uses only pure
    text measurement + injected context. The measurement logic was ported verbatim
    from source.py `_evaluate` (diff-0).
    
    NOTE: the schema matchers below key on the live English wiki schema (## Key Claims·
    ## Connections·[fact]/[analysis]/[forecast]·## Key Quotes·## Representative Evidence)
    and fire normally. Only the Korean-prose matchers are dormant on an English corpus —
    G3 (와/및 conjunctions), G5 (Korean verb endings) — and fire under WIKI_LANG=ko.
    """
    
    from __future__ import annotations
    
    import re
    
    
    # ── cit measurement regexes (ported verbatim from source.py) ──
    # Grade markers [fact]/[analysis]/[forecast] — the live English source.md schema.
    GRADE_MARKER_RE = re.compile(r"^-\s*\[(fact|analysis|forecast)\]", re.MULTILINE)
    CLAIM_LINE_RE = re.compile(r"^-\s+(.+?)\s*$", re.MULTILINE)
    # What G2 measures is whether the claimant is **named** (WP:ASF). The wikilink is
    # one means to that end, so when a speaker falls below the page-creation threshold
    # the plain-text real name is the terminal form — accepting only wikilinks leaves
    # every page citing such a speaker permanently FAILing, and the author reaches for
    # some nearby entity to get through (the misattribution path). Plain text is
    # admitted, and the three cases below are screened out instead. Whether a short
    # plain head is the speaker or the topic is not machine-decidable and stays with
    # the Desk.
    GRADE_HEAD_RE = re.compile(r"^\[(?:fact|analysis|forecast)\]\s+(.*)$")
    # (1) Anonymous / collective subjects are barred in plain text too — the
    # unattributed phrasing WP:ASF forbids. In English the discriminator is case: a
    # named source carries a capitalized final word (`Free Software Foundation`),
    # a weasel does not (`the foundation`). Keying on the word alone without that
    # boundary would drop the very names the guideline offers as correct examples.
    WEASEL_LAST_WORD = frozenset({
        "government", "industry", "market", "media", "authorities", "officials",
        "sources", "experts", "observers", "insiders", "critics", "supporters",
        "analysts", "community", "foundation", "public", "regulators", "researchers",
        "commentators", "advocates", "developers", "companies", "vendors",
    })
    _LEADING_ARTICLE_RE = re.compile(r"^(?:the|a|an)\s+", re.I)
    
    
    def _is_weasel_head(head: str) -> bool:
        words = _LEADING_ARTICLE_RE.sub("", head).split()
        if not words:
            return True
        last = words[-1]
        return last.lower().strip(".,;:") in WEASEL_LAST_WORD and not last[:1].isupper()
    
    
    # (2) Blocks pushing content into the head slot. Longest measured in-corpus: 28.
    CLAIMANT_HEAD_MAX = 40
    CITATION_PREFIX_RE = re.compile(r"^-\s*(cites|references|contradicts|defines)\s*:", re.MULTILINE)
    CONNECT_LINE_RE = re.compile(r"^-\s+.+?$", re.MULTILINE)
    QUOTE_LINE_RE = re.compile(r"^>\s+[\"“”].+", re.MULTILINE)
    # A2 speaker attribution — quote marks pair up, so the outermost quotation closes at an
    # **even-indexed** mark. The separator is the first *spaced* dash between such a mark and
    # the next mark, and the attribution is everything after it. Every clause there was bought
    # by a class that a cheaper rule got wrong: keyed on any dash, a dash inside the quotation
    # passes as an attribution and exempts a quote that names nobody; keyed on the text after
    # the last mark, a quoted title in the attribution (`— [[X]], author of "Y"`) swallows it
    # and reports a linked speaker as unattributed; ignoring parity, a nested term
    # (`"they call it "open source" — …`) outranks the real separator and credits a link from
    # inside the quotation as the speaker; requiring the dash to sit *immediately* after the
    # mark, anything between the two (`"…" (emphasis added) — X`, `"…," she said — X`, or a
    # Korean particle) reads as naming nobody; and allowing an unspaced dash, a hyphenated
    # date (`"…" 2024-10-28`) poses as a separator. The dash *character class* matches the
    # sibling parser in `tools/_lint/cited_speakers.py`, which permits an unspaced dash where
    # this does not.
    #
    # The premise is narrower than "marks pair up": it is that no even-indexed mark before the
    # outer closer carries a spaced dash inside its own window. Two things break it — an
    # unbalanced mark shifts every later parity, and in a balanced nested quotation the inner
    # *opening* mark sits at an even index, so a spaced dash within the nested phrase wins.
    # Scarcity is not the backstop: 20 of the 3,198 quote lines in the sibling corpus that
    # shares this schema carry a wikilink inside the quotation body, so the window bound is.
    #
    # Known cost, measured on that corpus: the in-window search is a loosening. 15 lines read
    # differently than under an adjacent-dash rule and every one moves toward PASS or exempt —
    # 4 correctly, while 7 newly credit a concept hub or an employer org as the speaker (the
    # blind spot step 5 forbids editorially and no check can see) and 2 lose a correct FAIL.
    QUOTE_MARK_RE = re.compile(r"[\"“”]")
    SPEAKER_DASH_RE = re.compile(r"\s+[—–-]+\s+")
    SPEAKER_LINK_RE = re.compile(r"\[\[([^\]|#]+)(?:#[^\]|]+)?(?:\|[^\]]+)?\]\]")
    H2_RE = re.compile(r"^##\s+(.+?)\s*$", re.MULTILINE)
    # (dormant: G3 keys on the Korean "and" conjunctions 와/및 joining two hubs; never
    #  fires on English prose. An English equivalent would detect "and" between two
    #  [[hub]] links. See FLAG.)
    G3_COMPOSITE_RE = re.compile(r"\[\[[^\]]+\]\][^\n]*\s(와|및)\s[^\n]*\[\[[^\]]+\]\]")
    # (dormant: G5 keys on Korean verb endings 했고…했다/한다/이다/된다; never fires on
    #  English prose. An English equivalent would detect verb-phrase juxtaposition like
    #  "…did, and …did". See FLAG.)
    G5_VERB_SPLIT_RE = re.compile(r"했고\s*,\s*[^\n]*?(했다|한다|이다|된다)")
    ANCHOR_LINK_RE = re.compile(r"\[\[([^\]|#]+)#([^\]|]+)(?:\|[^\]]+)?\]\]")
    # Keys on the live English grade markers [fact]/[analysis]/[forecast].
    CLAIMANT_AFTER_GRADE_RE = re.compile(
        r"^-\s*\[(?:fact|analysis|forecast)\]\s+\[\[([^\]|#]+)(?:\|[^\]]+)?\]\]", re.MULTILINE
    )
    CONNECT_PREFIX_TARGET_RE = re.compile(
        r"^-\s*(cites|references|contradicts|defines)\s*:\s*\[\[([^\]|#]+)(?:\|[^\]]+)?\]\]",
        re.MULTILINE,
    )
    
    
    # ── overview·contradiction schema meta-use (legacy G1·G2) ──
    # Ported from the _contradiction_meta_patterns module shared by overview.py·
    # contradiction.py. advisory count of whether the EDITOR body of a cluster overview·
    # theme/aggregate contradiction makes meta-use of the source schema's evidence grade
    # (fact/analysis/forecast)·citation type (cites/references/contradicts/defines).
    # Measures whether a synthesis page reflects the Phase 2 schema (attached to 1283
    # sources) — scite Smart Citations·evidence grading craft.
    
    # G1 — evidence grade reflection.
    # (partly dormant: several patterns key on Korean meta-phrases — 1차/2차/3차 fact
    #  (primary/secondary/tertiary source), 발화 주체/발화자 (claimant), 직접 인용
    #  (direct quote), fact급/analysis급/forecast급, [fact]/[analysis]/[forecast], 증거 등급. The English
    #  patterns (attribution·grade [ABC]·evidence grade) DO fire. An English equivalent
    #  would add "primary/secondary source", "claimant", "direct quote", "[fact]/
    #  [analysis]/[forecast]". See FLAG.)
    GRADE_META_PATTERNS = [
        re.compile(r"1차\s*fact"),
        re.compile(r"2차\s*fact"),
        re.compile(r"3차\s*fact"),
        re.compile(r"발화\s*주체"),
        re.compile(r"발화자"),
        re.compile(r"직접\s*인용"),
        re.compile(r"attribution", re.IGNORECASE),
        re.compile(r"\bgrade\s*[ABC]\b", re.IGNORECASE),
        re.compile(r"fact급"),
        re.compile(r"analysis급"),
        re.compile(r"forecast급"),
        re.compile(r"\[fact\]"),
        re.compile(r"\[analysis\]"),
        re.compile(r"\[forecast\]"),
        re.compile(r"evidence\s*grade", re.IGNORECASE),
        re.compile(r"증거\s*grade", re.IGNORECASE),
        re.compile(r"증거\s*등급"),
    ]
    
    # G2 — citation-type meta-distinction.
    # (partly dormant: several patterns key on Korean meta-phrases — 정의/반박/인용
    #  attribution, 맥락 참조 (context reference), cite/인용 강도 (cite strength), 강한
    #  결합/약한 참조 (strong coupling/weak reference). The English literals (cites:·
    #  references:·contradicts:·defines:) also fire, but the lint callers judge G2 only
    #  under WIKI_LANG=ko and show it as not applicable (—) in English. See FLAG.)
    CITATION_TYPE_META_PATTERNS = [
        re.compile(r"정의\s*attribution", re.IGNORECASE),
        re.compile(r"반박\s*attribution", re.IGNORECASE),
        re.compile(r"인용\s*attribution", re.IGNORECASE),
        re.compile(r"맥락\s*참조"),
        re.compile(r"cite\s*강도", re.IGNORECASE),
        re.compile(r"인용\s*강도"),
        re.compile(r"강한\s*결합"),
        re.compile(r"약한\s*참조"),
        re.compile(r"\bcites\s*:"),
        re.compile(r"\breferences\s*:"),
        re.compile(r"\bcontradicts\s*:"),
        re.compile(r"\bdefines\s*:"),
    ]
    
    
    def count_grade_meta(content: str) -> int:
        """cit.grade-meta algorithm — count of evidence-grade meta expressions (legacy G1).
    
        content is the EDITOR region (AUTO blocks excluded) extracted and injected by the
        orchestrator — this function does not know the content type. The threshold VALUE
        (≥2) is manifest-injected."""
        return sum(len(p.findall(content)) for p in GRADE_META_PATTERNS)
    
    
    def count_cite_type_meta(content: str) -> int:
        """cit.cite-type-meta algorithm — count of citation-type meta-distinction expressions (legacy G2)."""
        return sum(len(p.findall(content)) for p in CITATION_TYPE_META_PATTERNS)
    
    
    def _section_body(content: str, header: str) -> str:
        """Return the body of the H2 section (`## <header>`), or '' if absent. Ported
        verbatim from source.py."""
        pattern = re.compile(rf"^{re.escape(header)}\s*$", re.MULTILINE)
        m = pattern.search(content)
        if not m:
            return ""
        body_start = m.end()
        next_h2 = H2_RE.search(content, body_start)
        body_end = next_h2.start() if next_h2 else len(content)
        return content[body_start:body_end]
    
    
    def evaluate_citation(
        body: str,
        *,
        page_index: dict,
        section_titles_fn,
        c2_ref_ratio_max: float = 0.95,
        c2_min_lines: int = 5,
    ) -> dict:
        """Measure the source page's cit.* criteria (G1–G5·C1–C3·A1–A3).
    
        body is the strip_code'd body. page_index ({slug: (rel, hub_type)})·
        section_titles_fn (rel → section-title set) are wiki-global state injected by the
        orchestrator. The c2 threshold is manifest-injected (content-type-agnostic). The
        returned dict is byte-identical to source.py `_evaluate`'s key tuples.
        """
        # G1, G2, G3 — atomic units in `## Key Claims`.
        # `## Key Claims` is the live English source.md section header.
        claims_body = _section_body(body, "## Key Claims")
        claim_lines = [m.group(1) for m in CLAIM_LINE_RE.finditer(claims_body)]
        claim_total = len(claim_lines)
    
        grade_count = len(GRADE_MARKER_RE.findall(claims_body))
        g1_pass = (claim_total == 0) or (grade_count == claim_total)
        g1_missing_samples = []
        if not g1_pass:
            for line in claim_lines:
                if not re.match(r"\[(fact|analysis|forecast)\]", line):
                    g1_missing_samples.append(line[:60])
                    if len(g1_missing_samples) >= 3:
                        break
    
        claimant_slot_filled = 0
        g2_plain = 0
        g2_missing_samples = []
        for line in claim_lines:
            m = GRADE_HEAD_RE.match(line)
            if not m:
                continue  # a missing grade marker is G1's business
            rest = m.group(1)
            if rest.startswith("[["):
                claimant_slot_filled += 1
                continue
            head = rest.split("—", 1)[0].strip() if "—" in rest else ""
            # (3) A plain head that names a real page is evasion — it could have been linked.
            if (
                head
                and len(head) <= CLAIMANT_HEAD_MAX
                and not _is_weasel_head(head)
                and head not in page_index
            ):
                claimant_slot_filled += 1
                g2_plain += 1
            elif len(g2_missing_samples) < 3:
                g2_missing_samples.append(line[:60])
        g2_pass = (claim_total == 0) or (claimant_slot_filled == claim_total)
    
        g3_violations = len(G3_COMPOSITE_RE.findall(claims_body))
        g3_pass = g3_violations == 0
    
        # G4 — claimant wikilink target validity.
        claimant_targets = CLAIMANT_AFTER_GRADE_RE.findall(claims_body)
        claimant_total = len(claimant_targets)
        claimant_valid = sum(1 for t in claimant_targets if t.strip() in page_index)
        g4_pass = claimant_total == 0 or claimant_valid == claimant_total
        g4_invalid_samples = [t for t in claimant_targets if t.strip() not in page_index][:3]
    
        # G5 — composite verb-clause split.
        g5_violations = 0
        for line in claim_lines:
            if G5_VERB_SPLIT_RE.search(line):
                g5_violations += 1
        g5_pass = g5_violations == 0
    
        # C1, C2 — citation-type prefix in `## Connections` (live English header).
        connect_body = _section_body(body, "## Connections")
        connect_lines = [m.group(0) for m in CONNECT_LINE_RE.finditer(connect_body)]
        connect_total = len(connect_lines)
        prefix_count = len(CITATION_PREFIX_RE.findall(connect_body))
        c1_pass = (connect_total == 0) or (prefix_count == connect_total)
    
        ref_count = len(
            [m for m in CITATION_PREFIX_RE.finditer(connect_body) if m.group(1) == "references"]
        )
        if connect_total < c2_min_lines:
            c2_pass = True
            c2_ratio_pct = -1  # exempt sentinel
        else:
            c2_ratio = ref_count / connect_total if connect_total else 0
            c2_ratio_pct = round(c2_ratio * 100)
            c2_pass = c2_ratio <= c2_ref_ratio_max
    
        # C3 — type-hub matching. `defines:` requires concept hub.
        c3_violations: list[str] = []
        c3_total = 0
        for m in CONNECT_PREFIX_TARGET_RE.finditer(connect_body):
            c3_total += 1
            prefix, target = m.group(1), m.group(2).strip()
            entry = page_index.get(target)
            if entry is None:
                continue  # broken link is G4/structure's job
            _, hub_type = entry
            if prefix == "defines" and hub_type != "concept":
                c3_violations.append(f"{prefix}: [[{target}]] (hub_type={hub_type})")
        c3_pass = len(c3_violations) == 0
    
        # A1 — anchor presence in `[fact]`·`[analysis]` claim lines (advisory only).
        # Keys on the live English grade markers [fact]/[analysis].
        anchor_eligible_lines = [ln for ln in claim_lines if re.match(r"\[(fact|analysis)\]", ln)]
        anchored_lines = [
            ln for ln in anchor_eligible_lines if re.search(r"\[\[[^\]]*#[^\]]+\]\]", ln)
        ]
        a1_eligible = len(anchor_eligible_lines)
        a1_anchored = len(anchored_lines)
        a1_pass = True  # advisory — always PASS
    
        # A2 — blockquote speaker attribution in `## Key Quotes` (live English header).
        # Two things only are decided mechanically: is a speaker named at all, and if one
        # is linked, does the target exist. Whether a plain-text speaker should have been
        # linked is not machine-decidable — the org in a role line is not the speaker
        # (`Matt Garman, AWS CEO`), so name-matching would manufacture misattribution.
        # Plain-text speakers therefore leave the denominator, and link *validity* is
        # counted in place of link presence.
        quotes_body = _section_body(body, "## Key Quotes")
        quote_lines = QUOTE_LINE_RE.findall(quotes_body)
        a2_total = 0
        a2_with = 0
        for line in quote_lines:
            marks = list(QUOTE_MARK_RE.finditer(line))
            attribution = ""
            for i, mark in enumerate(marks, 1):
                if i % 2:
                    continue  # an opening mark — a dash after it sits inside the quotation
                stop = marks[i].start() if i < len(marks) else len(line)
                dash = SPEAKER_DASH_RE.search(line, mark.end(), stop)
                if dash:
                    attribution = line[dash.end():].strip()
                    break
            if not attribution:
                a2_total += 1  # nobody named — there is no attribution to exempt
                continue
            link = SPEAKER_LINK_RE.search(attribution)
            if link:
                a2_total += 1
                if link.group(1).strip() in page_index:
                    a2_with += 1
            elif attribution.split(",", 1)[0].strip() in page_index:
                a2_total += 1  # a plain name that has a page is link evasion, as in G2 (3)
            # else: plain-text speaker with no page — not judged
        a2_pass = a2_with == a2_total
    
        # A3 — anchor wikilink validity (`[[<slug>#<section>]]`).
        a3_total = 0
        a3_valid = 0
        a3_invalid_samples: list[str] = []
        for m in ANCHOR_LINK_RE.finditer(claims_body):
            a3_total += 1
            target_slug, section = m.group(1).strip(), m.group(2).strip()
            entry = page_index.get(target_slug)
            if entry is None:
                a3_invalid_samples.append(f"{target_slug}#{section} (slug not found)")
                continue
            target_rel, _ = entry
            if section in section_titles_fn(target_rel):
                a3_valid += 1
            else:
                a3_invalid_samples.append(f"{target_slug}#{section} (section not found)")
        a3_pass = a3_total == 0 or a3_valid == a3_total
    
        return {
            "g1": (g1_pass, grade_count, claim_total, g1_missing_samples),
            "g2": (g2_pass, claimant_slot_filled, claim_total, g2_missing_samples, g2_plain),
            "g3": (g3_pass, g3_violations),
            "g4": (g4_pass, claimant_valid, claimant_total, g4_invalid_samples),
            "g5": (g5_pass, g5_violations),
            "c1": (c1_pass, prefix_count, connect_total),
            "c2": (c2_pass, c2_ratio_pct, connect_total),
            "c3": (c3_pass, c3_total - len(c3_violations), c3_total, c3_violations[:3]),
            "a1": (a1_pass, a1_anchored, a1_eligible),
            "a2": (a2_pass, a2_with, a2_total),
            "a3": (a3_pass, a3_valid, a3_total, a3_invalid_samples[:3]),
        }
    
    
    # ── contradiction theme cit.* (L2 cite-consistency·L3 grounding·L4 anchor) ──
    # Ported verbatim from contradiction.py. wiki-global (sources_dir)·shared parsing
    # (claim_sources·evidence_slugs) are orchestrator-injected.
    
    L3_ITEM_MIN_CHARS = 20       # minimum item length to register
    L3_SUBSTRING_CHARS = 20      # leading chars used for body substring match
    _QUOTE_BLOCK_RE = re.compile(r'>\s*"([^"]+)"', re.MULTILINE)
    # Source-evidence extractors for L3 grounding: match the live English source-page
    # headers `## Key Quotes` / `## Key Claims` (source.md schema). The Korean headers
    # `## 주요 인용` / `## 주요 주장` fire under WIKI_LANG=ko.
    _SECTION_QUOTES_RE = re.compile(r'##\s*(?:Key Quotes|주요\s*인용)\s*\n(.*?)(?=\n##\s|\Z)', re.DOTALL)
    _SECTION_CLAIMS_RE = re.compile(r'##\s*(?:Key Claims|주요\s*주장)\s*\n(.*?)(?=\n##\s|\Z)', re.DOTALL)
    _BULLET_RE = re.compile(r'^\s*-\s+(.+?)$', re.MULTILINE)
    # Keys on the live English grade markers [fact]/[analysis]/[forecast].
    _CLAIM_PREFIX_RE = re.compile(r'^\s*\[(?:fact|analysis|forecast)\]\s*\[\[[^\]]+\]\]\s*[—–-]\s*')
    _SMART_QUOTE_TRANS = str.maketrans({
        "“": '"', "”": '"',
        "‘": "'", "’": "'",
    })
    _QUOTE_IN_BULLET_RE = re.compile(
        r'["“”][^"“”\n]{3,}?["“”]|[「『][^」』\n]{3,}?[」』]|^>\s*["“]',
        re.MULTILINE,
    )
    # The live English anchor-target section titles (an anchored bullet points to one
    # of these source.md sections).
    _ANCHOR_ALLOWED_SECTIONS = {"Key Quotes", "Key Claims", "Summary"}
    
    
    def _extract_section(content: str, heading: str) -> str:
        """Body between `## <heading>` and the next H2 (or EOF). Verbatim from
        contradiction.py."""
        pattern = re.compile(
            r"^##\s+" + re.escape(heading) + r".*?$(.*?)(?=^##\s|\Z)",
            re.DOTALL | re.MULTILINE,
        )
        m = pattern.search(content)
        return m.group(1) if m else ""
    
    
    def _normalize_for_match(s: str) -> str:
        s = s.translate(_SMART_QUOTE_TRANS)
        s = re.sub(r'[\"\'`*]', '', s)
        s = re.sub(r'\s+', ' ', s)
        return s.strip()
    
    
    def _extract_source_evidence_items(source_md_path) -> list:
        try:
            text = source_md_path.read_text(encoding="utf-8", errors="replace")
        except OSError:
            return []
        items: list[str] = []
        sect_quote = _SECTION_QUOTES_RE.search(text)
        if sect_quote:
            items.extend(_QUOTE_BLOCK_RE.findall(sect_quote.group(1)))
        sect_claim = _SECTION_CLAIMS_RE.search(text)
        if sect_claim:
            claim_body = sect_claim.group(1)
            for bullet in _BULLET_RE.findall(claim_body):
                items.append(bullet)
                prose = _CLAIM_PREFIX_RE.sub("", bullet, count=1)
                if prose != bullet:
                    items.append(prose)
        return [i.strip() for i in items if len(i.strip()) >= L3_ITEM_MIN_CHARS]
    
    
    def _cite_consistency(body: str, frontmatter_sources: list, claim_sources: set):
        """L2 — (missing_count, missing_slugs). body ⊆ frontmatter for claim sources."""
        fm_set = set(frontmatter_sources or [])
        body_cited = [s for s in claim_sources if s in body]
        missing = sorted(s for s in body_cited if s not in fm_set)
        return len(missing), missing
    
    
    def _quote_grounding(body: str, evidence_slugs: set, sources_dir):
        """L3 — (grounded, total_with_evidence, missing_slugs)."""
        grounded = 0
        total_with_evidence = 0
        missing: list[str] = []
        body_norm = _normalize_for_match(body)
        for slug in sorted(evidence_slugs):
            src = sources_dir / f"{slug}.md"
            if not src.exists():
                continue
            items = _extract_source_evidence_items(src)
            if not items:
                continue
            total_with_evidence += 1
            any_grounded = False
            for item in items:
                snippet = _normalize_for_match(item)[:L3_SUBSTRING_CHARS]
                if snippet and snippet in body_norm:
                    any_grounded = True
                    break
            if any_grounded:
                grounded += 1
            else:
                missing.append(slug)
        return grounded, total_with_evidence, missing
    
    
    def _evidence_anchor_check(body: str, source_slugs: set):
        """L4 advisory — (anchored, quoted_total, unanchored_samples)."""
        # `Representative Evidence` is the live English contradiction.md section header.
        evidence_section = _extract_section(body, "Representative Evidence")
        if not evidence_section:
            return 0, 0, []
        parts = re.split(r"(?:^|\n)\s*-\s+", "\n" + evidence_section.strip())
        bullets = [p for p in parts if p.strip()]
        quoted_total = 0
        anchored = 0
        unanchored_samples: list[str] = []
        for bullet in bullets:
            if not _QUOTE_IN_BULLET_RE.search(bullet):
                continue
            quoted_total += 1
            bullet_anchored = False
            for m in ANCHOR_LINK_RE.finditer(bullet):
                stem = m.group(1).strip().split("/")[-1].removesuffix(".md")
                section = m.group(2).strip()
                if section not in _ANCHOR_ALLOWED_SECTIONS:
                    continue
                if source_slugs and stem not in source_slugs:
                    continue
                bullet_anchored = True
                break
            if bullet_anchored:
                anchored += 1
            else:
                preview = bullet.strip().splitlines()[0][:60]
                unanchored_samples.append(preview + ("…" if len(bullet.strip()) > 60 else ""))
        return anchored, quoted_total, unanchored_samples
    
    
    def evaluate_contradiction_citation(
        body: str,
        *,
        fm_sources: list,
        claim_sources: set,
        evidence_slugs: set,
        source_slugs: set,
        sources_dir,
    ) -> dict:
        """Measure the contradiction theme's cit.* (L2·L3·L4). claim_sources (X1)·
        evidence_slugs (S2)·sources_dir are orchestrator-injected. The returned dict is
        byte-identical to the corresponding keys of the original _rubric_metrics (verbatim
        port)."""
        l2_missing_count, l2_missing_slugs = _cite_consistency(body, fm_sources, claim_sources)
        l3_grounded, l3_total_quotes, l3_missing_grounding = _quote_grounding(
            body, evidence_slugs, sources_dir
        )
        l4_anchored, l4_quoted_total, l4_unanchored_samples = _evidence_anchor_check(
            body, source_slugs
        )
        return {
            "L2_missing_count": l2_missing_count,
            "L2_missing_slugs": l2_missing_slugs,
            "L3_grounded": l3_grounded,
            "L3_total_with_quotes": l3_total_quotes,
            "L3_missing_grounding": l3_missing_grounding,
            "L4_anchored": l4_anchored,
            "L4_quoted_total": l4_quoted_total,
            "L4_unanchored_samples": l4_unanchored_samples,
        }
    
  • criteria.json 9.8 KB
    {
      "skill": "scholarly-citation",
      "_note": "Single SoT for verifiable-attribution craft criteria (cit.*). legacy = old content-type-local IDs. The measurement function is given by each `algorithm` field — source G/C/A use `evaluate_citation`; contradiction synthesis integrity (L2·L3·L4) uses `evaluate_contradiction_citation`; schema meta (G1·G2) use `count_grade_meta`·`count_cite_type_meta`. Wiki-global state (page_index·section title·claim_sources·EDITOR region) is orchestrator-injected. All judge=A (deterministic). Any Korean literals are the wiki's actual schema markers and section headers (several matchers key on Korean schema markers — see checks.py FLAGs).",
      "criteria": {
        "cit.grade-marker": {
          "name": "Grade marker present", "dimension": "Claim atomization", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "==", "default_threshold": "claim_total",
          "legacy": {"source": "G1"},
          "source": "Elicit·scite citation classification", "source_url": "https://scite.ai/",
          "note": "Every line of ## Key Claims carries one of [fact]·[analysis]·[forecast]."
        },
        "cit.claimant-link": {
          "name": "Claimant named", "dimension": "Attribution", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "==", "default_threshold": "claim_total",
          "legacy": {"source": "G2"},
          "source": "Wikipedia:ASF", "source_url": "https://en.wikipedia.org/wiki/Wikipedia:Manual_of_Style/Words_to_watch#Attribution",
          "note": "The claimant right after the grade marker must be named: a [[wikilink]] when the speaker has a page, plain text with their role when below the page-creation threshold. Fails when the subject is absent, generic or anonymous, longer than a name (content pushed into the slot), or plain text naming a page that does exist (link evasion). Whether a short plain head is the speaker or the topic is not machine-decidable and stays with editorial review."
        },
        "cit.atomic": {
          "name": "Atomic enforcement (multi-hub juxtaposition)", "dimension": "Claim atomization", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "==", "default_threshold": 0,
          "legacy": {"source": "G3"},
          "source": "Toulmin argument model", "source_url": "https://owl.purdue.edu/owl/general_writing/academic_writing/historical_perspectives_on_argumentation/toulmin_argument.html",
          "note": "No two subjects' claims on one line (Korean 와/및 'and' conjunction + two [[hub]]). NOTE the matcher (G3_COMPOSITE_RE) keys on the Korean conjunction and is dormant on English prose — see checks.py FLAG."
        },
        "cit.claimant-valid": {
          "name": "Claimant validity", "dimension": "Attribution", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "==", "default_threshold": "claimant_total",
          "legacy": {"source": "G4"},
          "source": "Internal standard (broken-link prevention)", "source_url": "",
          "note": "The claimant wikilink target exists as a real page (page_index injected)."
        },
        "cit.composite-split": {
          "name": "Composite split (verb-phrase)", "dimension": "Claim atomization", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "==", "default_threshold": 0,
          "legacy": {"source": "G5"},
          "source": "Toulmin argument model", "source_url": "https://owl.purdue.edu/owl/general_writing/academic_writing/historical_perspectives_on_argumentation/toulmin_argument.html",
          "note": "No '~했고, ~했다' (Korean '…did, and …did') verb-phrase juxtaposition on one line. NOTE the matcher (G5_VERB_SPLIT_RE) keys on Korean verb endings and is dormant on English prose — see checks.py FLAG."
        },
        "cit.cite-type": {
          "name": "Type prefix present", "dimension": "Citation type", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "==", "default_threshold": "connect_total",
          "legacy": {"source": "C1"},
          "source": "scite Smart Citations", "source_url": "https://scite.ai/",
          "note": "Every line of ## Connections starts with a cites:·references:·contradicts:·defines: prefix."
        },
        "cit.cite-distribution": {
          "name": "Type distribution sanity", "dimension": "Citation type", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "<=", "default_threshold": 0.95,
          "legacy": {"source": "C2"},
          "source": "scite Smart Citations", "source_url": "https://scite.ai/",
          "note": "references ratio >95% → advisory (default-safe avoidance). <5 lines exempt."
        },
        "cit.cite-type-hub": {
          "name": "Type-hub consistency", "dimension": "Citation type", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "==", "default_threshold": 0,
          "legacy": {"source": "C3"},
          "source": "Internal standard (decision-tree consistency)", "source_url": "",
          "note": "After defines: only a concept hub (an entity hub = misclassification)."
        },
        "cit.anchor": {
          "name": "Evidence anchor — source claim line (advisory)", "dimension": "Anchor", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "advisory", "default_threshold": null,
          "legacy": {"source": "A1"},
          "source": "Project Xanadu citation anchoring(Nelson 1965)·Hyper-G(Maurer 1996)", "source_url": "https://en.wikipedia.org/wiki/Project_Xanadu",
          "note": "[[slug#section]] anchor on [fact]·[analysis] lines. 0% also PASS (advisory)."
        },
        "cit.speaker-link": {
          "name": "Quote speaker attribution", "dimension": "Anchor", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "==", "default_threshold": "judged_quote_total",
          "legacy": {"source": "A2"},
          "source": "APA in-text citation", "source_url": "https://apastyle.apa.org/style-grammar-guidelines/citations/basic-principles",
          "note": "A ## Key Quotes blockquote names its speaker, wikilinked when that speaker has a page. Denominator = quotes that link a speaker + quotes naming none; the metric scores whether a speaker is named at all and whether the link resolves. A link to a missing page fails. A plain-text speaker is exempt: whether it should have been linked is not machine-decidable, since the org in a role line is not the speaker ('Matt Garman, AWS CEO' is not AWS) and name-matching would manufacture misattribution. Section absence exempt."
        },
        "cit.anchor-valid": {
          "name": "Anchor validity", "dimension": "Anchor", "judge": "A",
          "algorithm": "evaluate_citation", "comparator": "==", "default_threshold": "anchor_total",
          "legacy": {"source": "A3"},
          "source": "Internal standard (broken-link prevention)", "source_url": "",
          "note": "The [[slug#section]] anchor's slug and section both exist (page_index·section title injected)."
        },
        "cit.cite-consistency": {
          "name": "Cite consistency (body ⊆ frontmatter)", "dimension": "Synthesis citation", "judge": "A",
          "algorithm": "evaluate_contradiction_citation", "comparator": "==", "default_threshold": 0,
          "legacy": {"contradiction-theme": "L2"},
          "source": "Internal standard (Select→Read→Cite consistency)", "source_url": "",
          "note": "Every claim source cited in the body is registered in frontmatter sources: (body ⊆ frontmatter, 0 omissions). claim_sources orchestrator-injected."
        },
        "cit.grounding": {
          "name": "Evidence grounding (read discipline)", "dimension": "Synthesis citation", "judge": "A",
          "algorithm": "evaluate_contradiction_citation", "comparator": ">=", "default_threshold": 1,
          "legacy": {"contradiction-theme": "L3"},
          "source": "Internal standard (objective verification of read discipline)", "source_url": "",
          "note": "A representative-evidence source's key-quote/key-claim item appears in the body as a substring (grounded≥1 when total≥1). evidence_slugs·sources_dir orchestrator-injected."
        },
        "cit.anchor-evidence": {
          "name": "Evidence anchor — synthesis representative evidence (advisory)", "dimension": "Synthesis citation", "judge": "A",
          "algorithm": "evaluate_contradiction_citation", "comparator": "advisory", "default_threshold": null,
          "legacy": {"contradiction-theme": "L4"},
          "source": "Project Xanadu citation anchoring(Nelson 1965)", "source_url": "https://en.wikipedia.org/wiki/Project_Xanadu",
          "note": "A representative-evidence citation bullet anchors its source wikilink to a source section such as #Key Quotes. 0% also PASS (advisory·migration). NOTE the allowed-section set is the live English {Key Quotes, Key Claims, Summary}; the Korean section titles fire under WIKI_LANG=ko."
        },
        "cit.grade-meta": {
          "name": "Evidence-grade meta-use", "dimension": "Schema meta-use", "judge": "A",
          "algorithm": "count_grade_meta", "comparator": ">=", "default_threshold": 2,
          "legacy": {"overview-cluster": "G1", "overview-aggregate": "G1", "contradiction-theme": "G1", "contradiction-aggregate": "G1"},
          "source": "scite Smart Citations", "source_url": "https://scite.ai/",
          "note": "The synthesis page's EDITOR body reflects sources' evidence grade (fact/analysis/forecast)·attribution strength as meta (≥2 advisory). content = AUTO-block-excluded EDITOR region, orchestrator-injected. NOTE the meta-phrase patterns (GRADE_META_PATTERNS) are largely Korean — see checks.py FLAG."
        },
        "cit.cite-type-meta": {
          "name": "Citation-type meta-use", "dimension": "Schema meta-use", "judge": "A",
          "algorithm": "count_cite_type_meta", "comparator": ">=", "default_threshold": 1,
          "legacy": {"overview-cluster": "G2", "overview-aggregate": "G2", "contradiction-theme": "G2", "contradiction-aggregate": "G2"},
          "source": "scite Smart Citations", "source_url": "https://scite.ai/",
          "note": "The synthesis page's EDITOR body reflects cites/references/contradicts/defines meaning distinction as meta (≥1 advisory). The overview and contradiction lints judge it only under WIKI_LANG=ko and show it as not applicable (—) in English."
        }
      }
    }
    
  • SKILL.md 6.7 KB
    ---
    name: scholarly-citation
    description: Verifiable-attribution and citation-discipline craft — atomic claim decomposition, evidence grading (primary/analysis/forecast tiers), claimant attribution, citation typing (cites/references/contradicts/defines), source anchoring (Xanadu). Use when writing or reviewing a reference or synthesis page that attributes and cites source-based claims so they stay traceable, or when claim attribution, citation typing, or evidence grading is needed.
    ---
    
    # scholarly-citation
    
    Verifiable-attribution craft drawn from scholarly/scientific citation traditions (citation classification, typing, anchoring). It ensures every claim in this wiki is atomized, evidence-graded, attributed to a speaker, typed by citation meaning, and anchored to a source location — and therefore traceable. Unlike the three borrowed traditions (narrative, exec summary, encyclopedic neutrality), this craft synthesizes external sources into an atomic-claim schema and is the project's signature. `criteria.json` is the SoT for each criterion's definition, comparator, and source; wiki-global state (page existence·hub type·section titles) is injected by the orchestrator (`tools/_lint/source.py`). The schema literals (`## Key Claims`, `[fact]`, `cites:` etc.) are the wiki's actual schema markers and section headers.
    
    ## Claim atomization & grading (cit.grade-marker · cit.atomic · cit.composite-split)
    
    One claim, one line, one utterance. Each line of `## Key Claims` (Key Claims) marks an evidence grade — `[fact]` (primary source), `[analysis]` (secondary analysis), `[forecast]` (forecast). Putting two subjects' claims on one line (joined by "and"/"&" or a comma series spanning two hubs; ko: the conjunctions `와`/`및`), or chaining verb phrases on one line ("…did, and …did"; ko: `~했고, ~했다`), breaks atomicity.
    
    ## Claimant attribution (cit.claimant-link · cit.claimant-valid · cit.speaker-link)
    
    Every claim is bound to who said it. The slot right after the grade marker names the speaker — a `[[wikilink]]` when they have a page, their plain-text real name and role when they fall below the page-creation threshold. What WP:ASF requires is the named source; the wikilink is one means to it, so demanding a link from a speaker who cannot have a page only pushes the author to link some other entity. Anonymous or collective subjects ("the government"·"industry sources"; ko `정부는`·`업계에서는`) stay forbidden in either form ("avoid stating opinions as facts"; attribute them to a named source), as does naming in plain text a page that does exist. A blockquote names its speaker (APA author-date principle — identify the cited author), wikilinked when that speaker has a page. A speaker with no page stays plain text with their role spelled out. Any target that is linked must exist as a real page.
    
    ## Citation typing (cit.cite-type · cit.cite-distribution · cit.cite-type-hub)
    
    Each line of `## Connections` (Connections) types the citation's meaning — `cites:` (direct grounds)·`references:` (context)·`contradicts:` (dispute)·`defines:` (definition). This extends scite Smart Citations' Supporting / Mentioning / Contrasting classes (with `defines:` added). If every line collapses to the default-safe `references:`, that signals avoidance of meaning classification. `defines:` points only to a concept hub.
    
    ## Source anchoring (cit.anchor · cit.anchor-valid)
    
    An external-evidence citation anchors to the source's section — `[[slug#section]]` (ko: `[[slug#섹션]]`). This applies Xanadu citation anchoring (visible links to the precise source location, Ted Nelson) at the source level; a gradual migration, so 0% still passes (advisory), but the anchor target's slug and section must exist.
    
    ## Synthesis citation integrity (cit.cite-consistency · cit.grounding · cit.anchor-evidence)
    
    A piece synthesizing many sources follows Select→Read→Cite discipline — the declared source set (frontmatter) is the top set, and both the sources read and the sources cited in the body are subsets of it.
    
    - **cit.cite-consistency** — every source cited in the body is registered in the declared source list (body ⊆ declaration). Citing an undeclared source is an editorial slip (omission).
    - **cit.grounding** — a source offered as representative evidence must actually be read and used. When at least one of the source's direct-quote / key-claim items appears in the body as a substring, that is objective evidence of having read it (writing from a one-line summary by memory/guess risks divergence from the source). The deterministic lint blocks only the "never read at all" extreme; full read-through rests on editor self-discipline.
    - **cit.anchor-evidence** — when a direct quote is placed in representative evidence, mark the citation location with a source section anchor (`[[slug#section]]`; ko `[[slug#섹션]]`) — the same Xanadu principle applied to synthesis evidence. A gradual migration, so advisory.
    
    ## Schema meta-use (cit.grade-meta · cit.cite-type-meta)
    
    The synthesizing piece's narrative consciously reflects the evidence-grade (`[fact]`·`[analysis]`·`[forecast]`) and citation-type (`cites:`·`references:`·`contradicts:`·`defines:`) schema attached to sources — noting which grades the evidence base rests on, the dominant claimant, strong coupling (cites) vs context (references) — to deepen attribution. Confine attribution strength to factual statements; do not let it spill into verdict conclusions. Deterministic measurement (`count_grade_meta`·`count_cite_type_meta`) counts grade/cite-type meta expressions as advisory; the lints judge the cite-type count only under WIKI_LANG=ko.
    
    ## Sources
    
    Each URL points to the relevant page as of the last verification.
    
    - [Toulmin Argument — Purdue OWL](https://owl.purdue.edu/owl/general_writing/academic_writing/historical_perspectives_on_argumentation/toulmin_argument.html) — Claim·Data·Warrant·Qualifier decomposition (claim atomization)
    - [scite citation classification](https://scite.ai/) · [Elicit](https://elicit.com/) — claim atomization + citation-context classification (evidence grading·citation typing)
    - [Wikipedia:Manual of Style/Words to watch — WP:ASF](https://en.wikipedia.org/wiki/Wikipedia:Manual_of_Style/Words_to_watch#Attribution) — Attribute Statements to Facts (claimant attribution)
    - [scite Smart Citations](https://scite.ai/) — Supporting/Mentioning/Contrasting classification (→ extended with defines)
    - [Project Xanadu citation anchoring (Nelson 1965)](https://en.wikipedia.org/wiki/Project_Xanadu) · [Hyper-G typed edges (Maurer 1996)](https://www.iicm.tugraz.at/) — citation-location anchoring
    - [APA in-text citation guidelines](https://apastyle.apa.org/style-grammar-guidelines/citations/basic-principles) — name the cited speaker
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related