Claude Cursor Skill

verify-refs

Use when checking BibTeX reference records for missing fields, malformed identifiers, duplicate keys, or metadata mismatches before submission.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download yha9806-academic-writing-toolkit-.claude_skills_verify-refs-184e482.zip · 8 KB
Part of yha9806/academic-writing-toolkit — 21 skills

Install

skills CLI npx skills add https://github.com/yha9806/academic-writing-toolkit/tree/main/.claude/skills/verify-refs
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install yha9806-academic-writing-toolkit@llmmart
Git git clone https://github.com/yha9806/academic-writing-toolkit.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole yha9806/academic-writing-toolkit collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

/verify-refs — Reference Authenticity Check

Running Python helpers

Choose the interpreter before running the examples. For a globally installed copy, use the private runtime recorded by its installer. In a checkout or linked workspace, use AWT_PYTHON when set, otherwise the toolkit's .venv (follow the skill directory link back to the toolkit): Scripts/python.exe on Windows, bin/python on macOS/Linux. Without that environment, check that python (Windows) or python3 (macOS/Linux) actually runs and has the helper's dependencies. Replace the example's python3 with that executable. In PowerShell, prefix a quoted executable with &; keep commands on one line and quote file paths.

Purpose

Check reference records for missing required fields, duplicate keys, malformed DOI values, malformed arXiv identifiers, and invalid URLs. The default mode is offline and deterministic.

Trigger Words

This skill activates on: verify refs, verify references, reference check, /verify-refs.

Workflow

  1. Identify the target Markdown or .bib file. If the user does not specify one, ask.
  2. Run: python3 .claude/skills/verify-refs/scripts/verify-refs.py --bib "{path}" --json
  3. Report issues by entry key and severity.
  4. For explicit metadata verification, run: python3 .claude/skills/verify-refs/scripts/verify-refs.py --bib "{path}" --json --online
  5. Use CrossRef for DOI metadata, Semantic Scholar as a secondary metadata source, and arXiv for preprint identifiers. Online checks must be explicit because they depend on network availability.
  6. In tests or offline review, use --metadata-dir "{dir}" to read CrossRef JSON, Semantic Scholar JSON, and arXiv Atom fixtures instead of live network calls.
  7. For a LaTeX manuscript, reconcile its citations with the .bib file both ways: python3 .claude/skills/verify-refs/scripts/reconcile-cites.py --bib "{bib}" --root "{repo}" --json "{main.tex}" ... It follows \input, \include and \subfile, reports keys cited but not defined and entries defined but cited nowhere, and treats \nocite{*} as citing everything. Exit 2 when nothing could be read.

Constraints

  1. Do not import project-specific reference rules from other repositories.
  2. Do not auto-fix reference records without user approval.
  3. Keep output plain Markdown with no emoji.
  4. Treat CrossRef, Semantic Scholar, and arXiv as verification sources, not as citation style authorities.
Files (academic-writing-toolkit)
  • scripts
    • reconcile-cites.py 5.1 KB
      #!/usr/bin/env python3
      """Reconcile a LaTeX manuscript's citations with its BibTeX file, both ways.
      
          python3 reconcile-cites.py --bib references.bib [--root DIR] [--json] FILE...
      
      Reads each FILE and every file it pulls in with \\input, \\include or \\subfile (relative to --root, ".tex" added
      when missing), and collects the keys of every citation command (\\cite, \\citet, \\citep, \\citeauthor, \\nocite,
      \\parencite, \\textcite, \\autocite ...). \\nocite{*} cites every entry. Reports:
      
      - cited-not-in-bib: a key the text cites that the bibliography does not define (with where it is first cited);
      - bib-not-cited:    an entry the bibliography defines that nothing reads.
      
      Duplicate keys and malformed entries are verify-refs.py's job and are not repeated here; this uses its parser.
      A \\input it cannot find is listed under unresolved_inputs, not counted as an issue (a generated file is often
      missing from a clean tree), and the report says so.
      
      Exit: 0 both lists empty; 1 at least one issue; 2 nothing to reconcile (no FILE could be read, or the
      bibliography is unreadable or defines no entry).
      """
      import argparse
      import importlib.util
      import json
      import re
      import sys
      from pathlib import Path
      
      HERE = Path(__file__).resolve().parent
      CITE = re.compile(r"\\[A-Za-z]*cite[A-Za-z]*\*?(?:\s*\[[^\]]*\]){0,2}\s*\{([^}]*)\}")
      INPUT = re.compile(r"\\(?:input|include|subfile)\s*\{([^}]+)\}")
      COMMENT = re.compile(r"(?<!\\)%.*")
      
      
      def _parser():
          spec = importlib.util.spec_from_file_location("verify_refs", HERE / "verify-refs.py")
          mod = importlib.util.module_from_spec(spec)
          saved = sys.argv
          sys.argv = ["verify-refs.py"]
          try:
              spec.loader.exec_module(mod)
          finally:
              sys.argv = saved
          return mod.parse_bibtex
      
      
      def read_tree(files, root):
          """[(path, text without comments)] for the files given and everything they input, each file once."""
          out, seen, unresolved = [], set(), []
          stack = [Path(f) for f in reversed(files)]
          while stack:
              p = stack.pop()
              key = str(p.resolve()) if p.exists() else str(p)
              if key in seen:
                  continue
              seen.add(key)
              try:
                  text = COMMENT.sub("", p.read_text(encoding="utf-8", errors="replace"))
              except OSError:
                  unresolved.append(str(p))
                  continue
              out.append((p, text))
              for m in reversed(list(INPUT.finditer(text))):
                  name = m.group(1).strip()
                  q = Path(root) / name
                  if not q.suffix:
                      q = q.with_suffix(".tex")
                  if q.exists():
                      stack.append(q)
                  else:
                      unresolved.append(name)
          return out, unresolved
      
      
      def main():
          ap = argparse.ArgumentParser(description="Reconcile LaTeX citations with a BibTeX file.")
          ap.add_argument("--bib", required=True)
          ap.add_argument("--root", default=".")
          ap.add_argument("--json", action="store_true")
          ap.add_argument("files", nargs="+")
          a = ap.parse_args()
          try:
              entries = _parser()(Path(a.bib).read_text(encoding="utf-8", errors="replace"))
          except OSError as e:
              sys.stderr.write(f"reconcile-cites: cannot read the bibliography {a.bib}: {e}\n")
              return 2
          defined = [e["key"] for e in entries]
          if not defined:
              sys.stderr.write(f"reconcile-cites: {a.bib} defines no entry\n")
              return 2
          files, unresolved = read_tree(a.files, a.root)
          if not files:
              sys.stderr.write("reconcile-cites: none of the files given could be read\n")
              return 2
          first = {}
          everything = False
          for path, text in files:
              for m in CITE.finditer(text):
                  line = text.count("\n", 0, m.start()) + 1
                  for k in (x.strip() for x in m.group(1).split(",")):
                      if k == "*":
                          everything = True
                      elif k:
                          first.setdefault(k, f"{path}:{line}")
          bib = set(defined)
          issues = [{"kind": "cited-not-in-bib", "severity": "high", "key": k, "location": loc,
                     "message": "Cited in the text, not defined in the bibliography."}
                    for k, loc in sorted(first.items()) if k not in bib]
          if not everything:
              issues += [{"kind": "bib-not-cited", "severity": "medium", "key": k, "location": a.bib,
                          "message": "Defined in the bibliography, cited nowhere in the files read."}
                         for k in sorted(bib - set(first))]
          payload = {"schema_version": 1, "bib_entries": len(bib), "cited_keys": len(first), "nocite_all": everything,
                     "files_read": [str(p) for p, _ in files], "unresolved_inputs": sorted(set(unresolved)),
                     "issues": issues, "issue_count": len(issues)}
          if a.json:
              print(json.dumps(payload, indent=2, ensure_ascii=False))
          else:
              for i in issues:
                  print(f"{i['location']}: {i['kind']}: {i['key']}")
              print(f"read {len(files)} file(s); {len(first)} cited key(s), {len(bib)} bibliography entries; "
                    f"{len(issues)} issue(s)" + (f"; unresolved inputs: {', '.join(sorted(set(unresolved)))}" if unresolved else ""))
          return 1 if issues else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • verify-refs.py 17.3 KB
      #!/usr/bin/env python3
      """Offline-first academic reference verifier.
      
      The default mode is deterministic and CI-safe: parse BibTeX, validate required
      fields, duplicate keys, DOI shape, URL shape, and arXiv identifier shape.
      Online metadata checks are available behind an explicit flag without changing
      the offline contract.
      """
      
      import argparse
      import json
      import re
      import sys
      import urllib.parse
      import urllib.request
      import xml.etree.ElementTree as ET
      from pathlib import Path
      from typing import Dict, List, Optional, Tuple
      
      REQUIRED = {
          "article": ["title", "author", "year", "journal"],
          "inproceedings": ["title", "author", "year", "booktitle"],
          "book": ["title", "author", "year", "publisher"],
          "misc": ["title", "year"],
          "phdthesis": ["title", "author", "year", "school"],
          "techreport": ["title", "author", "year", "institution"],
      }
      
      ARXIV_NEW = re.compile(r"^\d{4}\.\d{4,5}(v\d+)?$")
      ARXIV_OLD = re.compile(r"^[a-z-]+/\d{7}(v\d+)?$")
      DOI_RE = re.compile(r"^10\.\d{4,9}/\S+$", re.IGNORECASE)
      
      CROSSREF_BASE = "https://api.crossref.org/works/"
      SEMANTIC_SCHOLAR_BASE = "https://api.semanticscholar.org/graph/v1/paper/"
      ARXIV_BASE = "https://export.arxiv.org/api/query"
      
      
      def extract_bibtex(text: str) -> str:
          blocks = re.findall(r"```(?:bibtex|bib)\s*(.*?)```", text, flags=re.DOTALL | re.IGNORECASE)
          return "\n\n".join(blocks) if blocks else text
      
      
      def _balanced_group(text: str, open_index: int) -> Tuple[Optional[str], int]:
          """Return the content of the brace group starting at open_index ('{') and
          the index just past its closing brace, honouring nested braces."""
          depth = 0
          for i in range(open_index, len(text)):
              ch = text[i]
              if ch == "{":
                  depth += 1
              elif ch == "}":
                  depth -= 1
                  if depth == 0:
                      return text[open_index + 1 : i], i + 1
          return None, open_index
      
      
      _FIELD_NAME = re.compile(r"(?P<name>[A-Za-z_][A-Za-z0-9_-]*)\s*=\s*")
      
      
      def parse_bibtex(text: str) -> List[dict]:
          """Deterministic BibTeX reader: balanced-brace entry and value scanning.
      
          Replaces the earlier lazy-regex parser, which truncated brace-protected
          values at the first inner '}' and reported unbraced numeric fields
          (e.g. `year = 2024`) as missing. Inner braces are preserved verbatim;
          downstream text comparison normalises them away.
          """
          entries: List[dict] = []
          for m in re.finditer(r"@(?P<type>\w+)\s*\{", text):
              entry_type = m.group("type").lower()
              if entry_type in ("comment", "preamble", "string"):
                  continue
              body, _end = _balanced_group(text, m.end() - 1)
              if body is None or "," not in body:
                  continue
              key, rest = body.split(",", 1)
              fields: Dict[str, str] = {}
              i = 0
              while True:
                  fm = _FIELD_NAME.search(rest, i)
                  if not fm:
                      break
                  name = fm.group("name").lower()
                  j = fm.end()
                  if j < len(rest) and rest[j] == "{":
                      value, i = _balanced_group(rest, j)
                      if value is None:
                          break
                  elif j < len(rest) and rest[j] == '"':
                      closing = rest.find('"', j + 1)
                      if closing == -1:
                          break
                      value, i = rest[j + 1 : closing], closing + 1
                  else:
                      bare = re.match(r"[^,\n]*", rest[j:])
                      value = bare.group(0).strip() if bare else ""
                      i = j + (bare.end() if bare else 0)
                  fields[name] = " ".join(value.split())
              entries.append({"type": entry_type, "key": key.strip(), "fields": fields})
          return entries
      
      
      def audit_entries(entries: List[dict]) -> List[dict]:
          issues: List[dict] = []
          seen: Dict[str, int] = {}
          for index, entry in enumerate(entries, start=1):
              key = entry["key"]
              if key in seen:
                  issues.append({"kind": "duplicate-key", "severity": "high", "entry": key, "message": "BibTeX key is repeated."})
              seen[key] = index
              required = REQUIRED.get(entry["type"], ["title", "year"])
              for field in required:
                  if not entry["fields"].get(field):
                      issues.append({"kind": "missing-required-field", "severity": "high", "entry": key, "message": "Missing required field: {}".format(field)})
              doi = entry["fields"].get("doi")
              if doi and not DOI_RE.match(doi):
                  issues.append({"kind": "doi-invalid", "severity": "medium", "entry": key, "message": "DOI format is invalid."})
              arxiv = entry["fields"].get("arxiv_id") or entry["fields"].get("eprint")
              if arxiv and not (ARXIV_NEW.match(arxiv) or ARXIV_OLD.match(arxiv)):
                  issues.append({"kind": "arxiv-invalid", "severity": "medium", "entry": key, "message": "arXiv identifier format is invalid."})
              url = entry["fields"].get("url")
              if url and not re.match(r"^https?://", url):
                  issues.append({"kind": "url-invalid", "severity": "medium", "entry": key, "message": "URL must start with http:// or https://."})
          return issues
      
      
      def normalize_text(value: str) -> str:
          return " ".join(re.sub(r"[^a-z0-9]+", " ", value.lower()).split())
      
      
      def jaccard_similarity(a: str, b: str) -> float:
          left = set(normalize_text(a).split())
          right = set(normalize_text(b).split())
          if not left and not right:
              return 1.0
          if not left or not right:
              return 0.0
          return float(len(left & right)) / float(len(left | right))
      
      
      def bib_year(entry: dict) -> Optional[int]:
          raw = entry["fields"].get("year", "")
          m = re.search(r"\d{4}", raw)
          return int(m.group(0)) if m else None
      
      
      def bib_author_count(entry: dict) -> int:
          authors = entry["fields"].get("author", "")
          if not authors:
              return 0
          return len([p for p in re.split(r"\s+and\s+", authors) if p.strip()])
      
      
      def safe_id(value: str) -> str:
          return value.replace("/", "_").replace(":", "_")
      
      
      def read_json_fixture(metadata_dir: Optional[Path], source: str, identifier: str) -> Optional[dict]:
          if metadata_dir is None:
              return None
          path = metadata_dir / source / (safe_id(identifier) + ".json")
          if not path.is_file():
              return None
          return json.loads(path.read_text(encoding="utf-8"))
      
      
      def read_text_fixture(metadata_dir: Optional[Path], source: str, identifier: str, suffix: str) -> Optional[str]:
          if metadata_dir is None:
              return None
          path = metadata_dir / source / (safe_id(identifier) + suffix)
          if not path.is_file():
              return None
          return path.read_text(encoding="utf-8")
      
      
      def fetch_json(url: str) -> Optional[dict]:
          request = urllib.request.Request(url, headers={"User-Agent": "academic-writing-toolkit/1.0"})
          try:
              with urllib.request.urlopen(request, timeout=15) as response:
                  if response.status >= 400:
                      return None
                  return json.loads(response.read().decode("utf-8"))
          except Exception:
              return None
      
      
      def fetch_text(url: str) -> Optional[str]:
          request = urllib.request.Request(url, headers={"User-Agent": "academic-writing-toolkit/1.0"})
          try:
              with urllib.request.urlopen(request, timeout=15) as response:
                  if response.status >= 400:
                      return None
                  return response.read().decode("utf-8")
          except Exception:
              return None
      
      
      def crossref_metadata(entry: dict, metadata_dir: Optional[Path]) -> Optional[dict]:
          doi = entry["fields"].get("doi")
          if not doi:
              return None
          fixture = read_json_fixture(metadata_dir, "crossref", doi)
          if fixture is not None:
              return fixture
          url = CROSSREF_BASE + urllib.parse.quote(doi, safe="")
          return fetch_json(url)
      
      
      def semantic_scholar_metadata(entry: dict, metadata_dir: Optional[Path]) -> Optional[dict]:
          identifier = entry["fields"].get("doi") or entry["fields"].get("arxiv_id") or entry["fields"].get("eprint")
          if not identifier:
              return None
          fixture = read_json_fixture(metadata_dir, "semantic-scholar", identifier)
          if fixture is not None:
              return fixture
          if entry["fields"].get("doi"):
              paper_id = "DOI:" + identifier
          else:
              paper_id = "ARXIV:" + identifier
          url = SEMANTIC_SCHOLAR_BASE + urllib.parse.quote(paper_id, safe=":") + "?fields=title,authors,year,venue"
          return fetch_json(url)
      
      
      def arxiv_metadata(entry: dict, metadata_dir: Optional[Path]) -> Optional[str]:
          arxiv_id = entry["fields"].get("arxiv_id") or entry["fields"].get("eprint")
          if not arxiv_id:
              return None
          fixture = read_text_fixture(metadata_dir, "arxiv", arxiv_id, ".xml")
          if fixture is not None:
              return fixture
          url = ARXIV_BASE + "?id_list=" + urllib.parse.quote(arxiv_id)
          return fetch_text(url)
      
      
      def year_from_crossref(message: dict) -> Optional[int]:
          for key in ("published-print", "published-online", "published", "issued"):
              parts = message.get(key, {}).get("date-parts")
              if parts and parts[0]:
                  return int(parts[0][0])
          return None
      
      
      def compare_metadata(entry: dict, source: str, title: str, year: Optional[int], author_count: int) -> Tuple[List[dict], Optional[dict]]:
          issues: List[dict] = []
          key = entry["key"]
          bib_title = entry["fields"].get("title", "")
          sim = jaccard_similarity(bib_title, title)
          if sim < 0.60:
              issues.append({
                  "kind": "metadata-title-low-similarity",
                  "severity": "high",
                  "entry": key,
                  "source": source,
                  "message": "Title similarity to {} metadata is below 60%.".format(source),
              })
          elif sim < 0.90:
              issues.append({
                  "kind": "metadata-title-minor-diff",
                  "severity": "medium",
                  "entry": key,
                  "source": source,
                  "message": "Title similarity to {} metadata is below 90%.".format(source),
              })
          byear = bib_year(entry)
          if byear is not None and year is not None:
              diff = abs(byear - year)
              if diff > 2:
                  issues.append({
                      "kind": "metadata-year-mismatch",
                      "severity": "high",
                      "entry": key,
                      "source": source,
                      "message": "Year differs from {} metadata by more than two years.".format(source),
                  })
              elif diff > 0:
                  issues.append({
                      "kind": "metadata-year-mismatch",
                      "severity": "medium",
                      "entry": key,
                      "source": source,
                      "message": "Year differs from {} metadata.".format(source),
                  })
          bcount = bib_author_count(entry)
          if bcount and author_count:
              ratio = abs(bcount - author_count) / float(max(bcount, author_count))
              if ratio > 0.50:
                  issues.append({
                      "kind": "metadata-author-count-mismatch",
                      "severity": "high",
                      "entry": key,
                      "source": source,
                      "message": "Author count differs from {} metadata by more than 50%.".format(source),
                  })
              elif ratio > 0.20:
                  issues.append({
                      "kind": "metadata-author-count-mismatch",
                      "severity": "medium",
                      "entry": key,
                      "source": source,
                      "message": "Author count differs from {} metadata.".format(source),
                  })
          if issues:
              return issues, None
          return issues, {"entry": key, "source": source, "title_similarity": sim}
      
      
      def parse_crossref_result(payload: dict) -> Tuple[str, Optional[int], int]:
          message = payload.get("message", payload)
          title = (message.get("title") or [""])[0]
          year = year_from_crossref(message)
          author_count = len(message.get("author") or [])
          return title, year, author_count
      
      
      def parse_semantic_scholar_result(payload: dict) -> Tuple[str, Optional[int], int]:
          title = payload.get("title", "")
          year = payload.get("year")
          author_count = len(payload.get("authors") or [])
          return title, int(year) if year else None, author_count
      
      
      def parse_arxiv_result(xml_text: str) -> Optional[Tuple[str, Optional[int], int]]:
          root = ET.fromstring(xml_text)
          ns = {"atom": "http://www.w3.org/2005/Atom"}
          entry = root.find("atom:entry", ns)
          if entry is None:
              return None
          title_el = entry.find("atom:title", ns)
          published_el = entry.find("atom:published", ns)
          authors = entry.findall("atom:author", ns)
          title = " ".join((title_el.text or "").split()) if title_el is not None else ""
          year = None
          if published_el is not None and published_el.text:
              m = re.search(r"\d{4}", published_el.text)
              year = int(m.group(0)) if m else None
          return title, year, len(authors)
      
      
      def verify_online(entries: List[dict], metadata_dir: Optional[Path]) -> Tuple[List[dict], List[dict], List[dict]]:
          issues: List[dict] = []
          verified: List[dict] = []
          checks: List[dict] = []
          for entry in entries:
              key = entry["key"]
              checked = False
              crossref = crossref_metadata(entry, metadata_dir)
              if crossref is not None:
                  checked = True
                  title, year, author_count = parse_crossref_result(crossref)
                  new_issues, ok = compare_metadata(entry, "crossref", title, year, author_count)
                  issues.extend(new_issues)
                  checks.append({"entry": key, "source": "crossref", "status": "checked"})
                  if ok:
                      verified.append(ok)
              elif entry["fields"].get("doi"):
                  issues.append({"kind": "doi-not-found", "severity": "high", "entry": key, "source": "crossref", "message": "DOI was not found in Crossref metadata."})
      
              semantic = semantic_scholar_metadata(entry, metadata_dir)
              if semantic is not None:
                  checked = True
                  title, year, author_count = parse_semantic_scholar_result(semantic)
                  new_issues, ok = compare_metadata(entry, "semantic-scholar", title, year, author_count)
                  issues.extend(new_issues)
                  checks.append({"entry": key, "source": "semantic-scholar", "status": "checked"})
                  if ok:
                      verified.append(ok)
      
              arxiv_xml = arxiv_metadata(entry, metadata_dir)
              if arxiv_xml is not None:
                  checked = True
                  parsed = parse_arxiv_result(arxiv_xml)
                  checks.append({"entry": key, "source": "arxiv", "status": "checked"})
                  if parsed is None:
                      issues.append({"kind": "arxiv-not-found", "severity": "high", "entry": key, "source": "arxiv", "message": "arXiv metadata did not contain an entry."})
                  else:
                      title, year, author_count = parsed
                      new_issues, ok = compare_metadata(entry, "arxiv", title, year, author_count)
                      issues.extend(new_issues)
                      if ok:
                          verified.append(ok)
              elif entry["fields"].get("arxiv_id") or entry["fields"].get("eprint"):
                  issues.append({"kind": "arxiv-not-found", "severity": "high", "entry": key, "source": "arxiv", "message": "arXiv identifier was not found."})
      
              if not checked:
                  checks.append({"entry": key, "source": "none", "status": "skipped"})
          return issues, verified, checks
      
      
      def main() -> int:
          parser = argparse.ArgumentParser(description="Verify academic references from BibTeX.")
          parser.add_argument("path", nargs="?")
          parser.add_argument("--bib", dest="bib_path")
          parser.add_argument("--json", action="store_true", dest="emit_json")
          parser.add_argument("--online", action="store_true", help="Verify records against external metadata sources.")
          parser.add_argument("--metadata-dir", help="Read metadata fixtures from DIR instead of live APIs when present.")
          parser.add_argument("--allow-empty", action="store_true",
                              help="a bibliography with no entries is expected here, so exit 0 rather than 2")
          args = parser.parse_args()
          target = Path(args.bib_path or args.path or "")
          if not target.is_file():
              sys.stderr.write("error: input file not found\n")
              return 2
          text = extract_bibtex(target.read_text(encoding="utf-8"))
          entries = parse_bibtex(text)
          issues = audit_entries(entries)
          verified: List[dict] = []
          metadata_checks: List[dict] = []
          if args.online:
              metadata_dir = Path(args.metadata_dir) if args.metadata_dir else None
              online_issues, verified, metadata_checks = verify_online(entries, metadata_dir)
              issues.extend(online_issues)
          # A bibliography with no entries verifies no reference, so "0 issues" here
          # states only that nothing was examined. Reported as a pass it reads as
          # "the references were checked and are sound", which is the opposite.
          nothing_checked = not entries
          payload = {
              "schema_version": 1,
              "entries": len(entries),
              "nothing_checked": nothing_checked,
              "issues": issues,
              "issue_count": len(issues),
              "verified": verified,
              "metadata_checks": metadata_checks,
              "online_sources": sorted(set(c["source"] for c in metadata_checks if c["source"] != "none")),
          }
          if args.emit_json:
              print(json.dumps(payload, indent=2))
          else:
              print("entries: {}".format(len(entries)))
              for issue in issues:
                  print("{entry}: {kind}: {message}".format(**issue))
              if nothing_checked:
                  print("NOTHING CHECKED: the bibliography holds no entries. This is not a pass.")
          if issues:
              return 1
          return 2 if nothing_checked and not args.allow_empty else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • SKILL.md 2.6 KB
    ---
    name: verify-refs
    description: Use when checking BibTeX reference records for missing fields, malformed identifiers, duplicate keys, or metadata mismatches before submission.
    allowed-tools: Read, Glob, Bash, WebFetch
    ---
    
    # /verify-refs — Reference Authenticity Check
    
    ## Running Python helpers
    
    Choose the interpreter before running the examples. For a globally installed
    copy, use the private runtime recorded by its installer. In a checkout or linked
    workspace, use `AWT_PYTHON` when set, otherwise the toolkit's `.venv` (follow the
    skill directory link back to the toolkit): `Scripts/python.exe` on Windows,
    `bin/python` on macOS/Linux. Without that environment, check that `python`
    (Windows) or `python3` (macOS/Linux) actually runs and has the helper's dependencies.
    Replace the example's `python3` with that executable. In PowerShell, prefix a
    quoted executable with `&`; keep commands on one line and quote file paths.
    
    ## Purpose
    
    Check reference records for missing required fields, duplicate keys, malformed DOI values, malformed arXiv identifiers, and invalid URLs. The default mode is offline and deterministic.
    
    ## Trigger Words
    
    This skill activates on: `verify refs`, `verify references`, `reference check`, `/verify-refs`.
    
    ## Workflow
    
    1. Identify the target Markdown or `.bib` file. If the user does not specify one, ask.
    2. Run:
       `python3 .claude/skills/verify-refs/scripts/verify-refs.py --bib "{path}" --json`
    3. Report issues by entry key and severity.
    4. For explicit metadata verification, run:
       `python3 .claude/skills/verify-refs/scripts/verify-refs.py --bib "{path}" --json --online`
    5. Use CrossRef for DOI metadata, Semantic Scholar as a secondary metadata source, and arXiv for preprint identifiers. Online checks must be explicit because they depend on network availability.
    6. In tests or offline review, use `--metadata-dir "{dir}"` to read CrossRef JSON, Semantic Scholar JSON, and arXiv Atom fixtures instead of live network calls.
    7. For a LaTeX manuscript, reconcile its citations with the `.bib` file both ways:
       `python3 .claude/skills/verify-refs/scripts/reconcile-cites.py --bib "{bib}" --root "{repo}" --json "{main.tex}" ...`
       It follows `\input`, `\include` and `\subfile`, reports keys cited but not defined and entries defined but cited
       nowhere, and treats `\nocite{*}` as citing everything. Exit 2 when nothing could be read.
    
    ## Constraints
    
    1. Do not import project-specific reference rules from other repositories.
    2. Do not auto-fix reference records without user approval.
    3. Keep output plain Markdown with no emoji.
    4. Treat CrossRef, Semantic Scholar, and arXiv as verification sources, not as citation style authorities.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related