Claude Cursor Skill

export

Convert thesis chapters and reading notes from Markdown to Word (.docx) and package for submission. Use when preparing materials for supervisors or examiners.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download yha9806-academic-writing-toolkit-.claude_skills_export-184e482.zip · 7 KB
Part of yha9806/academic-writing-toolkit — 21 skills

Install

skills CLI npx skills add https://github.com/yha9806/academic-writing-toolkit/tree/main/.claude/skills/export
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install yha9806-academic-writing-toolkit@llmmart
Git git clone https://github.com/yha9806/academic-writing-toolkit.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole yha9806/academic-writing-toolkit collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

/export — Document Export Skill

Running Python helpers

Choose the interpreter before running the examples. For a globally installed copy, use the private runtime recorded by its installer. In a checkout or linked workspace, use AWT_PYTHON when set, otherwise the toolkit's .venv (follow the skill directory link back to the toolkit): Scripts/python.exe on Windows, bin/python on macOS/Linux. Without that environment, check that python (Windows) or python3 (macOS/Linux) actually runs and has the helper's dependencies. Replace the example's python3 with that executable. In PowerShell, prefix a quoted executable with &; keep commands on one line and quote file paths.

Purpose

Convert thesis chapters and reading notes from Markdown to Word (.docx) format and package them into a ZIP archive for submission to supervisors or examiners. This skill is user-invoked only and will not be triggered automatically.

Trigger Words

This skill activates on: /export. It is never invoked automatically by the model.

Parameters

  • $ARGUMENTS[0] -- scope: chapters | notes | all (default: all)
  • $ARGUMENTS[1] -- language filter: en-only | all (default: all)

Examples:

  • /export chapters en-only -- export only chapter files, skip files with significant CJK content
  • /export notes all -- export only reading notes, include all languages
  • /export -- export everything, all languages

Workflow

  1. Resolve every source first (Advisory here; Enforced in the AWT app). Before converting, confirm that every author-year citation in chapters/ has a lint-conforming notes file under literature/reading_notes/, and that references.bib (when present) neither contains entries nothing cites nor lacks an entry a chapter cites:

    python3 .claude/skills/audit/scripts/audit-claim-positioning.py --base-dir chapters --bib references.bib --json
    

    Stop and report if anything is unresolved. On this surface nothing blocks you from exporting anyway — say so if you do. In the AWT app the same condition is a typed denial (EXPORT_SOURCES_UNRESOLVED) on the export_docx tool, and the export does not run.

  2. Check dependencies. Verify that pypandoc is installed AND a working pandoc binary is on PATH (the script smoke-probes via pypandoc.get_pandoc_version()). If pypandoc/pandoc are unavailable, fall back to python-docx + markdown. Report which conversion method is being used.

    Run from the project root.

  3. Run the conversion script.

    python3 .claude/skills/export/scripts/convert_to_docx.py --base-dir "{project_root}" --output-dir "{project_root}/final_output" --scope {scope} --lang-filter {lang_filter}
    
  4. Report results.

    ## Export Complete
    
    | Item | Value |
    |------|-------|
    | Files converted | {N} |
    | Files skipped (language filter) | {N} |
    | Output directory | {path} |
    | ZIP archive | {path} |
    | Total size | {size} |
    
    Conversion method: {pandoc | python-docx fallback}
    

Constraints

  1. User-invoked only. This skill has disable-model-invocation: true and must not run unless the user explicitly calls /export.
  2. No hardcoded paths. All paths are derived from arguments or project root.
  3. No emoji in output.
  4. Preserve formatting as much as possible during conversion -- headings, tables, block quotes, inline code.
Files (academic-writing-toolkit)
  • scripts
    • convert_to_docx.py 16 KB
      #!/usr/bin/env python3
      """Convert thesis chapters and reading notes from Markdown to Word (.docx).
      
      Usage:
          python convert_to_docx.py --base-dir /path/to/project --output-dir /path/to/output
          python convert_to_docx.py --base-dir . --output-dir ./export --scope chapters --lang-filter en-only
      
      Conversion priority:
          1. pypandoc (if available)
          2. python-docx + markdown (fallback)
      """
      from __future__ import annotations
      
      import argparse
      import os
      import re
      import sys
      import zipfile
      from datetime import datetime
      from pathlib import Path
      
      # ---------------------------------------------------------------------------
      # Conversion backends
      # ---------------------------------------------------------------------------
      
      USE_PANDOC = False
      
      try:
          import pypandoc
          # Smoke probe: get_pandoc_version() shells out to `pandoc --version`
          # internally and raises if the binary is missing or broken
          # (Snap isolation, missing libs, broken wrapper).
          pypandoc.get_pandoc_version()
          USE_PANDOC = True
      except (ImportError, OSError, RuntimeError):
          pass
      
      # Always try to import python-docx — needed both as primary path when
      # USE_PANDOC=False, and as runtime fallback when pypandoc fails on a
      # specific file (see convert_file / create_cover try/except below).
      HAS_DOCX = False
      try:
          import markdown
          from docx import Document
          from docx.shared import Inches, Pt
          HAS_DOCX = True
      except ImportError:
          pass
      
      # A bare `pip install` is refused as externally-managed by any PEP 668
      # interpreter — the default on Homebrew Python and on current Debian and
      # Ubuntu — so the remedy leads with a virtual environment. Naming a command
      # the reader cannot run is the same as naming none.
      VENV_COMMAND = r'.\.venv\Scripts\python.exe' if os.name == 'nt' else '.venv/bin/python'
      PYTHON_COMMAND = 'python' if os.name == 'nt' else 'python3'
      BACKEND_ERROR = (
          "Error: No conversion backend available.\n"
          "  - pypandoc is unavailable (not installed, or `pandoc` binary missing/broken on PATH).\n"
          "    Note: the `pandoc` binary alone is not enough; pypandoc is the Python binding.\n"
          "  - python-docx + markdown fallback is also missing.\n"
          "Install one backend into a virtual environment, then run this converter with it:\n"
          f"  {PYTHON_COMMAND} -m venv .venv\n"
          f"  {VENV_COMMAND} -m pip install python-docx markdown\n"
          f'  {VENV_COMMAND} "<this script>" --check\n'
          "Alternatives, if your interpreter is not externally managed:\n"
          f"  {PYTHON_COMMAND} -m pip install --user python-docx markdown\n"
          "  pipx install pypandoc       (also requires the `pandoc` binary on PATH)"
      )
      
      
      def backend_name() -> str:
          """Which backend this interpreter would convert with, or the empty string."""
          if USE_PANDOC:
              return "pypandoc (pandoc)"
          if HAS_DOCX:
              return "python-docx + markdown"
          return ""
      
      
      def ensure_conversion_backend() -> None:
          """Exit with a clear message when no Markdown-to-docx backend is available."""
          if not USE_PANDOC and not HAS_DOCX:
              sys.exit(BACKEND_ERROR)
      
      
      # ---------------------------------------------------------------------------
      # Language detection
      # ---------------------------------------------------------------------------
      
      CJK_PATTERN = re.compile(r"[\u4e00-\u9fff]")
      
      
      def cjk_ratio(text: str) -> float:
          """Return the proportion of CJK characters in *text*."""
          if not text:
              return 0.0
          cjk_count = len(CJK_PATTERN.findall(text))
          # Count only non-whitespace characters for the denominator
          total = len(re.sub(r"\s", "", text))
          if total == 0:
              return 0.0
          return cjk_count / total
      
      
      def should_skip(text: str, lang_filter: str) -> bool:
          """Return True if the file should be skipped under *lang_filter*."""
          if lang_filter != "en-only":
              return False
          return cjk_ratio(text) > 0.01
      
      
      # ---------------------------------------------------------------------------
      # File conversion
      # ---------------------------------------------------------------------------
      
      
      def convert_file(src: Path, dst: Path) -> None:
          """Convert a single Markdown file to .docx."""
          dst.parent.mkdir(parents=True, exist_ok=True)
      
          if USE_PANDOC:
              try:
                  pypandoc.convert_file(
                      str(src),
                      "docx",
                      outputfile=str(dst),
                      extra_args=["--wrap=none"],
                  )
                  return
              except (OSError, RuntimeError) as exc:
                  print(f"  Warning: pandoc failed on {src.name}: {exc}; falling back to python-docx")
          _convert_with_docx(src, dst)
      
      
      def _convert_with_docx(src: Path, dst: Path) -> None:
          """Fallback converter using python-docx + markdown."""
          if not HAS_DOCX:
              raise RuntimeError(
                  "python-docx fallback not available; install python-docx + markdown"
              )
          md_text = src.read_text(encoding="utf-8")
          html = markdown.markdown(
              md_text,
              extensions=["tables", "fenced_code", "footnotes", "toc"],
          )
      
          doc = Document()
      
          # Apply a minimal style
          style = doc.styles["Normal"]
          font = style.font
          font.name = "Times New Roman"
          font.size = Pt(12)
      
          # Simple HTML-to-docx: split by block tags and add paragraphs.
          # This is intentionally simple -- for full fidelity, use pandoc.
          _html_blocks_to_docx(doc, html)
          doc.save(str(dst))
      
      
      def _html_blocks_to_docx(doc, html: str) -> None:
          """Minimal HTML block parser for python-docx fallback."""
          # Strip tags for a basic conversion; heading detection via markdown markers
          # We re-parse from a simple approach: line-by-line
          from html.parser import HTMLParser
      
          class _Parser(HTMLParser):
              def __init__(self, document):
                  super().__init__()
                  self.doc = document
                  self.current_text = ""
                  self.in_heading = 0  # heading level, 0 = not in heading
                  self.in_table = False
                  self.table_rows = []
                  self.current_row = []
                  self.in_cell = False
      
              def handle_starttag(self, tag, attrs):
                  if tag in ("h1", "h2", "h3", "h4", "h5", "h6"):
                      self.in_heading = int(tag[1])
                      self.current_text = ""
                  elif tag == "table":
                      self.in_table = True
                      self.table_rows = []
                  elif tag == "tr":
                      self.current_row = []
                  elif tag in ("td", "th"):
                      self.in_cell = True
                      self.current_text = ""
                  elif tag == "p":
                      self.current_text = ""
                  elif tag == "blockquote":
                      self.current_text = ""
                  elif tag == "br":
                      self.current_text += "\n"
      
              def handle_endtag(self, tag):
                  if tag in ("h1", "h2", "h3", "h4", "h5", "h6"):
                      level = min(self.in_heading, 4)  # docx supports heading 1-4 well
                      self.doc.add_heading(self.current_text.strip(), level=level)
                      self.in_heading = 0
                      self.current_text = ""
                  elif tag == "p":
                      text = self.current_text.strip()
                      if text:
                          self.doc.add_paragraph(text)
                      self.current_text = ""
                  elif tag == "blockquote":
                      text = self.current_text.strip()
                      if text:
                          p = self.doc.add_paragraph(text)
                          p.style = self.doc.styles["Quote"] if "Quote" in [
                              s.name for s in self.doc.styles
                          ] else self.doc.styles["Normal"]
                      self.current_text = ""
                  elif tag in ("td", "th"):
                      self.in_cell = False
                      self.current_row.append(self.current_text.strip())
                      self.current_text = ""
                  elif tag == "tr":
                      self.table_rows.append(self.current_row)
                  elif tag == "table":
                      self.in_table = False
                      if self.table_rows:
                          cols = max(len(r) for r in self.table_rows)
                          table = self.doc.add_table(
                              rows=len(self.table_rows), cols=cols
                          )
                          table.style = "Table Grid"
                          for i, row_data in enumerate(self.table_rows):
                              for j, cell_text in enumerate(row_data):
                                  if j < cols:
                                      table.rows[i].cells[j].text = cell_text
      
              def handle_data(self, data):
                  if self.in_heading or self.in_cell:
                      self.current_text += data
                  else:
                      self.current_text += data
      
          parser = _Parser(doc)
          parser.feed(html)
      
      
      # ---------------------------------------------------------------------------
      # Batch operations
      # ---------------------------------------------------------------------------
      
      
      def convert_chapters(base: Path, out: Path) -> tuple[int, int]:
          """Convert all chapter files. Returns (converted, skipped)."""
          chapters_dir = base / "chapters"
          if not chapters_dir.is_dir():
              print(f"Warning: {chapters_dir} not found, skipping chapters.")
              return 0, 0
          return _convert_dir(chapters_dir, out / "chapters", "ch*.md")
      
      
      def convert_notes(base: Path, out: Path, lang_filter: str) -> tuple[int, int]:
          """Convert reading notes. Returns (converted, skipped)."""
          notes_dir = base / "literature" / "reading_notes"
          if not notes_dir.is_dir():
              print(f"Warning: {notes_dir} not found, skipping notes.")
              return 0, 0
          return _convert_dir(notes_dir, out / "notes", "*_NOTES.md", lang_filter)
      
      
      def _convert_dir(
          src_dir: Path, dst_dir: Path, pattern: str, lang_filter: str = "all"
      ) -> tuple[int, int]:
          """Convert all matching files in a directory. Returns (converted, skipped)."""
          converted = 0
          skipped = 0
          for md_file in sorted(src_dir.glob(pattern)):
              text = md_file.read_text(encoding="utf-8")
              if should_skip(text, lang_filter):
                  print(f"  Skipped (language filter): {md_file.name}")
                  skipped += 1
                  continue
              dst = dst_dir / (md_file.stem + ".docx")
              print(f"  Converting: {md_file.name} -> {dst.name}")
              convert_file(md_file, dst)
              converted += 1
          return converted, skipped
      
      
      def create_cover(metadata: dict, out: Path) -> Path:
          """Create a simple cover page .docx with thesis metadata."""
          dst = out / "00_cover.docx"
          dst.parent.mkdir(parents=True, exist_ok=True)
      
          if USE_PANDOC:
              try:
                  cover_md = f"# {metadata.get('title', 'Thesis')}\n\n"
                  cover_md += f"**Author**: {metadata.get('author', 'Unknown')}\n\n"
                  cover_md += f"**Date**: {metadata.get('date', datetime.now().strftime('%Y-%m-%d'))}\n\n"
                  cover_md += f"**Institution**: {metadata.get('institution', '')}\n\n"
                  if metadata.get("abstract"):
                      cover_md += f"## Abstract\n\n{metadata['abstract']}\n"
                  tmp = out / "_cover_tmp.md"
                  tmp.write_text(cover_md, encoding="utf-8")
                  pypandoc.convert_file(str(tmp), "docx", outputfile=str(dst))
                  tmp.unlink()
                  return dst
              except (OSError, RuntimeError) as exc:
                  print(f"  Warning: pandoc failed on cover: {exc}; falling back to python-docx")
                  if tmp.exists():
                      tmp.unlink()
      
          # Fallback path
          if not HAS_DOCX:
              raise RuntimeError("python-docx fallback not available")
          doc = Document()
          doc.add_heading(metadata.get("title", "Thesis"), level=0)
          doc.add_paragraph(f"Author: {metadata.get('author', 'Unknown')}")
          doc.add_paragraph(
              f"Date: {metadata.get('date', datetime.now().strftime('%Y-%m-%d'))}"
          )
          doc.add_paragraph(f"Institution: {metadata.get('institution', '')}")
          if metadata.get("abstract"):
              doc.add_heading("Abstract", level=1)
              doc.add_paragraph(metadata["abstract"])
          doc.save(str(dst))
      
          return dst
      
      
      def package_zip(out_dir: Path, zip_path: Path) -> None:
          """Create a ZIP archive of the output directory with UTF-8 filenames."""
          zip_path.parent.mkdir(parents=True, exist_ok=True)
          with zipfile.ZipFile(zip_path, "w", zipfile.ZIP_DEFLATED) as zf:
              for root, _dirs, files in os.walk(out_dir):
                  for fname in sorted(files):
                      if fname.endswith(".zip"):
                          continue
                      filepath = Path(root) / fname
                      # Use forward slashes and English folder names
                      arcname = filepath.relative_to(out_dir).as_posix()
                      # Set UTF-8 flag (bit 11) via ZipInfo
                      info = zipfile.ZipInfo(arcname)
                      info.flag_bits |= 0x800  # bit 11 = UTF-8
                      info.compress_type = zipfile.ZIP_DEFLATED
                      with open(filepath, "rb") as f:
                          zf.writestr(info, f.read())
          print(f"  ZIP created: {zip_path} ({zip_path.stat().st_size / 1024:.1f} KB)")
      
      
      # ---------------------------------------------------------------------------
      # CLI
      # ---------------------------------------------------------------------------
      
      
      def main():
          # A piped Windows process otherwise uses the system code page, which
          # cannot represent many valid manuscript paths. Match our UTF-8 inputs
          # and the agent's captured-output decoder for both messages and errors.
          for stream in (sys.stdout, sys.stderr):
              if hasattr(stream, "reconfigure"):
                  stream.reconfigure(encoding="utf-8")
          parser = argparse.ArgumentParser(
              description="Convert thesis Markdown files to Word (.docx)"
          )
          parser.add_argument(
              "--check",
              action="store_true",
              help="Report whether this interpreter has a conversion backend, and exit. "
                   "Converts nothing and needs no project directory.",
          )
          parser.add_argument(
              "--base-dir",
              type=Path,
              required=True,
              help="Project root directory containing chapters/ and literature/",
          )
          parser.add_argument(
              "--output-dir",
              type=Path,
              required=True,
              help="Output directory for converted files",
          )
          parser.add_argument(
              "--scope",
              choices=["chapters", "notes", "all"],
              default="all",
              help="What to convert (default: all)",
          )
          parser.add_argument(
              "--lang-filter",
              choices=["en-only", "all"],
              default="all",
              help="Language filter for notes (default: all)",
          )
          if "--check" in sys.argv[1:]:
              # Answered before parse_args, whose required arguments describe a
              # conversion this mode never performs.
              found = backend_name()
              if not found:
                  print(BACKEND_ERROR, file=sys.stderr)
                  return 1
              print(f"conversion backend available: {found}")
              return 0
      
          args = parser.parse_args()
          ensure_conversion_backend()
      
          base = args.base_dir.resolve()
          out = args.output_dir.resolve()
      
          print(f"Base directory: {base}")
          print(f"Output directory: {out}")
          print(f"Scope: {args.scope}")
          print(f"Language filter: {args.lang_filter}")
          print(f"Conversion method: {'pypandoc (pandoc)' if USE_PANDOC else 'python-docx + markdown (fallback)'}")
          print()
      
          total_converted = 0
          total_skipped = 0
      
          if args.scope in ("chapters", "all"):
              print("--- Chapters ---")
              c, s = convert_chapters(base, out)
              total_converted += c
              total_skipped += s
              print(f"  Chapters: {c} converted, {s} skipped\n")
      
          if args.scope in ("notes", "all"):
              print("--- Reading Notes ---")
              c, s = convert_notes(base, out, args.lang_filter)
              total_converted += c
              total_skipped += s
              print(f"  Notes: {c} converted, {s} skipped\n")
      
          if total_converted > 0:
              # Package into ZIP
              timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
              zip_path = out / f"thesis_export_{timestamp}.zip"
              print("--- Packaging ---")
              package_zip(out, zip_path)
      
          print(f"\nDone. {total_converted} files converted, {total_skipped} skipped.")
      
      
      if __name__ == "__main__":
          # main() returns a status only for --check; a conversion falls off the end
          # and returns None, which sys.exit reads as success.
          sys.exit(main())
      
    • requirements.txt 1 KB
      # Conversion backends for /export. One of the two groups is enough.
      #
      # Declared because nothing else did: CI installed these for itself while no
      # manifest told a reader they existed, so `/export` failed on a machine whose
      # environment check reported healthy. Ask the converter which group you have:
      #
      #   python3 .claude/skills/export/scripts/convert_to_docx.py --check
      #
      # On a PEP 668 interpreter (Homebrew Python, current Debian and Ubuntu), a bare
      # `pip install` is refused. Install into a virtual environment and run the
      # converter with it:
      #
      #   python3 -m venv .venv
      #   .venv/bin/python -m pip install -r .claude/skills/export/scripts/requirements.txt
      #
      # Windows PowerShell uses .\.venv\Scripts\python.exe -m pip instead.
      # From a toolkit checkout, node scripts/setup.mjs prepares the native venv.
      
      # Group 1 — the fallback backend, no external binary needed.
      python-docx
      markdown
      
      # Group 2 — the pandoc backend. Uncomment BOTH the binding below and install
      # the `pandoc` binary separately; the binary alone is not enough.
      # pypandoc
      
  • SKILL.md 3.6 KB
    ---
    name: export
    description: Convert thesis chapters and reading notes from Markdown to Word (.docx) and package for submission. Use when preparing materials for supervisors or examiners.
    disable-model-invocation: true
    allowed-tools: Bash, Read, Glob, Write
    ---
    
    # /export — Document Export Skill
    
    ## Running Python helpers
    
    Choose the interpreter before running the examples. For a globally installed
    copy, use the private runtime recorded by its installer. In a checkout or linked
    workspace, use `AWT_PYTHON` when set, otherwise the toolkit's `.venv` (follow the
    skill directory link back to the toolkit): `Scripts/python.exe` on Windows,
    `bin/python` on macOS/Linux. Without that environment, check that `python`
    (Windows) or `python3` (macOS/Linux) actually runs and has the helper's dependencies.
    Replace the example's `python3` with that executable. In PowerShell, prefix a
    quoted executable with `&`; keep commands on one line and quote file paths.
    
    ## Purpose
    
    Convert thesis chapters and reading notes from Markdown to Word (.docx) format and package them into a ZIP archive for submission to supervisors or examiners. This skill is user-invoked only and will not be triggered automatically.
    
    ## Trigger Words
    
    This skill activates on: `/export`. It is never invoked automatically by the model.
    
    ## Parameters
    
    - `$ARGUMENTS[0]` -- **scope**: `chapters` | `notes` | `all` (default: `all`)
    - `$ARGUMENTS[1]` -- **language filter**: `en-only` | `all` (default: `all`)
    
    Examples:
    - `/export chapters en-only` -- export only chapter files, skip files with significant CJK content
    - `/export notes all` -- export only reading notes, include all languages
    - `/export` -- export everything, all languages
    
    ## Workflow
    
    0. **Resolve every source first (Advisory here; Enforced in the AWT app).**
       Before converting, confirm that every author-year citation in `chapters/`
       has a lint-conforming notes file under `literature/reading_notes/`, and
       that `references.bib` (when present) neither contains entries nothing
       cites nor lacks an entry a chapter cites:
    
       ```
       python3 .claude/skills/audit/scripts/audit-claim-positioning.py --base-dir chapters --bib references.bib --json
       ```
    
       Stop and report if anything is unresolved. On this surface nothing blocks
       you from exporting anyway — say so if you do. In the AWT app the same
       condition is a typed denial (`EXPORT_SOURCES_UNRESOLVED`) on the
       `export_docx` tool, and the export does not run.
    
    1. **Check dependencies.** Verify that `pypandoc` is installed AND a working `pandoc` binary is on PATH (the script smoke-probes via `pypandoc.get_pandoc_version()`). If pypandoc/pandoc are unavailable, fall back to `python-docx` + `markdown`. Report which conversion method is being used.
    
       **Run from the project root.**
    
    2. **Run the conversion script.**
       ```
       python3 .claude/skills/export/scripts/convert_to_docx.py --base-dir "{project_root}" --output-dir "{project_root}/final_output" --scope {scope} --lang-filter {lang_filter}
       ```
    
    3. **Report results.**
       ```
       ## Export Complete
    
       | Item | Value |
       |------|-------|
       | Files converted | {N} |
       | Files skipped (language filter) | {N} |
       | Output directory | {path} |
       | ZIP archive | {path} |
       | Total size | {size} |
    
       Conversion method: {pandoc | python-docx fallback}
       ```
    
    ## Constraints
    
    1. **User-invoked only.** This skill has `disable-model-invocation: true` and must not run unless the user explicitly calls `/export`.
    2. **No hardcoded paths.** All paths are derived from arguments or project root.
    3. **No emoji** in output.
    4. **Preserve formatting** as much as possible during conversion -- headings, tables, block quotes, inline code.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related