Claude Skill

docx

Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction. When Claude needs to work with professional documents (.docx files) for: (1) Creating new documents, (2) Modifying or editing content,

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download shajith003-awesome-claude-skills-document-skills_docx-3d83198.zip · 166 KB
Part of shajith003/awesome-claude-skills — 12 skills

Install

skills CLI npx skills add https://github.com/shajith003/awesome-claude-skills/tree/main/document-skills/docx
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install shajith003-awesome-claude-skills@llmmart
Git git clone https://github.com/shajith003/awesome-claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole shajith003/awesome-claude-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

DOCX creation, editing, and analysis

Overview

A user may ask you to create, edit, or analyze the contents of a .docx file. A .docx file is essentially a ZIP archive containing XML files and other resources that you can read or edit. You have different tools and workflows available for different tasks.

Workflow Decision Tree

Reading/Analyzing Content

Use "Text extraction" or "Raw XML access" sections below

Creating New Document

Use "Creating a new Word document" workflow

Editing Existing Document

  • Your own document + simple changes Use "Basic OOXML editing" workflow

  • Someone else's document Use "Redlining workflow" (recommended default)

  • Legal, academic, business, or government docs Use "Redlining workflow" (required)

Reading and analyzing content

Text extraction

If you just need to read the text contents of a document, you should convert the document to markdown using pandoc. Pandoc provides excellent support for preserving document structure and can show tracked changes:

# Convert document to markdown with tracked changes
pandoc --track-changes=all path-to-file.docx -o output.md
# Options: --track-changes=accept/reject/all

Raw XML access

You need raw XML access for: comments, complex formatting, document structure, embedded media, and metadata. For any of these features, you'll need to unpack a document and read its raw XML contents.

Unpacking a file

python ooxml/scripts/unpack.py <office_file> <output_directory>

Key file structures

  • word/document.xml - Main document contents
  • word/comments.xml - Comments referenced in document.xml
  • word/media/ - Embedded images and media files
  • Tracked changes use <w:ins> (insertions) and <w:del> (deletions) tags

Creating a new Word document

When creating a new Word document from scratch, use docx-js, which allows you to create Word documents using JavaScript/TypeScript.

Workflow

  1. MANDATORY - READ ENTIRE FILE: Read docx-js.md (~500 lines) completely from start to finish. NEVER set any range limits when reading this file. Read the full file content for detailed syntax, critical formatting rules, and best practices before proceeding with document creation.
  2. Create a JavaScript/TypeScript file using Document, Paragraph, TextRun components (You can assume all dependencies are installed, but if not, refer to the dependencies section below)
  3. Export as .docx using Packer.toBuffer()

Editing an existing Word document

When editing an existing Word document, use the Document library (a Python library for OOXML manipulation). The library automatically handles infrastructure setup and provides methods for document manipulation. For complex scenarios, you can access the underlying DOM directly through the library.

Workflow

  1. MANDATORY - READ ENTIRE FILE: Read ooxml.md (~600 lines) completely from start to finish. NEVER set any range limits when reading this file. Read the full file content for the Document library API and XML patterns for directly editing document files.
  2. Unpack the document: python ooxml/scripts/unpack.py <office_file> <output_directory>
  3. Create and run a Python script using the Document library (see "Document Library" section in ooxml.md)
  4. Pack the final document: python ooxml/scripts/pack.py <input_directory> <office_file>

The Document library provides both high-level methods for common operations and direct DOM access for complex scenarios.

Redlining workflow for document review

This workflow allows you to plan comprehensive tracked changes using markdown before implementing them in OOXML. CRITICAL: For complete tracked changes, you must implement ALL changes systematically.

Batching Strategy: Group related changes into batches of 3-10 changes. This makes debugging manageable while maintaining efficiency. Test each batch before moving to the next.

Principle: Minimal, Precise Edits When implementing tracked changes, only mark text that actually changes. Repeating unchanged text makes edits harder to review and appears unprofessional. Break replacements into: [unchanged text] + [deletion] + [insertion] + [unchanged text]. Preserve the original run's RSID for unchanged text by extracting the <w:r> element from the original and reusing it.

Example - Changing "30 days" to "60 days" in a sentence:

# BAD - Replaces entire sentence
'<w:del><w:r><w:delText>The term is 30 days.</w:delText></w:r></w:del><w:ins><w:r><w:t>The term is 60 days.</w:t></w:r></w:ins>'

# GOOD - Only marks what changed, preserves original <w:r> for unchanged text
'<w:r w:rsidR="00AB12CD"><w:t>The term is </w:t></w:r><w:del><w:r><w:delText>30</w:delText></w:r></w:del><w:ins><w:r><w:t>60</w:t></w:r></w:ins><w:r w:rsidR="00AB12CD"><w:t> days.</w:t></w:r>'

Tracked changes workflow

  1. Get markdown representation: Convert document to markdown with tracked changes preserved:

    pandoc --track-changes=all path-to-file.docx -o current.md
    
  2. Identify and group changes: Review the document and identify ALL changes needed, organizing them into logical batches:

    Location methods (for finding changes in XML):

    • Section/heading numbers (e.g., "Section 3.2", "Article IV")
    • Paragraph identifiers if numbered
    • Grep patterns with unique surrounding text
    • Document structure (e.g., "first paragraph", "signature block")
    • DO NOT use markdown line numbers - they don't map to XML structure

    Batch organization (group 3-10 related changes per batch):

    • By section: "Batch 1: Section 2 amendments", "Batch 2: Section 5 updates"
    • By type: "Batch 1: Date corrections", "Batch 2: Party name changes"
    • By complexity: Start with simple text replacements, then tackle complex structural changes
    • Sequential: "Batch 1: Pages 1-3", "Batch 2: Pages 4-6"
  3. Read documentation and unpack:

    • MANDATORY - READ ENTIRE FILE: Read ooxml.md (~600 lines) completely from start to finish. NEVER set any range limits when reading this file. Pay special attention to the "Document Library" and "Tracked Change Patterns" sections.
    • Unpack the document: python ooxml/scripts/unpack.py <file.docx> <dir>
    • Note the suggested RSID: The unpack script will suggest an RSID to use for your tracked changes. Copy this RSID for use in step 4b.
  4. Implement changes in batches: Group changes logically (by section, by type, or by proximity) and implement them together in a single script. This approach:

    • Makes debugging easier (smaller batch = easier to isolate errors)
    • Allows incremental progress
    • Maintains efficiency (batch size of 3-10 changes works well)

    Suggested batch groupings:

    • By document section (e.g., "Section 3 changes", "Definitions", "Termination clause")
    • By change type (e.g., "Date changes", "Party name updates", "Legal term replacements")
    • By proximity (e.g., "Changes on pages 1-3", "Changes in first half of document")

    For each batch of related changes:

    a. Map text to XML: Grep for text in word/document.xml to verify how text is split across <w:r> elements.

    b. Create and run script: Use get_node to find nodes, implement changes, then doc.save(). See "Document Library" section in ooxml.md for patterns.

    Note: Always grep word/document.xml immediately before writing a script to get current line numbers and verify text content. Line numbers change after each script run.

  5. Pack the document: After all batches are complete, convert the unpacked directory back to .docx:

    python ooxml/scripts/pack.py unpacked reviewed-document.docx
    
  6. Final verification: Do a comprehensive check of the complete document:

    • Convert final document to markdown:
      pandoc --track-changes=all reviewed-document.docx -o verification.md
      
    • Verify ALL changes were applied correctly:
      grep "original phrase" verification.md  # Should NOT find it
      grep "replacement phrase" verification.md  # Should find it
      
    • Check that no unintended changes were introduced

Converting Documents to Images

To visually analyze Word documents, convert them to images using a two-step process:

  1. Convert DOCX to PDF:

    soffice --headless --convert-to pdf document.docx
    
  2. Convert PDF pages to JPEG images:

    pdftoppm -jpeg -r 150 document.pdf page
    

    This creates files like page-1.jpg, page-2.jpg, etc.

Options:

  • -r 150: Sets resolution to 150 DPI (adjust for quality/size balance)
  • -jpeg: Output JPEG format (use -png for PNG if preferred)
  • -f N: First page to convert (e.g., -f 2 starts from page 2)
  • -l N: Last page to convert (e.g., -l 5 stops at page 5)
  • page: Prefix for output files

Example for specific range:

pdftoppm -jpeg -r 150 -f 2 -l 5 document.pdf page  # Converts only pages 2-5

Code Style Guidelines

IMPORTANT: When generating code for DOCX operations:

  • Write concise code
  • Avoid verbose variable names and redundant operations
  • Avoid unnecessary print statements

Dependencies

Required dependencies (install if not available):

  • pandoc: sudo apt-get install pandoc (for text extraction)
  • docx: npm install -g docx (for creating new documents)
  • LibreOffice: sudo apt-get install libreoffice (for PDF conversion)
  • Poppler: sudo apt-get install poppler-utils (for pdftoppm to convert PDF to images)
  • defusedxml: pip install defusedxml (for secure XML parsing)
Files (awesome-claude-skills)
  • ooxml
    • schemas
      • ecma
        • fouth-edition
          • opc-contentTypes.xsd 1.9 KB · in bundle
          • opc-coreProperties.xsd 2.5 KB · in bundle
          • opc-digSig.xsd 2.8 KB · in bundle
          • opc-relationships.xsd 1.3 KB · in bundle
      • ISO-IEC29500-4_2016
        • dml-chart.xsd 73.2 KB · in bundle
        • dml-chartDrawing.xsd 6.8 KB · in bundle
        • dml-diagram.xsd 50.1 KB · in bundle
        • dml-lockedCanvas.xsd 624 B · in bundle
        • dml-main.xsd 148.5 KB · in bundle
        • dml-picture.xsd 1.2 KB · in bundle
        • dml-spreadsheetDrawing.xsd 8.7 KB · in bundle
        • dml-wordprocessingDrawing.xsd 14.4 KB · in bundle
        • pml.xsd 81.7 KB · in bundle
        • shared-additionalCharacteristics.xsd 1.2 KB · in bundle
        • shared-bibliography.xsd 7.2 KB · in bundle
        • shared-commonSimpleTypes.xsd 6.2 KB · in bundle
        • shared-customXmlDataProperties.xsd 1.2 KB · in bundle
        • shared-customXmlSchemaProperties.xsd 880 B · in bundle
        • shared-documentPropertiesCustom.xsd 2.5 KB · in bundle
        • shared-documentPropertiesExtended.xsd 3.4 KB · in bundle
        • shared-documentPropertiesVariantTypes.xsd 7.3 KB · in bundle
        • shared-math.xsd 22.8 KB · in bundle
        • shared-relationshipReference.xsd 1.3 KB · in bundle
        • sml.xsd 236.6 KB · in bundle
        • vml-main.xsd 25.5 KB · in bundle
        • vml-officeDrawing.xsd 24.7 KB · in bundle
        • vml-presentationDrawing.xsd 535 B · in bundle
        • vml-spreadsheetDrawing.xsd 5.6 KB · in bundle
        • vml-wordprocessingDrawing.xsd 3.9 KB · in bundle
        • wml.xsd 167.4 KB · in bundle
        • xml.xsd 4.5 KB · in bundle
      • mce
        • mc.xsd 3.1 KB · in bundle
      • microsoft
        • wml-2010.xsd 25.9 KB · in bundle
        • wml-2012.xsd 3.7 KB · in bundle
        • wml-2018.xsd 901 B · in bundle
        • wml-cex-2018.xsd 1.7 KB · in bundle
        • wml-cid-2016.xsd 1002 B · in bundle
        • wml-sdtdatahash-2020.xsd 600 B · in bundle
        • wml-symex-2015.xsd 745 B · in bundle
    • scripts
      • validation
        • base.py 39 KB
          """
          Base validator with common validation logic for document files.
          """
          
          import re
          from pathlib import Path
          
          import lxml.etree
          
          
          class BaseSchemaValidator:
              """Base validator with common validation logic for document files."""
          
              # Elements whose 'id' attributes must be unique within their file
              # Format: element_name -> (attribute_name, scope)
              # scope can be 'file' (unique within file) or 'global' (unique across all files)
              UNIQUE_ID_REQUIREMENTS = {
                  # Word elements
                  "comment": ("id", "file"),  # Comment IDs in comments.xml
                  "commentrangestart": ("id", "file"),  # Must match comment IDs
                  "commentrangeend": ("id", "file"),  # Must match comment IDs
                  "bookmarkstart": ("id", "file"),  # Bookmark start IDs
                  "bookmarkend": ("id", "file"),  # Bookmark end IDs
                  # Note: ins and del (track changes) can share IDs when part of same revision
                  # PowerPoint elements
                  "sldid": ("id", "file"),  # Slide IDs in presentation.xml
                  "sldmasterid": ("id", "global"),  # Slide master IDs must be globally unique
                  "sldlayoutid": ("id", "global"),  # Slide layout IDs must be globally unique
                  "cm": ("authorid", "file"),  # Comment author IDs
                  # Excel elements
                  "sheet": ("sheetid", "file"),  # Sheet IDs in workbook.xml
                  "definedname": ("id", "file"),  # Named range IDs
                  # Drawing/Shape elements (all formats)
                  "cxnsp": ("id", "file"),  # Connection shape IDs
                  "sp": ("id", "file"),  # Shape IDs
                  "pic": ("id", "file"),  # Picture IDs
                  "grpsp": ("id", "file"),  # Group shape IDs
              }
          
              # Mapping of element names to expected relationship types
              # Subclasses should override this with format-specific mappings
              ELEMENT_RELATIONSHIP_TYPES = {}
          
              # Unified schema mappings for all Office document types
              SCHEMA_MAPPINGS = {
                  # Document type specific schemas
                  "word": "ISO-IEC29500-4_2016/wml.xsd",  # Word documents
                  "ppt": "ISO-IEC29500-4_2016/pml.xsd",  # PowerPoint presentations
                  "xl": "ISO-IEC29500-4_2016/sml.xsd",  # Excel spreadsheets
                  # Common file types
                  "[Content_Types].xml": "ecma/fouth-edition/opc-contentTypes.xsd",
                  "app.xml": "ISO-IEC29500-4_2016/shared-documentPropertiesExtended.xsd",
                  "core.xml": "ecma/fouth-edition/opc-coreProperties.xsd",
                  "custom.xml": "ISO-IEC29500-4_2016/shared-documentPropertiesCustom.xsd",
                  ".rels": "ecma/fouth-edition/opc-relationships.xsd",
                  # Word-specific files
                  "people.xml": "microsoft/wml-2012.xsd",
                  "commentsIds.xml": "microsoft/wml-cid-2016.xsd",
                  "commentsExtensible.xml": "microsoft/wml-cex-2018.xsd",
                  "commentsExtended.xml": "microsoft/wml-2012.xsd",
                  # Chart files (common across document types)
                  "chart": "ISO-IEC29500-4_2016/dml-chart.xsd",
                  # Theme files (common across document types)
                  "theme": "ISO-IEC29500-4_2016/dml-main.xsd",
                  # Drawing and media files
                  "drawing": "ISO-IEC29500-4_2016/dml-main.xsd",
              }
          
              # Unified namespace constants
              MC_NAMESPACE = "http://schemas.openxmlformats.org/markup-compatibility/2006"
              XML_NAMESPACE = "http://www.w3.org/XML/1998/namespace"
          
              # Common OOXML namespaces used across validators
              PACKAGE_RELATIONSHIPS_NAMESPACE = (
                  "http://schemas.openxmlformats.org/package/2006/relationships"
              )
              OFFICE_RELATIONSHIPS_NAMESPACE = (
                  "http://schemas.openxmlformats.org/officeDocument/2006/relationships"
              )
              CONTENT_TYPES_NAMESPACE = (
                  "http://schemas.openxmlformats.org/package/2006/content-types"
              )
          
              # Folders where we should clean ignorable namespaces
              MAIN_CONTENT_FOLDERS = {"word", "ppt", "xl"}
          
              # All allowed OOXML namespaces (superset of all document types)
              OOXML_NAMESPACES = {
                  "http://schemas.openxmlformats.org/officeDocument/2006/math",
                  "http://schemas.openxmlformats.org/officeDocument/2006/relationships",
                  "http://schemas.openxmlformats.org/schemaLibrary/2006/main",
                  "http://schemas.openxmlformats.org/drawingml/2006/main",
                  "http://schemas.openxmlformats.org/drawingml/2006/chart",
                  "http://schemas.openxmlformats.org/drawingml/2006/chartDrawing",
                  "http://schemas.openxmlformats.org/drawingml/2006/diagram",
                  "http://schemas.openxmlformats.org/drawingml/2006/picture",
                  "http://schemas.openxmlformats.org/drawingml/2006/spreadsheetDrawing",
                  "http://schemas.openxmlformats.org/drawingml/2006/wordprocessingDrawing",
                  "http://schemas.openxmlformats.org/wordprocessingml/2006/main",
                  "http://schemas.openxmlformats.org/presentationml/2006/main",
                  "http://schemas.openxmlformats.org/spreadsheetml/2006/main",
                  "http://schemas.openxmlformats.org/officeDocument/2006/sharedTypes",
                  "http://www.w3.org/XML/1998/namespace",
              }
          
              def __init__(self, unpacked_dir, original_file, verbose=False):
                  self.unpacked_dir = Path(unpacked_dir).resolve()
                  self.original_file = Path(original_file)
                  self.verbose = verbose
          
                  # Set schemas directory
                  self.schemas_dir = Path(__file__).parent.parent.parent / "schemas"
          
                  # Get all XML and .rels files
                  patterns = ["*.xml", "*.rels"]
                  self.xml_files = [
                      f for pattern in patterns for f in self.unpacked_dir.rglob(pattern)
                  ]
          
                  if not self.xml_files:
                      print(f"Warning: No XML files found in {self.unpacked_dir}")
          
              def validate(self):
                  """Run all validation checks and return True if all pass."""
                  raise NotImplementedError("Subclasses must implement the validate method")
          
              def validate_xml(self):
                  """Validate that all XML files are well-formed."""
                  errors = []
          
                  for xml_file in self.xml_files:
                      try:
                          # Try to parse the XML file
                          lxml.etree.parse(str(xml_file))
                      except lxml.etree.XMLSyntaxError as e:
                          errors.append(
                              f"  {xml_file.relative_to(self.unpacked_dir)}: "
                              f"Line {e.lineno}: {e.msg}"
                          )
                      except Exception as e:
                          errors.append(
                              f"  {xml_file.relative_to(self.unpacked_dir)}: "
                              f"Unexpected error: {str(e)}"
                          )
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} XML violations:")
                      for error in errors:
                          print(error)
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - All XML files are well-formed")
                      return True
          
              def validate_namespaces(self):
                  """Validate that namespace prefixes in Ignorable attributes are declared."""
                  errors = []
          
                  for xml_file in self.xml_files:
                      try:
                          root = lxml.etree.parse(str(xml_file)).getroot()
                          declared = set(root.nsmap.keys()) - {None}  # Exclude default namespace
          
                          for attr_val in [
                              v for k, v in root.attrib.items() if k.endswith("Ignorable")
                          ]:
                              undeclared = set(attr_val.split()) - declared
                              errors.extend(
                                  f"  {xml_file.relative_to(self.unpacked_dir)}: "
                                  f"Namespace '{ns}' in Ignorable but not declared"
                                  for ns in undeclared
                              )
                      except lxml.etree.XMLSyntaxError:
                          continue
          
                  if errors:
                      print(f"FAILED - {len(errors)} namespace issues:")
                      for error in errors:
                          print(error)
                      return False
                  if self.verbose:
                      print("PASSED - All namespace prefixes properly declared")
                  return True
          
              def validate_unique_ids(self):
                  """Validate that specific IDs are unique according to OOXML requirements."""
                  errors = []
                  global_ids = {}  # Track globally unique IDs across all files
          
                  for xml_file in self.xml_files:
                      try:
                          root = lxml.etree.parse(str(xml_file)).getroot()
                          file_ids = {}  # Track IDs that must be unique within this file
          
                          # Remove all mc:AlternateContent elements from the tree
                          mc_elements = root.xpath(
                              ".//mc:AlternateContent", namespaces={"mc": self.MC_NAMESPACE}
                          )
                          for elem in mc_elements:
                              elem.getparent().remove(elem)
          
                          # Now check IDs in the cleaned tree
                          for elem in root.iter():
                              # Get the element name without namespace
                              tag = (
                                  elem.tag.split("}")[-1].lower()
                                  if "}" in elem.tag
                                  else elem.tag.lower()
                              )
          
                              # Check if this element type has ID uniqueness requirements
                              if tag in self.UNIQUE_ID_REQUIREMENTS:
                                  attr_name, scope = self.UNIQUE_ID_REQUIREMENTS[tag]
          
                                  # Look for the specified attribute
                                  id_value = None
                                  for attr, value in elem.attrib.items():
                                      attr_local = (
                                          attr.split("}")[-1].lower()
                                          if "}" in attr
                                          else attr.lower()
                                      )
                                      if attr_local == attr_name:
                                          id_value = value
                                          break
          
                                  if id_value is not None:
                                      if scope == "global":
                                          # Check global uniqueness
                                          if id_value in global_ids:
                                              prev_file, prev_line, prev_tag = global_ids[
                                                  id_value
                                              ]
                                              errors.append(
                                                  f"  {xml_file.relative_to(self.unpacked_dir)}: "
                                                  f"Line {elem.sourceline}: Global ID '{id_value}' in <{tag}> "
                                                  f"already used in {prev_file} at line {prev_line} in <{prev_tag}>"
                                              )
                                          else:
                                              global_ids[id_value] = (
                                                  xml_file.relative_to(self.unpacked_dir),
                                                  elem.sourceline,
                                                  tag,
                                              )
                                      elif scope == "file":
                                          # Check file-level uniqueness
                                          key = (tag, attr_name)
                                          if key not in file_ids:
                                              file_ids[key] = {}
          
                                          if id_value in file_ids[key]:
                                              prev_line = file_ids[key][id_value]
                                              errors.append(
                                                  f"  {xml_file.relative_to(self.unpacked_dir)}: "
                                                  f"Line {elem.sourceline}: Duplicate {attr_name}='{id_value}' in <{tag}> "
                                                  f"(first occurrence at line {prev_line})"
                                              )
                                          else:
                                              file_ids[key][id_value] = elem.sourceline
          
                      except (lxml.etree.XMLSyntaxError, Exception) as e:
                          errors.append(
                              f"  {xml_file.relative_to(self.unpacked_dir)}: Error: {e}"
                          )
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} ID uniqueness violations:")
                      for error in errors:
                          print(error)
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - All required IDs are unique")
                      return True
          
              def validate_file_references(self):
                  """
                  Validate that all .rels files properly reference files and that all files are referenced.
                  """
                  errors = []
          
                  # Find all .rels files
                  rels_files = list(self.unpacked_dir.rglob("*.rels"))
          
                  if not rels_files:
                      if self.verbose:
                          print("PASSED - No .rels files found")
                      return True
          
                  # Get all files in the unpacked directory (excluding reference files)
                  all_files = []
                  for file_path in self.unpacked_dir.rglob("*"):
                      if (
                          file_path.is_file()
                          and file_path.name != "[Content_Types].xml"
                          and not file_path.name.endswith(".rels")
                      ):  # This file is not referenced by .rels
                          all_files.append(file_path.resolve())
          
                  # Track all files that are referenced by any .rels file
                  all_referenced_files = set()
          
                  if self.verbose:
                      print(
                          f"Found {len(rels_files)} .rels files and {len(all_files)} target files"
                      )
          
                  # Check each .rels file
                  for rels_file in rels_files:
                      try:
                          # Parse relationships file
                          rels_root = lxml.etree.parse(str(rels_file)).getroot()
          
                          # Get the directory where this .rels file is located
                          rels_dir = rels_file.parent
          
                          # Find all relationships and their targets
                          referenced_files = set()
                          broken_refs = []
          
                          for rel in rels_root.findall(
                              ".//ns:Relationship",
                              namespaces={"ns": self.PACKAGE_RELATIONSHIPS_NAMESPACE},
                          ):
                              target = rel.get("Target")
                              if target and not target.startswith(
                                  ("http", "mailto:")
                              ):  # Skip external URLs
                                  # Resolve the target path relative to the .rels file location
                                  if rels_file.name == ".rels":
                                      # Root .rels file - targets are relative to unpacked_dir
                                      target_path = self.unpacked_dir / target
                                  else:
                                      # Other .rels files - targets are relative to their parent's parent
                                      # e.g., word/_rels/document.xml.rels -> targets relative to word/
                                      base_dir = rels_dir.parent
                                      target_path = base_dir / target
          
                                  # Normalize the path and check if it exists
                                  try:
                                      target_path = target_path.resolve()
                                      if target_path.exists() and target_path.is_file():
                                          referenced_files.add(target_path)
                                          all_referenced_files.add(target_path)
                                      else:
                                          broken_refs.append((target, rel.sourceline))
                                  except (OSError, ValueError):
                                      broken_refs.append((target, rel.sourceline))
          
                          # Report broken references
                          if broken_refs:
                              rel_path = rels_file.relative_to(self.unpacked_dir)
                              for broken_ref, line_num in broken_refs:
                                  errors.append(
                                      f"  {rel_path}: Line {line_num}: Broken reference to {broken_ref}"
                                  )
          
                      except Exception as e:
                          rel_path = rels_file.relative_to(self.unpacked_dir)
                          errors.append(f"  Error parsing {rel_path}: {e}")
          
                  # Check for unreferenced files (files that exist but are not referenced anywhere)
                  unreferenced_files = set(all_files) - all_referenced_files
          
                  if unreferenced_files:
                      for unref_file in sorted(unreferenced_files):
                          unref_rel_path = unref_file.relative_to(self.unpacked_dir)
                          errors.append(f"  Unreferenced file: {unref_rel_path}")
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} relationship validation errors:")
                      for error in errors:
                          print(error)
                      print(
                          "CRITICAL: These errors will cause the document to appear corrupt. "
                          + "Broken references MUST be fixed, "
                          + "and unreferenced files MUST be referenced or removed."
                      )
                      return False
                  else:
                      if self.verbose:
                          print(
                              "PASSED - All references are valid and all files are properly referenced"
                          )
                      return True
          
              def validate_all_relationship_ids(self):
                  """
                  Validate that all r:id attributes in XML files reference existing IDs
                  in their corresponding .rels files, and optionally validate relationship types.
                  """
                  import lxml.etree
          
                  errors = []
          
                  # Process each XML file that might contain r:id references
                  for xml_file in self.xml_files:
                      # Skip .rels files themselves
                      if xml_file.suffix == ".rels":
                          continue
          
                      # Determine the corresponding .rels file
                      # For dir/file.xml, it's dir/_rels/file.xml.rels
                      rels_dir = xml_file.parent / "_rels"
                      rels_file = rels_dir / f"{xml_file.name}.rels"
          
                      # Skip if there's no corresponding .rels file (that's okay)
                      if not rels_file.exists():
                          continue
          
                      try:
                          # Parse the .rels file to get valid relationship IDs and their types
                          rels_root = lxml.etree.parse(str(rels_file)).getroot()
                          rid_to_type = {}
          
                          for rel in rels_root.findall(
                              f".//{{{self.PACKAGE_RELATIONSHIPS_NAMESPACE}}}Relationship"
                          ):
                              rid = rel.get("Id")
                              rel_type = rel.get("Type", "")
                              if rid:
                                  # Check for duplicate rIds
                                  if rid in rid_to_type:
                                      rels_rel_path = rels_file.relative_to(self.unpacked_dir)
                                      errors.append(
                                          f"  {rels_rel_path}: Line {rel.sourceline}: "
                                          f"Duplicate relationship ID '{rid}' (IDs must be unique)"
                                      )
                                  # Extract just the type name from the full URL
                                  type_name = (
                                      rel_type.split("/")[-1] if "/" in rel_type else rel_type
                                  )
                                  rid_to_type[rid] = type_name
          
                          # Parse the XML file to find all r:id references
                          xml_root = lxml.etree.parse(str(xml_file)).getroot()
          
                          # Find all elements with r:id attributes
                          for elem in xml_root.iter():
                              # Check for r:id attribute (relationship ID)
                              rid_attr = elem.get(f"{{{self.OFFICE_RELATIONSHIPS_NAMESPACE}}}id")
                              if rid_attr:
                                  xml_rel_path = xml_file.relative_to(self.unpacked_dir)
                                  elem_name = (
                                      elem.tag.split("}")[-1] if "}" in elem.tag else elem.tag
                                  )
          
                                  # Check if the ID exists
                                  if rid_attr not in rid_to_type:
                                      errors.append(
                                          f"  {xml_rel_path}: Line {elem.sourceline}: "
                                          f"<{elem_name}> references non-existent relationship '{rid_attr}' "
                                          f"(valid IDs: {', '.join(sorted(rid_to_type.keys())[:5])}{'...' if len(rid_to_type) > 5 else ''})"
                                      )
                                  # Check if we have type expectations for this element
                                  elif self.ELEMENT_RELATIONSHIP_TYPES:
                                      expected_type = self._get_expected_relationship_type(
                                          elem_name
                                      )
                                      if expected_type:
                                          actual_type = rid_to_type[rid_attr]
                                          # Check if the actual type matches or contains the expected type
                                          if expected_type not in actual_type.lower():
                                              errors.append(
                                                  f"  {xml_rel_path}: Line {elem.sourceline}: "
                                                  f"<{elem_name}> references '{rid_attr}' which points to '{actual_type}' "
                                                  f"but should point to a '{expected_type}' relationship"
                                              )
          
                      except Exception as e:
                          xml_rel_path = xml_file.relative_to(self.unpacked_dir)
                          errors.append(f"  Error processing {xml_rel_path}: {e}")
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} relationship ID reference errors:")
                      for error in errors:
                          print(error)
                      print("\nThese ID mismatches will cause the document to appear corrupt!")
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - All relationship ID references are valid")
                      return True
          
              def _get_expected_relationship_type(self, element_name):
                  """
                  Get the expected relationship type for an element.
                  First checks the explicit mapping, then tries pattern detection.
                  """
                  # Normalize element name to lowercase
                  elem_lower = element_name.lower()
          
                  # Check explicit mapping first
                  if elem_lower in self.ELEMENT_RELATIONSHIP_TYPES:
                      return self.ELEMENT_RELATIONSHIP_TYPES[elem_lower]
          
                  # Try pattern detection for common patterns
                  # Pattern 1: Elements ending in "Id" often expect a relationship of the prefix type
                  if elem_lower.endswith("id") and len(elem_lower) > 2:
                      # e.g., "sldId" -> "sld", "sldMasterId" -> "sldMaster"
                      prefix = elem_lower[:-2]  # Remove "id"
                      # Check if this might be a compound like "sldMasterId"
                      if prefix.endswith("master"):
                          return prefix.lower()
                      elif prefix.endswith("layout"):
                          return prefix.lower()
                      else:
                          # Simple case like "sldId" -> "slide"
                          # Common transformations
                          if prefix == "sld":
                              return "slide"
                          return prefix.lower()
          
                  # Pattern 2: Elements ending in "Reference" expect a relationship of the prefix type
                  if elem_lower.endswith("reference") and len(elem_lower) > 9:
                      prefix = elem_lower[:-9]  # Remove "reference"
                      return prefix.lower()
          
                  return None
          
              def validate_content_types(self):
                  """Validate that all content files are properly declared in [Content_Types].xml."""
                  errors = []
          
                  # Find [Content_Types].xml file
                  content_types_file = self.unpacked_dir / "[Content_Types].xml"
                  if not content_types_file.exists():
                      print("FAILED - [Content_Types].xml file not found")
                      return False
          
                  try:
                      # Parse and get all declared parts and extensions
                      root = lxml.etree.parse(str(content_types_file)).getroot()
                      declared_parts = set()
                      declared_extensions = set()
          
                      # Get Override declarations (specific files)
                      for override in root.findall(
                          f".//{{{self.CONTENT_TYPES_NAMESPACE}}}Override"
                      ):
                          part_name = override.get("PartName")
                          if part_name is not None:
                              declared_parts.add(part_name.lstrip("/"))
          
                      # Get Default declarations (by extension)
                      for default in root.findall(
                          f".//{{{self.CONTENT_TYPES_NAMESPACE}}}Default"
                      ):
                          extension = default.get("Extension")
                          if extension is not None:
                              declared_extensions.add(extension.lower())
          
                      # Root elements that require content type declaration
                      declarable_roots = {
                          "sld",
                          "sldLayout",
                          "sldMaster",
                          "presentation",  # PowerPoint
                          "document",  # Word
                          "workbook",
                          "worksheet",  # Excel
                          "theme",  # Common
                      }
          
                      # Common media file extensions that should be declared
                      media_extensions = {
                          "png": "image/png",
                          "jpg": "image/jpeg",
                          "jpeg": "image/jpeg",
                          "gif": "image/gif",
                          "bmp": "image/bmp",
                          "tiff": "image/tiff",
                          "wmf": "image/x-wmf",
                          "emf": "image/x-emf",
                      }
          
                      # Get all files in the unpacked directory
                      all_files = list(self.unpacked_dir.rglob("*"))
                      all_files = [f for f in all_files if f.is_file()]
          
                      # Check all XML files for Override declarations
                      for xml_file in self.xml_files:
                          path_str = str(xml_file.relative_to(self.unpacked_dir)).replace(
                              "\\", "/"
                          )
          
                          # Skip non-content files
                          if any(
                              skip in path_str
                              for skip in [".rels", "[Content_Types]", "docProps/", "_rels/"]
                          ):
                              continue
          
                          try:
                              root_tag = lxml.etree.parse(str(xml_file)).getroot().tag
                              root_name = root_tag.split("}")[-1] if "}" in root_tag else root_tag
          
                              if root_name in declarable_roots and path_str not in declared_parts:
                                  errors.append(
                                      f"  {path_str}: File with <{root_name}> root not declared in [Content_Types].xml"
                                  )
          
                          except Exception:
                              continue  # Skip unparseable files
          
                      # Check all non-XML files for Default extension declarations
                      for file_path in all_files:
                          # Skip XML files and metadata files (already checked above)
                          if file_path.suffix.lower() in {".xml", ".rels"}:
                              continue
                          if file_path.name == "[Content_Types].xml":
                              continue
                          if "_rels" in file_path.parts or "docProps" in file_path.parts:
                              continue
          
                          extension = file_path.suffix.lstrip(".").lower()
                          if extension and extension not in declared_extensions:
                              # Check if it's a known media extension that should be declared
                              if extension in media_extensions:
                                  relative_path = file_path.relative_to(self.unpacked_dir)
                                  errors.append(
                                      f'  {relative_path}: File with extension \'{extension}\' not declared in [Content_Types].xml - should add: <Default Extension="{extension}" ContentType="{media_extensions[extension]}"/>'
                                  )
          
                  except Exception as e:
                      errors.append(f"  Error parsing [Content_Types].xml: {e}")
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} content type declaration errors:")
                      for error in errors:
                          print(error)
                      return False
                  else:
                      if self.verbose:
                          print(
                              "PASSED - All content files are properly declared in [Content_Types].xml"
                          )
                      return True
          
              def validate_file_against_xsd(self, xml_file, verbose=False):
                  """Validate a single XML file against XSD schema, comparing with original.
          
                  Args:
                      xml_file: Path to XML file to validate
                      verbose: Enable verbose output
          
                  Returns:
                      tuple: (is_valid, new_errors_set) where is_valid is True/False/None (skipped)
                  """
                  # Resolve both paths to handle symlinks
                  xml_file = Path(xml_file).resolve()
                  unpacked_dir = self.unpacked_dir.resolve()
          
                  # Validate current file
                  is_valid, current_errors = self._validate_single_file_xsd(
                      xml_file, unpacked_dir
                  )
          
                  if is_valid is None:
                      return None, set()  # Skipped
                  elif is_valid:
                      return True, set()  # Valid, no errors
          
                  # Get errors from original file for this specific file
                  original_errors = self._get_original_file_errors(xml_file)
          
                  # Compare with original (both are guaranteed to be sets here)
                  assert current_errors is not None
                  new_errors = current_errors - original_errors
          
                  if new_errors:
                      if verbose:
                          relative_path = xml_file.relative_to(unpacked_dir)
                          print(f"FAILED - {relative_path}: {len(new_errors)} new error(s)")
                          for error in list(new_errors)[:3]:
                              truncated = error[:250] + "..." if len(error) > 250 else error
                              print(f"  - {truncated}")
                      return False, new_errors
                  else:
                      # All errors existed in original
                      if verbose:
                          print(
                              f"PASSED - No new errors (original had {len(current_errors)} errors)"
                          )
                      return True, set()
          
              def validate_against_xsd(self):
                  """Validate XML files against XSD schemas, showing only new errors compared to original."""
                  new_errors = []
                  original_error_count = 0
                  valid_count = 0
                  skipped_count = 0
          
                  for xml_file in self.xml_files:
                      relative_path = str(xml_file.relative_to(self.unpacked_dir))
                      is_valid, new_file_errors = self.validate_file_against_xsd(
                          xml_file, verbose=False
                      )
          
                      if is_valid is None:
                          skipped_count += 1
                          continue
                      elif is_valid and not new_file_errors:
                          valid_count += 1
                          continue
                      elif is_valid:
                          # Had errors but all existed in original
                          original_error_count += 1
                          valid_count += 1
                          continue
          
                      # Has new errors
                      new_errors.append(f"  {relative_path}: {len(new_file_errors)} new error(s)")
                      for error in list(new_file_errors)[:3]:  # Show first 3 errors
                          new_errors.append(
                              f"    - {error[:250]}..." if len(error) > 250 else f"    - {error}"
                          )
          
                  # Print summary
                  if self.verbose:
                      print(f"Validated {len(self.xml_files)} files:")
                      print(f"  - Valid: {valid_count}")
                      print(f"  - Skipped (no schema): {skipped_count}")
                      if original_error_count:
                          print(f"  - With original errors (ignored): {original_error_count}")
                      print(
                          f"  - With NEW errors: {len(new_errors) > 0 and len([e for e in new_errors if not e.startswith('    ')]) or 0}"
                      )
          
                  if new_errors:
                      print("\nFAILED - Found NEW validation errors:")
                      for error in new_errors:
                          print(error)
                      return False
                  else:
                      if self.verbose:
                          print("\nPASSED - No new XSD validation errors introduced")
                      return True
          
              def _get_schema_path(self, xml_file):
                  """Determine the appropriate schema path for an XML file."""
                  # Check exact filename match
                  if xml_file.name in self.SCHEMA_MAPPINGS:
                      return self.schemas_dir / self.SCHEMA_MAPPINGS[xml_file.name]
          
                  # Check .rels files
                  if xml_file.suffix == ".rels":
                      return self.schemas_dir / self.SCHEMA_MAPPINGS[".rels"]
          
                  # Check chart files
                  if "charts/" in str(xml_file) and xml_file.name.startswith("chart"):
                      return self.schemas_dir / self.SCHEMA_MAPPINGS["chart"]
          
                  # Check theme files
                  if "theme/" in str(xml_file) and xml_file.name.startswith("theme"):
                      return self.schemas_dir / self.SCHEMA_MAPPINGS["theme"]
          
                  # Check if file is in a main content folder and use appropriate schema
                  if xml_file.parent.name in self.MAIN_CONTENT_FOLDERS:
                      return self.schemas_dir / self.SCHEMA_MAPPINGS[xml_file.parent.name]
          
                  return None
          
              def _clean_ignorable_namespaces(self, xml_doc):
                  """Remove attributes and elements not in allowed namespaces."""
                  # Create a clean copy
                  xml_string = lxml.etree.tostring(xml_doc, encoding="unicode")
                  xml_copy = lxml.etree.fromstring(xml_string)
          
                  # Remove attributes not in allowed namespaces
                  for elem in xml_copy.iter():
                      attrs_to_remove = []
          
                      for attr in elem.attrib:
                          # Check if attribute is from a namespace other than allowed ones
                          if "{" in attr:
                              ns = attr.split("}")[0][1:]
                              if ns not in self.OOXML_NAMESPACES:
                                  attrs_to_remove.append(attr)
          
                      # Remove collected attributes
                      for attr in attrs_to_remove:
                          del elem.attrib[attr]
          
                  # Remove elements not in allowed namespaces
                  self._remove_ignorable_elements(xml_copy)
          
                  return lxml.etree.ElementTree(xml_copy)
          
              def _remove_ignorable_elements(self, root):
                  """Recursively remove all elements not in allowed namespaces."""
                  elements_to_remove = []
          
                  # Find elements to remove
                  for elem in list(root):
                      # Skip non-element nodes (comments, processing instructions, etc.)
                      if not hasattr(elem, "tag") or callable(elem.tag):
                          continue
          
                      tag_str = str(elem.tag)
                      if tag_str.startswith("{"):
                          ns = tag_str.split("}")[0][1:]
                          if ns not in self.OOXML_NAMESPACES:
                              elements_to_remove.append(elem)
                              continue
          
                      # Recursively clean child elements
                      self._remove_ignorable_elements(elem)
          
                  # Remove collected elements
                  for elem in elements_to_remove:
                      root.remove(elem)
          
              def _preprocess_for_mc_ignorable(self, xml_doc):
                  """Preprocess XML to handle mc:Ignorable attribute properly."""
                  # Remove mc:Ignorable attributes before validation
                  root = xml_doc.getroot()
          
                  # Remove mc:Ignorable attribute from root
                  if f"{{{self.MC_NAMESPACE}}}Ignorable" in root.attrib:
                      del root.attrib[f"{{{self.MC_NAMESPACE}}}Ignorable"]
          
                  return xml_doc
          
              def _validate_single_file_xsd(self, xml_file, base_path):
                  """Validate a single XML file against XSD schema. Returns (is_valid, errors_set)."""
                  schema_path = self._get_schema_path(xml_file)
                  if not schema_path:
                      return None, None  # Skip file
          
                  try:
                      # Load schema
                      with open(schema_path, "rb") as xsd_file:
                          parser = lxml.etree.XMLParser()
                          xsd_doc = lxml.etree.parse(
                              xsd_file, parser=parser, base_url=str(schema_path)
                          )
                          schema = lxml.etree.XMLSchema(xsd_doc)
          
                      # Load and preprocess XML
                      with open(xml_file, "r") as f:
                          xml_doc = lxml.etree.parse(f)
          
                      xml_doc, _ = self._remove_template_tags_from_text_nodes(xml_doc)
                      xml_doc = self._preprocess_for_mc_ignorable(xml_doc)
          
                      # Clean ignorable namespaces if needed
                      relative_path = xml_file.relative_to(base_path)
                      if (
                          relative_path.parts
                          and relative_path.parts[0] in self.MAIN_CONTENT_FOLDERS
                      ):
                          xml_doc = self._clean_ignorable_namespaces(xml_doc)
          
                      # Validate
                      if schema.validate(xml_doc):
                          return True, set()
                      else:
                          errors = set()
                          for error in schema.error_log:
                              # Store normalized error message (without line numbers for comparison)
                              errors.add(error.message)
                          return False, errors
          
                  except Exception as e:
                      return False, {str(e)}
          
              def _get_original_file_errors(self, xml_file):
                  """Get XSD validation errors from a single file in the original document.
          
                  Args:
                      xml_file: Path to the XML file in unpacked_dir to check
          
                  Returns:
                      set: Set of error messages from the original file
                  """
                  import tempfile
                  import zipfile
          
                  # Resolve both paths to handle symlinks (e.g., /var vs /private/var on macOS)
                  xml_file = Path(xml_file).resolve()
                  unpacked_dir = self.unpacked_dir.resolve()
                  relative_path = xml_file.relative_to(unpacked_dir)
          
                  with tempfile.TemporaryDirectory() as temp_dir:
                      temp_path = Path(temp_dir)
          
                      # Extract original file
                      with zipfile.ZipFile(self.original_file, "r") as zip_ref:
                          zip_ref.extractall(temp_path)
          
                      # Find corresponding file in original
                      original_xml_file = temp_path / relative_path
          
                      if not original_xml_file.exists():
                          # File didn't exist in original, so no original errors
                          return set()
          
                      # Validate the specific file in original
                      is_valid, errors = self._validate_single_file_xsd(
                          original_xml_file, temp_path
                      )
                      return errors if errors else set()
          
              def _remove_template_tags_from_text_nodes(self, xml_doc):
                  """Remove template tags from XML text nodes and collect warnings.
          
                  Template tags follow the pattern {{ ... }} and are used as placeholders
                  for content replacement. They should be removed from text content before
                  XSD validation while preserving XML structure.
          
                  Returns:
                      tuple: (cleaned_xml_doc, warnings_list)
                  """
                  warnings = []
                  template_pattern = re.compile(r"\{\{[^}]*\}\}")
          
                  # Create a copy of the document to avoid modifying the original
                  xml_string = lxml.etree.tostring(xml_doc, encoding="unicode")
                  xml_copy = lxml.etree.fromstring(xml_string)
          
                  def process_text_content(text, content_type):
                      if not text:
                          return text
                      matches = list(template_pattern.finditer(text))
                      if matches:
                          for match in matches:
                              warnings.append(
                                  f"Found template tag in {content_type}: {match.group()}"
                              )
                          return template_pattern.sub("", text)
                      return text
          
                  # Process all text nodes in the document
                  for elem in xml_copy.iter():
                      # Skip processing if this is a w:t element
                      if not hasattr(elem, "tag") or callable(elem.tag):
                          continue
                      tag_str = str(elem.tag)
                      if tag_str.endswith("}t") or tag_str == "t":
                          continue
          
                      elem.text = process_text_content(elem.text, "text content")
                      elem.tail = process_text_content(elem.tail, "tail content")
          
                  return lxml.etree.ElementTree(xml_copy), warnings
          
          
          if __name__ == "__main__":
              raise RuntimeError("This module should not be run directly.")
          
        • docx.py 9.8 KB
          """
          Validator for Word document XML files against XSD schemas.
          """
          
          import re
          import tempfile
          import zipfile
          
          import lxml.etree
          
          from .base import BaseSchemaValidator
          
          
          class DOCXSchemaValidator(BaseSchemaValidator):
              """Validator for Word document XML files against XSD schemas."""
          
              # Word-specific namespace
              WORD_2006_NAMESPACE = "http://schemas.openxmlformats.org/wordprocessingml/2006/main"
          
              # Word-specific element to relationship type mappings
              # Start with empty mapping - add specific cases as we discover them
              ELEMENT_RELATIONSHIP_TYPES = {}
          
              def validate(self):
                  """Run all validation checks and return True if all pass."""
                  # Test 0: XML well-formedness
                  if not self.validate_xml():
                      return False
          
                  # Test 1: Namespace declarations
                  all_valid = True
                  if not self.validate_namespaces():
                      all_valid = False
          
                  # Test 2: Unique IDs
                  if not self.validate_unique_ids():
                      all_valid = False
          
                  # Test 3: Relationship and file reference validation
                  if not self.validate_file_references():
                      all_valid = False
          
                  # Test 4: Content type declarations
                  if not self.validate_content_types():
                      all_valid = False
          
                  # Test 5: XSD schema validation
                  if not self.validate_against_xsd():
                      all_valid = False
          
                  # Test 6: Whitespace preservation
                  if not self.validate_whitespace_preservation():
                      all_valid = False
          
                  # Test 7: Deletion validation
                  if not self.validate_deletions():
                      all_valid = False
          
                  # Test 8: Insertion validation
                  if not self.validate_insertions():
                      all_valid = False
          
                  # Test 9: Relationship ID reference validation
                  if not self.validate_all_relationship_ids():
                      all_valid = False
          
                  # Count and compare paragraphs
                  self.compare_paragraph_counts()
          
                  return all_valid
          
              def validate_whitespace_preservation(self):
                  """
                  Validate that w:t elements with whitespace have xml:space='preserve'.
                  """
                  errors = []
          
                  for xml_file in self.xml_files:
                      # Only check document.xml files
                      if xml_file.name != "document.xml":
                          continue
          
                      try:
                          root = lxml.etree.parse(str(xml_file)).getroot()
          
                          # Find all w:t elements
                          for elem in root.iter(f"{{{self.WORD_2006_NAMESPACE}}}t"):
                              if elem.text:
                                  text = elem.text
                                  # Check if text starts or ends with whitespace
                                  if re.match(r"^\s.*", text) or re.match(r".*\s$", text):
                                      # Check if xml:space="preserve" attribute exists
                                      xml_space_attr = f"{{{self.XML_NAMESPACE}}}space"
                                      if (
                                          xml_space_attr not in elem.attrib
                                          or elem.attrib[xml_space_attr] != "preserve"
                                      ):
                                          # Show a preview of the text
                                          text_preview = (
                                              repr(text)[:50] + "..."
                                              if len(repr(text)) > 50
                                              else repr(text)
                                          )
                                          errors.append(
                                              f"  {xml_file.relative_to(self.unpacked_dir)}: "
                                              f"Line {elem.sourceline}: w:t element with whitespace missing xml:space='preserve': {text_preview}"
                                          )
          
                      except (lxml.etree.XMLSyntaxError, Exception) as e:
                          errors.append(
                              f"  {xml_file.relative_to(self.unpacked_dir)}: Error: {e}"
                          )
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} whitespace preservation violations:")
                      for error in errors:
                          print(error)
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - All whitespace is properly preserved")
                      return True
          
              def validate_deletions(self):
                  """
                  Validate that w:t elements are not within w:del elements.
                  For some reason, XSD validation does not catch this, so we do it manually.
                  """
                  errors = []
          
                  for xml_file in self.xml_files:
                      # Only check document.xml files
                      if xml_file.name != "document.xml":
                          continue
          
                      try:
                          root = lxml.etree.parse(str(xml_file)).getroot()
          
                          # Find all w:t elements that are descendants of w:del elements
                          namespaces = {"w": self.WORD_2006_NAMESPACE}
                          xpath_expression = ".//w:del//w:t"
                          problematic_t_elements = root.xpath(
                              xpath_expression, namespaces=namespaces
                          )
                          for t_elem in problematic_t_elements:
                              if t_elem.text:
                                  # Show a preview of the text
                                  text_preview = (
                                      repr(t_elem.text)[:50] + "..."
                                      if len(repr(t_elem.text)) > 50
                                      else repr(t_elem.text)
                                  )
                                  errors.append(
                                      f"  {xml_file.relative_to(self.unpacked_dir)}: "
                                      f"Line {t_elem.sourceline}: <w:t> found within <w:del>: {text_preview}"
                                  )
          
                      except (lxml.etree.XMLSyntaxError, Exception) as e:
                          errors.append(
                              f"  {xml_file.relative_to(self.unpacked_dir)}: Error: {e}"
                          )
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} deletion validation violations:")
                      for error in errors:
                          print(error)
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - No w:t elements found within w:del elements")
                      return True
          
              def count_paragraphs_in_unpacked(self):
                  """Count the number of paragraphs in the unpacked document."""
                  count = 0
          
                  for xml_file in self.xml_files:
                      # Only check document.xml files
                      if xml_file.name != "document.xml":
                          continue
          
                      try:
                          root = lxml.etree.parse(str(xml_file)).getroot()
                          # Count all w:p elements
                          paragraphs = root.findall(f".//{{{self.WORD_2006_NAMESPACE}}}p")
                          count = len(paragraphs)
                      except Exception as e:
                          print(f"Error counting paragraphs in unpacked document: {e}")
          
                  return count
          
              def count_paragraphs_in_original(self):
                  """Count the number of paragraphs in the original docx file."""
                  count = 0
          
                  try:
                      # Create temporary directory to unpack original
                      with tempfile.TemporaryDirectory() as temp_dir:
                          # Unpack original docx
                          with zipfile.ZipFile(self.original_file, "r") as zip_ref:
                              zip_ref.extractall(temp_dir)
          
                          # Parse document.xml
                          doc_xml_path = temp_dir + "/word/document.xml"
                          root = lxml.etree.parse(doc_xml_path).getroot()
          
                          # Count all w:p elements
                          paragraphs = root.findall(f".//{{{self.WORD_2006_NAMESPACE}}}p")
                          count = len(paragraphs)
          
                  except Exception as e:
                      print(f"Error counting paragraphs in original document: {e}")
          
                  return count
          
              def validate_insertions(self):
                  """
                  Validate that w:delText elements are not within w:ins elements.
                  w:delText is only allowed in w:ins if nested within a w:del.
                  """
                  errors = []
          
                  for xml_file in self.xml_files:
                      if xml_file.name != "document.xml":
                          continue
          
                      try:
                          root = lxml.etree.parse(str(xml_file)).getroot()
                          namespaces = {"w": self.WORD_2006_NAMESPACE}
          
                          # Find w:delText in w:ins that are NOT within w:del
                          invalid_elements = root.xpath(
                              ".//w:ins//w:delText[not(ancestor::w:del)]",
                              namespaces=namespaces
                          )
          
                          for elem in invalid_elements:
                              text_preview = (
                                  repr(elem.text or "")[:50] + "..."
                                  if len(repr(elem.text or "")) > 50
                                  else repr(elem.text or "")
                              )
                              errors.append(
                                  f"  {xml_file.relative_to(self.unpacked_dir)}: "
                                  f"Line {elem.sourceline}: <w:delText> within <w:ins>: {text_preview}"
                              )
          
                      except (lxml.etree.XMLSyntaxError, Exception) as e:
                          errors.append(
                              f"  {xml_file.relative_to(self.unpacked_dir)}: Error: {e}"
                          )
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} insertion validation violations:")
                      for error in errors:
                          print(error)
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - No w:delText elements within w:ins elements")
                      return True
          
              def compare_paragraph_counts(self):
                  """Compare paragraph counts between original and new document."""
                  original_count = self.count_paragraphs_in_original()
                  new_count = self.count_paragraphs_in_unpacked()
          
                  diff = new_count - original_count
                  diff_str = f"+{diff}" if diff > 0 else str(diff)
                  print(f"\nParagraphs: {original_count} → {new_count} ({diff_str})")
          
          
          if __name__ == "__main__":
              raise RuntimeError("This module should not be run directly.")
          
        • pptx.py 12 KB
          """
          Validator for PowerPoint presentation XML files against XSD schemas.
          """
          
          import re
          
          from .base import BaseSchemaValidator
          
          
          class PPTXSchemaValidator(BaseSchemaValidator):
              """Validator for PowerPoint presentation XML files against XSD schemas."""
          
              # PowerPoint presentation namespace
              PRESENTATIONML_NAMESPACE = (
                  "http://schemas.openxmlformats.org/presentationml/2006/main"
              )
          
              # PowerPoint-specific element to relationship type mappings
              ELEMENT_RELATIONSHIP_TYPES = {
                  "sldid": "slide",
                  "sldmasterid": "slidemaster",
                  "notesmasterid": "notesmaster",
                  "sldlayoutid": "slidelayout",
                  "themeid": "theme",
                  "tablestyleid": "tablestyles",
              }
          
              def validate(self):
                  """Run all validation checks and return True if all pass."""
                  # Test 0: XML well-formedness
                  if not self.validate_xml():
                      return False
          
                  # Test 1: Namespace declarations
                  all_valid = True
                  if not self.validate_namespaces():
                      all_valid = False
          
                  # Test 2: Unique IDs
                  if not self.validate_unique_ids():
                      all_valid = False
          
                  # Test 3: UUID ID validation
                  if not self.validate_uuid_ids():
                      all_valid = False
          
                  # Test 4: Relationship and file reference validation
                  if not self.validate_file_references():
                      all_valid = False
          
                  # Test 5: Slide layout ID validation
                  if not self.validate_slide_layout_ids():
                      all_valid = False
          
                  # Test 6: Content type declarations
                  if not self.validate_content_types():
                      all_valid = False
          
                  # Test 7: XSD schema validation
                  if not self.validate_against_xsd():
                      all_valid = False
          
                  # Test 8: Notes slide reference validation
                  if not self.validate_notes_slide_references():
                      all_valid = False
          
                  # Test 9: Relationship ID reference validation
                  if not self.validate_all_relationship_ids():
                      all_valid = False
          
                  # Test 10: Duplicate slide layout references validation
                  if not self.validate_no_duplicate_slide_layouts():
                      all_valid = False
          
                  return all_valid
          
              def validate_uuid_ids(self):
                  """Validate that ID attributes that look like UUIDs contain only hex values."""
                  import lxml.etree
          
                  errors = []
                  # UUID pattern: 8-4-4-4-12 hex digits with optional braces/hyphens
                  uuid_pattern = re.compile(
                      r"^[\{\(]?[0-9A-Fa-f]{8}-?[0-9A-Fa-f]{4}-?[0-9A-Fa-f]{4}-?[0-9A-Fa-f]{4}-?[0-9A-Fa-f]{12}[\}\)]?$"
                  )
          
                  for xml_file in self.xml_files:
                      try:
                          root = lxml.etree.parse(str(xml_file)).getroot()
          
                          # Check all elements for ID attributes
                          for elem in root.iter():
                              for attr, value in elem.attrib.items():
                                  # Check if this is an ID attribute
                                  attr_name = attr.split("}")[-1].lower()
                                  if attr_name == "id" or attr_name.endswith("id"):
                                      # Check if value looks like a UUID (has the right length and pattern structure)
                                      if self._looks_like_uuid(value):
                                          # Validate that it contains only hex characters in the right positions
                                          if not uuid_pattern.match(value):
                                              errors.append(
                                                  f"  {xml_file.relative_to(self.unpacked_dir)}: "
                                                  f"Line {elem.sourceline}: ID '{value}' appears to be a UUID but contains invalid hex characters"
                                              )
          
                      except (lxml.etree.XMLSyntaxError, Exception) as e:
                          errors.append(
                              f"  {xml_file.relative_to(self.unpacked_dir)}: Error: {e}"
                          )
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} UUID ID validation errors:")
                      for error in errors:
                          print(error)
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - All UUID-like IDs contain valid hex values")
                      return True
          
              def _looks_like_uuid(self, value):
                  """Check if a value has the general structure of a UUID."""
                  # Remove common UUID delimiters
                  clean_value = value.strip("{}()").replace("-", "")
                  # Check if it's 32 hex-like characters (could include invalid hex chars)
                  return len(clean_value) == 32 and all(c.isalnum() for c in clean_value)
          
              def validate_slide_layout_ids(self):
                  """Validate that sldLayoutId elements in slide masters reference valid slide layouts."""
                  import lxml.etree
          
                  errors = []
          
                  # Find all slide master files
                  slide_masters = list(self.unpacked_dir.glob("ppt/slideMasters/*.xml"))
          
                  if not slide_masters:
                      if self.verbose:
                          print("PASSED - No slide masters found")
                      return True
          
                  for slide_master in slide_masters:
                      try:
                          # Parse the slide master file
                          root = lxml.etree.parse(str(slide_master)).getroot()
          
                          # Find the corresponding _rels file for this slide master
                          rels_file = slide_master.parent / "_rels" / f"{slide_master.name}.rels"
          
                          if not rels_file.exists():
                              errors.append(
                                  f"  {slide_master.relative_to(self.unpacked_dir)}: "
                                  f"Missing relationships file: {rels_file.relative_to(self.unpacked_dir)}"
                              )
                              continue
          
                          # Parse the relationships file
                          rels_root = lxml.etree.parse(str(rels_file)).getroot()
          
                          # Build a set of valid relationship IDs that point to slide layouts
                          valid_layout_rids = set()
                          for rel in rels_root.findall(
                              f".//{{{self.PACKAGE_RELATIONSHIPS_NAMESPACE}}}Relationship"
                          ):
                              rel_type = rel.get("Type", "")
                              if "slideLayout" in rel_type:
                                  valid_layout_rids.add(rel.get("Id"))
          
                          # Find all sldLayoutId elements in the slide master
                          for sld_layout_id in root.findall(
                              f".//{{{self.PRESENTATIONML_NAMESPACE}}}sldLayoutId"
                          ):
                              r_id = sld_layout_id.get(
                                  f"{{{self.OFFICE_RELATIONSHIPS_NAMESPACE}}}id"
                              )
                              layout_id = sld_layout_id.get("id")
          
                              if r_id and r_id not in valid_layout_rids:
                                  errors.append(
                                      f"  {slide_master.relative_to(self.unpacked_dir)}: "
                                      f"Line {sld_layout_id.sourceline}: sldLayoutId with id='{layout_id}' "
                                      f"references r:id='{r_id}' which is not found in slide layout relationships"
                                  )
          
                      except (lxml.etree.XMLSyntaxError, Exception) as e:
                          errors.append(
                              f"  {slide_master.relative_to(self.unpacked_dir)}: Error: {e}"
                          )
          
                  if errors:
                      print(f"FAILED - Found {len(errors)} slide layout ID validation errors:")
                      for error in errors:
                          print(error)
                      print(
                          "Remove invalid references or add missing slide layouts to the relationships file."
                      )
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - All slide layout IDs reference valid slide layouts")
                      return True
          
              def validate_no_duplicate_slide_layouts(self):
                  """Validate that each slide has exactly one slideLayout reference."""
                  import lxml.etree
          
                  errors = []
                  slide_rels_files = list(self.unpacked_dir.glob("ppt/slides/_rels/*.xml.rels"))
          
                  for rels_file in slide_rels_files:
                      try:
                          root = lxml.etree.parse(str(rels_file)).getroot()
          
                          # Find all slideLayout relationships
                          layout_rels = [
                              rel
                              for rel in root.findall(
                                  f".//{{{self.PACKAGE_RELATIONSHIPS_NAMESPACE}}}Relationship"
                              )
                              if "slideLayout" in rel.get("Type", "")
                          ]
          
                          if len(layout_rels) > 1:
                              errors.append(
                                  f"  {rels_file.relative_to(self.unpacked_dir)}: has {len(layout_rels)} slideLayout references"
                              )
          
                      except Exception as e:
                          errors.append(
                              f"  {rels_file.relative_to(self.unpacked_dir)}: Error: {e}"
                          )
          
                  if errors:
                      print("FAILED - Found slides with duplicate slideLayout references:")
                      for error in errors:
                          print(error)
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - All slides have exactly one slideLayout reference")
                      return True
          
              def validate_notes_slide_references(self):
                  """Validate that each notesSlide file is referenced by only one slide."""
                  import lxml.etree
          
                  errors = []
                  notes_slide_references = {}  # Track which slides reference each notesSlide
          
                  # Find all slide relationship files
                  slide_rels_files = list(self.unpacked_dir.glob("ppt/slides/_rels/*.xml.rels"))
          
                  if not slide_rels_files:
                      if self.verbose:
                          print("PASSED - No slide relationship files found")
                      return True
          
                  for rels_file in slide_rels_files:
                      try:
                          # Parse the relationships file
                          root = lxml.etree.parse(str(rels_file)).getroot()
          
                          # Find all notesSlide relationships
                          for rel in root.findall(
                              f".//{{{self.PACKAGE_RELATIONSHIPS_NAMESPACE}}}Relationship"
                          ):
                              rel_type = rel.get("Type", "")
                              if "notesSlide" in rel_type:
                                  target = rel.get("Target", "")
                                  if target:
                                      # Normalize the target path to handle relative paths
                                      normalized_target = target.replace("../", "")
          
                                      # Track which slide references this notesSlide
                                      slide_name = rels_file.stem.replace(
                                          ".xml", ""
                                      )  # e.g., "slide1"
          
                                      if normalized_target not in notes_slide_references:
                                          notes_slide_references[normalized_target] = []
                                      notes_slide_references[normalized_target].append(
                                          (slide_name, rels_file)
                                      )
          
                      except (lxml.etree.XMLSyntaxError, Exception) as e:
                          errors.append(
                              f"  {rels_file.relative_to(self.unpacked_dir)}: Error: {e}"
                          )
          
                  # Check for duplicate references
                  for target, references in notes_slide_references.items():
                      if len(references) > 1:
                          slide_names = [ref[0] for ref in references]
                          errors.append(
                              f"  Notes slide '{target}' is referenced by multiple slides: {', '.join(slide_names)}"
                          )
                          for slide_name, rels_file in references:
                              errors.append(f"    - {rels_file.relative_to(self.unpacked_dir)}")
          
                  if errors:
                      print(
                          f"FAILED - Found {len([e for e in errors if not e.startswith('    ')])} notes slide reference validation errors:"
                      )
                      for error in errors:
                          print(error)
                      print("Each slide may optionally have its own slide file.")
                      return False
                  else:
                      if self.verbose:
                          print("PASSED - All notes slide references are unique")
                      return True
          
          
          if __name__ == "__main__":
              raise RuntimeError("This module should not be run directly.")
          
        • redlining.py 10.9 KB
          """
          Validator for tracked changes in Word documents.
          """
          
          import subprocess
          import tempfile
          import zipfile
          from pathlib import Path
          
          
          class RedliningValidator:
              """Validator for tracked changes in Word documents."""
          
              def __init__(self, unpacked_dir, original_docx, verbose=False):
                  self.unpacked_dir = Path(unpacked_dir)
                  self.original_docx = Path(original_docx)
                  self.verbose = verbose
                  self.namespaces = {
                      "w": "http://schemas.openxmlformats.org/wordprocessingml/2006/main"
                  }
          
              def validate(self):
                  """Main validation method that returns True if valid, False otherwise."""
                  # Verify unpacked directory exists and has correct structure
                  modified_file = self.unpacked_dir / "word" / "document.xml"
                  if not modified_file.exists():
                      print(f"FAILED - Modified document.xml not found at {modified_file}")
                      return False
          
                  # First, check if there are any tracked changes by Claude to validate
                  try:
                      import xml.etree.ElementTree as ET
          
                      tree = ET.parse(modified_file)
                      root = tree.getroot()
          
                      # Check for w:del or w:ins tags authored by Claude
                      del_elements = root.findall(".//w:del", self.namespaces)
                      ins_elements = root.findall(".//w:ins", self.namespaces)
          
                      # Filter to only include changes by Claude
                      claude_del_elements = [
                          elem
                          for elem in del_elements
                          if elem.get(f"{{{self.namespaces['w']}}}author") == "Claude"
                      ]
                      claude_ins_elements = [
                          elem
                          for elem in ins_elements
                          if elem.get(f"{{{self.namespaces['w']}}}author") == "Claude"
                      ]
          
                      # Redlining validation is only needed if tracked changes by Claude have been used.
                      if not claude_del_elements and not claude_ins_elements:
                          if self.verbose:
                              print("PASSED - No tracked changes by Claude found.")
                          return True
          
                  except Exception:
                      # If we can't parse the XML, continue with full validation
                      pass
          
                  # Create temporary directory for unpacking original docx
                  with tempfile.TemporaryDirectory() as temp_dir:
                      temp_path = Path(temp_dir)
          
                      # Unpack original docx
                      try:
                          with zipfile.ZipFile(self.original_docx, "r") as zip_ref:
                              zip_ref.extractall(temp_path)
                      except Exception as e:
                          print(f"FAILED - Error unpacking original docx: {e}")
                          return False
          
                      original_file = temp_path / "word" / "document.xml"
                      if not original_file.exists():
                          print(
                              f"FAILED - Original document.xml not found in {self.original_docx}"
                          )
                          return False
          
                      # Parse both XML files using xml.etree.ElementTree for redlining validation
                      try:
                          import xml.etree.ElementTree as ET
          
                          modified_tree = ET.parse(modified_file)
                          modified_root = modified_tree.getroot()
                          original_tree = ET.parse(original_file)
                          original_root = original_tree.getroot()
                      except ET.ParseError as e:
                          print(f"FAILED - Error parsing XML files: {e}")
                          return False
          
                      # Remove Claude's tracked changes from both documents
                      self._remove_claude_tracked_changes(original_root)
                      self._remove_claude_tracked_changes(modified_root)
          
                      # Extract and compare text content
                      modified_text = self._extract_text_content(modified_root)
                      original_text = self._extract_text_content(original_root)
          
                      if modified_text != original_text:
                          # Show detailed character-level differences for each paragraph
                          error_message = self._generate_detailed_diff(
                              original_text, modified_text
                          )
                          print(error_message)
                          return False
          
                      if self.verbose:
                          print("PASSED - All changes by Claude are properly tracked")
                      return True
          
              def _generate_detailed_diff(self, original_text, modified_text):
                  """Generate detailed word-level differences using git word diff."""
                  error_parts = [
                      "FAILED - Document text doesn't match after removing Claude's tracked changes",
                      "",
                      "Likely causes:",
                      "  1. Modified text inside another author's <w:ins> or <w:del> tags",
                      "  2. Made edits without proper tracked changes",
                      "  3. Didn't nest <w:del> inside <w:ins> when deleting another's insertion",
                      "",
                      "For pre-redlined documents, use correct patterns:",
                      "  - To reject another's INSERTION: Nest <w:del> inside their <w:ins>",
                      "  - To restore another's DELETION: Add new <w:ins> AFTER their <w:del>",
                      "",
                  ]
          
                  # Show git word diff
                  git_diff = self._get_git_word_diff(original_text, modified_text)
                  if git_diff:
                      error_parts.extend(["Differences:", "============", git_diff])
                  else:
                      error_parts.append("Unable to generate word diff (git not available)")
          
                  return "\n".join(error_parts)
          
              def _get_git_word_diff(self, original_text, modified_text):
                  """Generate word diff using git with character-level precision."""
                  try:
                      with tempfile.TemporaryDirectory() as temp_dir:
                          temp_path = Path(temp_dir)
          
                          # Create two files
                          original_file = temp_path / "original.txt"
                          modified_file = temp_path / "modified.txt"
          
                          original_file.write_text(original_text, encoding="utf-8")
                          modified_file.write_text(modified_text, encoding="utf-8")
          
                          # Try character-level diff first for precise differences
                          result = subprocess.run(
                              [
                                  "git",
                                  "diff",
                                  "--word-diff=plain",
                                  "--word-diff-regex=.",  # Character-by-character diff
                                  "-U0",  # Zero lines of context - show only changed lines
                                  "--no-index",
                                  str(original_file),
                                  str(modified_file),
                              ],
                              capture_output=True,
                              text=True,
                          )
          
                          if result.stdout.strip():
                              # Clean up the output - remove git diff header lines
                              lines = result.stdout.split("\n")
                              # Skip the header lines (diff --git, index, +++, ---, @@)
                              content_lines = []
                              in_content = False
                              for line in lines:
                                  if line.startswith("@@"):
                                      in_content = True
                                      continue
                                  if in_content and line.strip():
                                      content_lines.append(line)
          
                              if content_lines:
                                  return "\n".join(content_lines)
          
                          # Fallback to word-level diff if character-level is too verbose
                          result = subprocess.run(
                              [
                                  "git",
                                  "diff",
                                  "--word-diff=plain",
                                  "-U0",  # Zero lines of context
                                  "--no-index",
                                  str(original_file),
                                  str(modified_file),
                              ],
                              capture_output=True,
                              text=True,
                          )
          
                          if result.stdout.strip():
                              lines = result.stdout.split("\n")
                              content_lines = []
                              in_content = False
                              for line in lines:
                                  if line.startswith("@@"):
                                      in_content = True
                                      continue
                                  if in_content and line.strip():
                                      content_lines.append(line)
                              return "\n".join(content_lines)
          
                  except (subprocess.CalledProcessError, FileNotFoundError, Exception):
                      # Git not available or other error, return None to use fallback
                      pass
          
                  return None
          
              def _remove_claude_tracked_changes(self, root):
                  """Remove tracked changes authored by Claude from the XML root."""
                  ins_tag = f"{{{self.namespaces['w']}}}ins"
                  del_tag = f"{{{self.namespaces['w']}}}del"
                  author_attr = f"{{{self.namespaces['w']}}}author"
          
                  # Remove w:ins elements
                  for parent in root.iter():
                      to_remove = []
                      for child in parent:
                          if child.tag == ins_tag and child.get(author_attr) == "Claude":
                              to_remove.append(child)
                      for elem in to_remove:
                          parent.remove(elem)
          
                  # Unwrap content in w:del elements where author is "Claude"
                  deltext_tag = f"{{{self.namespaces['w']}}}delText"
                  t_tag = f"{{{self.namespaces['w']}}}t"
          
                  for parent in root.iter():
                      to_process = []
                      for child in parent:
                          if child.tag == del_tag and child.get(author_attr) == "Claude":
                              to_process.append((child, list(parent).index(child)))
          
                      # Process in reverse order to maintain indices
                      for del_elem, del_index in reversed(to_process):
                          # Convert w:delText to w:t before moving
                          for elem in del_elem.iter():
                              if elem.tag == deltext_tag:
                                  elem.tag = t_tag
          
                          # Move all children of w:del to its parent before removing w:del
                          for child in reversed(list(del_elem)):
                              parent.insert(del_index, child)
                          parent.remove(del_elem)
          
              def _extract_text_content(self, root):
                  """Extract text content from Word XML, preserving paragraph structure.
          
                  Empty paragraphs are skipped to avoid false positives when tracked
                  insertions add only structural elements without text content.
                  """
                  p_tag = f"{{{self.namespaces['w']}}}p"
                  t_tag = f"{{{self.namespaces['w']}}}t"
          
                  paragraphs = []
                  for p_elem in root.findall(f".//{p_tag}"):
                      # Get all text elements within this paragraph
                      text_parts = []
                      for t_elem in p_elem.findall(f".//{t_tag}"):
                          if t_elem.text:
                              text_parts.append(t_elem.text)
                      paragraph_text = "".join(text_parts)
                      # Skip empty paragraphs - they don't affect content validation
                      if paragraph_text:
                          paragraphs.append(paragraph_text)
          
                  return "\n".join(paragraphs)
          
          
          if __name__ == "__main__":
              raise RuntimeError("This module should not be run directly.")
          
        • __init__.py 336 B
          """
          Validation modules for Word document processing.
          """
          
          from .base import BaseSchemaValidator
          from .docx import DOCXSchemaValidator
          from .pptx import PPTXSchemaValidator
          from .redlining import RedliningValidator
          
          __all__ = [
              "BaseSchemaValidator",
              "DOCXSchemaValidator",
              "PPTXSchemaValidator",
              "RedliningValidator",
          ]
          
      • pack.py 5.5 KB
        #!/usr/bin/env python3
        """
        Tool to pack a directory into a .docx, .pptx, or .xlsx file with XML formatting undone.
        
        Example usage:
            python pack.py <input_directory> <office_file> [--force]
        """
        
        import argparse
        import shutil
        import subprocess
        import sys
        import tempfile
        import defusedxml.minidom
        import zipfile
        from pathlib import Path
        
        
        def main():
            parser = argparse.ArgumentParser(description="Pack a directory into an Office file")
            parser.add_argument("input_directory", help="Unpacked Office document directory")
            parser.add_argument("output_file", help="Output Office file (.docx/.pptx/.xlsx)")
            parser.add_argument("--force", action="store_true", help="Skip validation")
            args = parser.parse_args()
        
            try:
                success = pack_document(
                    args.input_directory, args.output_file, validate=not args.force
                )
        
                # Show warning if validation was skipped
                if args.force:
                    print("Warning: Skipped validation, file may be corrupt", file=sys.stderr)
                # Exit with error if validation failed
                elif not success:
                    print("Contents would produce a corrupt file.", file=sys.stderr)
                    print("Please validate XML before repacking.", file=sys.stderr)
                    print("Use --force to skip validation and pack anyway.", file=sys.stderr)
                    sys.exit(1)
        
            except ValueError as e:
                sys.exit(f"Error: {e}")
        
        
        def pack_document(input_dir, output_file, validate=False):
            """Pack a directory into an Office file (.docx/.pptx/.xlsx).
        
            Args:
                input_dir: Path to unpacked Office document directory
                output_file: Path to output Office file
                validate: If True, validates with soffice (default: False)
        
            Returns:
                bool: True if successful, False if validation failed
            """
            input_dir = Path(input_dir)
            output_file = Path(output_file)
        
            if not input_dir.is_dir():
                raise ValueError(f"{input_dir} is not a directory")
            if output_file.suffix.lower() not in {".docx", ".pptx", ".xlsx"}:
                raise ValueError(f"{output_file} must be a .docx, .pptx, or .xlsx file")
        
            # Work in temporary directory to avoid modifying original
            with tempfile.TemporaryDirectory() as temp_dir:
                temp_content_dir = Path(temp_dir) / "content"
                shutil.copytree(input_dir, temp_content_dir)
        
                # Process XML files to remove pretty-printing whitespace
                for pattern in ["*.xml", "*.rels"]:
                    for xml_file in temp_content_dir.rglob(pattern):
                        condense_xml(xml_file)
        
                # Create final Office file as zip archive
                output_file.parent.mkdir(parents=True, exist_ok=True)
                with zipfile.ZipFile(output_file, "w", zipfile.ZIP_DEFLATED) as zf:
                    for f in temp_content_dir.rglob("*"):
                        if f.is_file():
                            zf.write(f, f.relative_to(temp_content_dir))
        
                # Validate if requested
                if validate:
                    if not validate_document(output_file):
                        output_file.unlink()  # Delete the corrupt file
                        return False
        
            return True
        
        
        def validate_document(doc_path):
            """Validate document by converting to HTML with soffice."""
            # Determine the correct filter based on file extension
            match doc_path.suffix.lower():
                case ".docx":
                    filter_name = "html:HTML"
                case ".pptx":
                    filter_name = "html:impress_html_Export"
                case ".xlsx":
                    filter_name = "html:HTML (StarCalc)"
        
            with tempfile.TemporaryDirectory() as temp_dir:
                try:
                    result = subprocess.run(
                        [
                            "soffice",
                            "--headless",
                            "--convert-to",
                            filter_name,
                            "--outdir",
                            temp_dir,
                            str(doc_path),
                        ],
                        capture_output=True,
                        timeout=10,
                        text=True,
                    )
                    if not (Path(temp_dir) / f"{doc_path.stem}.html").exists():
                        error_msg = result.stderr.strip() or "Document validation failed"
                        print(f"Validation error: {error_msg}", file=sys.stderr)
                        return False
                    return True
                except FileNotFoundError:
                    print("Warning: soffice not found. Skipping validation.", file=sys.stderr)
                    return True
                except subprocess.TimeoutExpired:
                    print("Validation error: Timeout during conversion", file=sys.stderr)
                    return False
                except Exception as e:
                    print(f"Validation error: {e}", file=sys.stderr)
                    return False
        
        
        def condense_xml(xml_file):
            """Strip unnecessary whitespace and remove comments."""
            with open(xml_file, "r", encoding="utf-8") as f:
                dom = defusedxml.minidom.parse(f)
        
            # Process each element to remove whitespace and comments
            for element in dom.getElementsByTagName("*"):
                # Skip w:t elements and their processing
                if element.tagName.endswith(":t"):
                    continue
        
                # Remove whitespace-only text nodes and comment nodes
                for child in list(element.childNodes):
                    if (
                        child.nodeType == child.TEXT_NODE
                        and child.nodeValue
                        and child.nodeValue.strip() == ""
                    ) or child.nodeType == child.COMMENT_NODE:
                        element.removeChild(child)
        
            # Write back the condensed XML
            with open(xml_file, "wb") as f:
                f.write(dom.toxml(encoding="UTF-8"))
        
        
        if __name__ == "__main__":
            main()
        
      • unpack.py 1 KB
        #!/usr/bin/env python3
        """Unpack and format XML contents of Office files (.docx, .pptx, .xlsx)"""
        
        import random
        import sys
        import defusedxml.minidom
        import zipfile
        from pathlib import Path
        
        # Get command line arguments
        assert len(sys.argv) == 3, "Usage: python unpack.py <office_file> <output_dir>"
        input_file, output_dir = sys.argv[1], sys.argv[2]
        
        # Extract and format
        output_path = Path(output_dir)
        output_path.mkdir(parents=True, exist_ok=True)
        zipfile.ZipFile(input_file).extractall(output_path)
        
        # Pretty print all XML files
        xml_files = list(output_path.rglob("*.xml")) + list(output_path.rglob("*.rels"))
        for xml_file in xml_files:
            content = xml_file.read_text(encoding="utf-8")
            dom = defusedxml.minidom.parseString(content)
            xml_file.write_bytes(dom.toprettyxml(indent="  ", encoding="ascii"))
        
        # For .docx files, suggest an RSID for tracked changes
        if input_file.endswith(".docx"):
            suggested_rsid = "".join(random.choices("0123456789ABCDEF", k=8))
            print(f"Suggested RSID for edit session: {suggested_rsid}")
        
      • validate.py 1.9 KB
        #!/usr/bin/env python3
        """
        Command line tool to validate Office document XML files against XSD schemas and tracked changes.
        
        Usage:
            python validate.py <dir> --original <original_file>
        """
        
        import argparse
        import sys
        from pathlib import Path
        
        from validation import DOCXSchemaValidator, PPTXSchemaValidator, RedliningValidator
        
        
        def main():
            parser = argparse.ArgumentParser(description="Validate Office document XML files")
            parser.add_argument(
                "unpacked_dir",
                help="Path to unpacked Office document directory",
            )
            parser.add_argument(
                "--original",
                required=True,
                help="Path to original file (.docx/.pptx/.xlsx)",
            )
            parser.add_argument(
                "-v",
                "--verbose",
                action="store_true",
                help="Enable verbose output",
            )
            args = parser.parse_args()
        
            # Validate paths
            unpacked_dir = Path(args.unpacked_dir)
            original_file = Path(args.original)
            file_extension = original_file.suffix.lower()
            assert unpacked_dir.is_dir(), f"Error: {unpacked_dir} is not a directory"
            assert original_file.is_file(), f"Error: {original_file} is not a file"
            assert file_extension in [".docx", ".pptx", ".xlsx"], (
                f"Error: {original_file} must be a .docx, .pptx, or .xlsx file"
            )
        
            # Run validations
            match file_extension:
                case ".docx":
                    validators = [DOCXSchemaValidator, RedliningValidator]
                case ".pptx":
                    validators = [PPTXSchemaValidator]
                case _:
                    print(f"Error: Validation not supported for file type {file_extension}")
                    sys.exit(1)
        
            # Run validators
            success = True
            for V in validators:
                validator = V(unpacked_dir, original_file, verbose=args.verbose)
                if not validator.validate():
                    success = False
        
            if success:
                print("All validations PASSED!")
        
            sys.exit(0 if success else 1)
        
        
        if __name__ == "__main__":
            main()
        
  • scripts
    • templates
      • comments.xml 2.6 KB · in bundle
      • commentsExtended.xml 2.6 KB · in bundle
      • commentsExtensible.xml 2.7 KB · in bundle
      • commentsIds.xml 2.6 KB · in bundle
      • people.xml 147 B · in bundle
    • document.py 49.2 KB
      #!/usr/bin/env python3
      """
      Library for working with Word documents: comments, tracked changes, and editing.
      
      Usage:
          from skills.docx.scripts.document import Document
      
          # Initialize
          doc = Document('workspace/unpacked')
          doc = Document('workspace/unpacked', author="John Doe", initials="JD")
      
          # Find nodes
          node = doc["word/document.xml"].get_node(tag="w:del", attrs={"w:id": "1"})
          node = doc["word/document.xml"].get_node(tag="w:p", line_number=10)
      
          # Add comments
          doc.add_comment(start=node, end=node, text="Comment text")
          doc.reply_to_comment(parent_comment_id=0, text="Reply text")
      
          # Suggest tracked changes
          doc["word/document.xml"].suggest_deletion(node)  # Delete content
          doc["word/document.xml"].revert_insertion(ins_node)  # Reject insertion
          doc["word/document.xml"].revert_deletion(del_node)  # Reject deletion
      
          # Save
          doc.save()
      """
      
      import html
      import random
      import shutil
      import tempfile
      from datetime import datetime, timezone
      from pathlib import Path
      
      from defusedxml import minidom
      from ooxml.scripts.pack import pack_document
      from ooxml.scripts.validation.docx import DOCXSchemaValidator
      from ooxml.scripts.validation.redlining import RedliningValidator
      
      from .utilities import XMLEditor
      
      # Path to template files
      TEMPLATE_DIR = Path(__file__).parent / "templates"
      
      
      class DocxXMLEditor(XMLEditor):
          """XMLEditor that automatically applies RSID, author, and date to new elements.
      
          Automatically adds attributes to elements that support them when inserting new content:
          - w:rsidR, w:rsidRDefault, w:rsidP (for w:p and w:r elements)
          - w:author and w:date (for w:ins, w:del, w:comment elements)
          - w:id (for w:ins and w:del elements)
      
          Attributes:
              dom (defusedxml.minidom.Document): The DOM document for direct manipulation
          """
      
          def __init__(
              self, xml_path, rsid: str, author: str = "Claude", initials: str = "C"
          ):
              """Initialize with required RSID and optional author.
      
              Args:
                  xml_path: Path to XML file to edit
                  rsid: RSID to automatically apply to new elements
                  author: Author name for tracked changes and comments (default: "Claude")
                  initials: Author initials (default: "C")
              """
              super().__init__(xml_path)
              self.rsid = rsid
              self.author = author
              self.initials = initials
      
          def _get_next_change_id(self):
              """Get the next available change ID by checking all tracked change elements."""
              max_id = -1
              for tag in ("w:ins", "w:del"):
                  elements = self.dom.getElementsByTagName(tag)
                  for elem in elements:
                      change_id = elem.getAttribute("w:id")
                      if change_id:
                          try:
                              max_id = max(max_id, int(change_id))
                          except ValueError:
                              pass
              return max_id + 1
      
          def _ensure_w16du_namespace(self):
              """Ensure w16du namespace is declared on the root element."""
              root = self.dom.documentElement
              if not root.hasAttribute("xmlns:w16du"):  # type: ignore
                  root.setAttribute(  # type: ignore
                      "xmlns:w16du",
                      "http://schemas.microsoft.com/office/word/2023/wordml/word16du",
                  )
      
          def _ensure_w16cex_namespace(self):
              """Ensure w16cex namespace is declared on the root element."""
              root = self.dom.documentElement
              if not root.hasAttribute("xmlns:w16cex"):  # type: ignore
                  root.setAttribute(  # type: ignore
                      "xmlns:w16cex",
                      "http://schemas.microsoft.com/office/word/2018/wordml/cex",
                  )
      
          def _ensure_w14_namespace(self):
              """Ensure w14 namespace is declared on the root element."""
              root = self.dom.documentElement
              if not root.hasAttribute("xmlns:w14"):  # type: ignore
                  root.setAttribute(  # type: ignore
                      "xmlns:w14",
                      "http://schemas.microsoft.com/office/word/2010/wordml",
                  )
      
          def _inject_attributes_to_nodes(self, nodes):
              """Inject RSID, author, and date attributes into DOM nodes where applicable.
      
              Adds attributes to elements that support them:
              - w:r: gets w:rsidR (or w:rsidDel if inside w:del)
              - w:p: gets w:rsidR, w:rsidRDefault, w:rsidP, w14:paraId, w14:textId
              - w:t: gets xml:space="preserve" if text has leading/trailing whitespace
              - w:ins, w:del: get w:id, w:author, w:date, w16du:dateUtc
              - w:comment: gets w:author, w:date, w:initials
              - w16cex:commentExtensible: gets w16cex:dateUtc
      
              Args:
                  nodes: List of DOM nodes to process
              """
              from datetime import datetime, timezone
      
              timestamp = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
      
              def is_inside_deletion(elem):
                  """Check if element is inside a w:del element."""
                  parent = elem.parentNode
                  while parent:
                      if parent.nodeType == parent.ELEMENT_NODE and parent.tagName == "w:del":
                          return True
                      parent = parent.parentNode
                  return False
      
              def add_rsid_to_p(elem):
                  if not elem.hasAttribute("w:rsidR"):
                      elem.setAttribute("w:rsidR", self.rsid)
                  if not elem.hasAttribute("w:rsidRDefault"):
                      elem.setAttribute("w:rsidRDefault", self.rsid)
                  if not elem.hasAttribute("w:rsidP"):
                      elem.setAttribute("w:rsidP", self.rsid)
                  # Add w14:paraId and w14:textId if not present
                  if not elem.hasAttribute("w14:paraId"):
                      self._ensure_w14_namespace()
                      elem.setAttribute("w14:paraId", _generate_hex_id())
                  if not elem.hasAttribute("w14:textId"):
                      self._ensure_w14_namespace()
                      elem.setAttribute("w14:textId", _generate_hex_id())
      
              def add_rsid_to_r(elem):
                  # Use w:rsidDel for <w:r> inside <w:del>, otherwise w:rsidR
                  if is_inside_deletion(elem):
                      if not elem.hasAttribute("w:rsidDel"):
                          elem.setAttribute("w:rsidDel", self.rsid)
                  else:
                      if not elem.hasAttribute("w:rsidR"):
                          elem.setAttribute("w:rsidR", self.rsid)
      
              def add_tracked_change_attrs(elem):
                  # Auto-assign w:id if not present
                  if not elem.hasAttribute("w:id"):
                      elem.setAttribute("w:id", str(self._get_next_change_id()))
                  if not elem.hasAttribute("w:author"):
                      elem.setAttribute("w:author", self.author)
                  if not elem.hasAttribute("w:date"):
                      elem.setAttribute("w:date", timestamp)
                  # Add w16du:dateUtc for tracked changes (same as w:date since we generate UTC timestamps)
                  if elem.tagName in ("w:ins", "w:del") and not elem.hasAttribute(
                      "w16du:dateUtc"
                  ):
                      self._ensure_w16du_namespace()
                      elem.setAttribute("w16du:dateUtc", timestamp)
      
              def add_comment_attrs(elem):
                  if not elem.hasAttribute("w:author"):
                      elem.setAttribute("w:author", self.author)
                  if not elem.hasAttribute("w:date"):
                      elem.setAttribute("w:date", timestamp)
                  if not elem.hasAttribute("w:initials"):
                      elem.setAttribute("w:initials", self.initials)
      
              def add_comment_extensible_date(elem):
                  # Add w16cex:dateUtc for comment extensible elements
                  if not elem.hasAttribute("w16cex:dateUtc"):
                      self._ensure_w16cex_namespace()
                      elem.setAttribute("w16cex:dateUtc", timestamp)
      
              def add_xml_space_to_t(elem):
                  # Add xml:space="preserve" to w:t if text has leading/trailing whitespace
                  if (
                      elem.firstChild
                      and elem.firstChild.nodeType == elem.firstChild.TEXT_NODE
                  ):
                      text = elem.firstChild.data
                      if text and (text[0].isspace() or text[-1].isspace()):
                          if not elem.hasAttribute("xml:space"):
                              elem.setAttribute("xml:space", "preserve")
      
              for node in nodes:
                  if node.nodeType != node.ELEMENT_NODE:
                      continue
      
                  # Handle the node itself
                  if node.tagName == "w:p":
                      add_rsid_to_p(node)
                  elif node.tagName == "w:r":
                      add_rsid_to_r(node)
                  elif node.tagName == "w:t":
                      add_xml_space_to_t(node)
                  elif node.tagName in ("w:ins", "w:del"):
                      add_tracked_change_attrs(node)
                  elif node.tagName == "w:comment":
                      add_comment_attrs(node)
                  elif node.tagName == "w16cex:commentExtensible":
                      add_comment_extensible_date(node)
      
                  # Process descendants (getElementsByTagName doesn't return the element itself)
                  for elem in node.getElementsByTagName("w:p"):
                      add_rsid_to_p(elem)
                  for elem in node.getElementsByTagName("w:r"):
                      add_rsid_to_r(elem)
                  for elem in node.getElementsByTagName("w:t"):
                      add_xml_space_to_t(elem)
                  for tag in ("w:ins", "w:del"):
                      for elem in node.getElementsByTagName(tag):
                          add_tracked_change_attrs(elem)
                  for elem in node.getElementsByTagName("w:comment"):
                      add_comment_attrs(elem)
                  for elem in node.getElementsByTagName("w16cex:commentExtensible"):
                      add_comment_extensible_date(elem)
      
          def replace_node(self, elem, new_content):
              """Replace node with automatic attribute injection."""
              nodes = super().replace_node(elem, new_content)
              self._inject_attributes_to_nodes(nodes)
              return nodes
      
          def insert_after(self, elem, xml_content):
              """Insert after with automatic attribute injection."""
              nodes = super().insert_after(elem, xml_content)
              self._inject_attributes_to_nodes(nodes)
              return nodes
      
          def insert_before(self, elem, xml_content):
              """Insert before with automatic attribute injection."""
              nodes = super().insert_before(elem, xml_content)
              self._inject_attributes_to_nodes(nodes)
              return nodes
      
          def append_to(self, elem, xml_content):
              """Append to with automatic attribute injection."""
              nodes = super().append_to(elem, xml_content)
              self._inject_attributes_to_nodes(nodes)
              return nodes
      
          def revert_insertion(self, elem):
              """Reject an insertion by wrapping its content in a deletion.
      
              Wraps all runs inside w:ins in w:del, converting w:t to w:delText.
              Can process a single w:ins element or a container element with multiple w:ins.
      
              Args:
                  elem: Element to process (w:ins, w:p, w:body, etc.)
      
              Returns:
                  list: List containing the processed element(s)
      
              Raises:
                  ValueError: If the element contains no w:ins elements
      
              Example:
                  # Reject a single insertion
                  ins = doc["word/document.xml"].get_node(tag="w:ins", attrs={"w:id": "5"})
                  doc["word/document.xml"].revert_insertion(ins)
      
                  # Reject all insertions in a paragraph
                  para = doc["word/document.xml"].get_node(tag="w:p", line_number=42)
                  doc["word/document.xml"].revert_insertion(para)
              """
              # Collect insertions
              ins_elements = []
              if elem.tagName == "w:ins":
                  ins_elements.append(elem)
              else:
                  ins_elements.extend(elem.getElementsByTagName("w:ins"))
      
              # Validate that there are insertions to reject
              if not ins_elements:
                  raise ValueError(
                      f"revert_insertion requires w:ins elements. "
                      f"The provided element <{elem.tagName}> contains no insertions. "
                  )
      
              # Process all insertions - wrap all children in w:del
              for ins_elem in ins_elements:
                  runs = list(ins_elem.getElementsByTagName("w:r"))
                  if not runs:
                      continue
      
                  # Create deletion wrapper
                  del_wrapper = self.dom.createElement("w:del")
      
                  # Process each run
                  for run in runs:
                      # Convert w:t → w:delText and w:rsidR → w:rsidDel
                      if run.hasAttribute("w:rsidR"):
                          run.setAttribute("w:rsidDel", run.getAttribute("w:rsidR"))
                          run.removeAttribute("w:rsidR")
                      elif not run.hasAttribute("w:rsidDel"):
                          run.setAttribute("w:rsidDel", self.rsid)
      
                      for t_elem in list(run.getElementsByTagName("w:t")):
                          del_text = self.dom.createElement("w:delText")
                          # Copy ALL child nodes (not just firstChild) to handle entities
                          while t_elem.firstChild:
                              del_text.appendChild(t_elem.firstChild)
                          for i in range(t_elem.attributes.length):
                              attr = t_elem.attributes.item(i)
                              del_text.setAttribute(attr.name, attr.value)
                          t_elem.parentNode.replaceChild(del_text, t_elem)
      
                  # Move all children from ins to del wrapper
                  while ins_elem.firstChild:
                      del_wrapper.appendChild(ins_elem.firstChild)
      
                  # Add del wrapper back to ins
                  ins_elem.appendChild(del_wrapper)
      
                  # Inject attributes to the deletion wrapper
                  self._inject_attributes_to_nodes([del_wrapper])
      
              return [elem]
      
          def revert_deletion(self, elem):
              """Reject a deletion by re-inserting the deleted content.
      
              Creates w:ins elements after each w:del, copying deleted content and
              converting w:delText back to w:t.
              Can process a single w:del element or a container element with multiple w:del.
      
              Args:
                  elem: Element to process (w:del, w:p, w:body, etc.)
      
              Returns:
                  list: If elem is w:del, returns [elem, new_ins]. Otherwise returns [elem].
      
              Raises:
                  ValueError: If the element contains no w:del elements
      
              Example:
                  # Reject a single deletion - returns [w:del, w:ins]
                  del_elem = doc["word/document.xml"].get_node(tag="w:del", attrs={"w:id": "3"})
                  nodes = doc["word/document.xml"].revert_deletion(del_elem)
      
                  # Reject all deletions in a paragraph - returns [para]
                  para = doc["word/document.xml"].get_node(tag="w:p", line_number=42)
                  nodes = doc["word/document.xml"].revert_deletion(para)
              """
              # Collect deletions FIRST - before we modify the DOM
              del_elements = []
              is_single_del = elem.tagName == "w:del"
      
              if is_single_del:
                  del_elements.append(elem)
              else:
                  del_elements.extend(elem.getElementsByTagName("w:del"))
      
              # Validate that there are deletions to reject
              if not del_elements:
                  raise ValueError(
                      f"revert_deletion requires w:del elements. "
                      f"The provided element <{elem.tagName}> contains no deletions. "
                  )
      
              # Track created insertion (only relevant if elem is a single w:del)
              created_insertion = None
      
              # Process all deletions - create insertions that copy the deleted content
              for del_elem in del_elements:
                  # Clone the deleted runs and convert them to insertions
                  runs = list(del_elem.getElementsByTagName("w:r"))
                  if not runs:
                      continue
      
                  # Create insertion wrapper
                  ins_elem = self.dom.createElement("w:ins")
      
                  for run in runs:
                      # Clone the run
                      new_run = run.cloneNode(True)
      
                      # Convert w:delText → w:t
                      for del_text in list(new_run.getElementsByTagName("w:delText")):
                          t_elem = self.dom.createElement("w:t")
                          # Copy ALL child nodes (not just firstChild) to handle entities
                          while del_text.firstChild:
                              t_elem.appendChild(del_text.firstChild)
                          for i in range(del_text.attributes.length):
                              attr = del_text.attributes.item(i)
                              t_elem.setAttribute(attr.name, attr.value)
                          del_text.parentNode.replaceChild(t_elem, del_text)
      
                      # Update run attributes: w:rsidDel → w:rsidR
                      if new_run.hasAttribute("w:rsidDel"):
                          new_run.setAttribute("w:rsidR", new_run.getAttribute("w:rsidDel"))
                          new_run.removeAttribute("w:rsidDel")
                      elif not new_run.hasAttribute("w:rsidR"):
                          new_run.setAttribute("w:rsidR", self.rsid)
      
                      ins_elem.appendChild(new_run)
      
                  # Insert the new insertion after the deletion
                  nodes = self.insert_after(del_elem, ins_elem.toxml())
      
                  # If processing a single w:del, track the created insertion
                  if is_single_del and nodes:
                      created_insertion = nodes[0]
      
              # Return based on input type
              if is_single_del and created_insertion:
                  return [elem, created_insertion]
              else:
                  return [elem]
      
          @staticmethod
          def suggest_paragraph(xml_content: str) -> str:
              """Transform paragraph XML to add tracked change wrapping for insertion.
      
              Wraps runs in <w:ins> and adds <w:ins/> to w:rPr in w:pPr for numbered lists.
      
              Args:
                  xml_content: XML string containing a <w:p> element
      
              Returns:
                  str: Transformed XML with tracked change wrapping
              """
              wrapper = f'<root xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">{xml_content}</root>'
              doc = minidom.parseString(wrapper)
              para = doc.getElementsByTagName("w:p")[0]
      
              # Ensure w:pPr exists
              pPr_list = para.getElementsByTagName("w:pPr")
              if not pPr_list:
                  pPr = doc.createElement("w:pPr")
                  para.insertBefore(
                      pPr, para.firstChild
                  ) if para.firstChild else para.appendChild(pPr)
              else:
                  pPr = pPr_list[0]
      
              # Ensure w:rPr exists in w:pPr
              rPr_list = pPr.getElementsByTagName("w:rPr")
              if not rPr_list:
                  rPr = doc.createElement("w:rPr")
                  pPr.appendChild(rPr)
              else:
                  rPr = rPr_list[0]
      
              # Add <w:ins/> to w:rPr
              ins_marker = doc.createElement("w:ins")
              rPr.insertBefore(
                  ins_marker, rPr.firstChild
              ) if rPr.firstChild else rPr.appendChild(ins_marker)
      
              # Wrap all non-pPr children in <w:ins>
              ins_wrapper = doc.createElement("w:ins")
              for child in [c for c in para.childNodes if c.nodeName != "w:pPr"]:
                  para.removeChild(child)
                  ins_wrapper.appendChild(child)
              para.appendChild(ins_wrapper)
      
              return para.toxml()
      
          def suggest_deletion(self, elem):
              """Mark a w:r or w:p element as deleted with tracked changes (in-place DOM manipulation).
      
              For w:r: wraps in <w:del>, converts <w:t> to <w:delText>, preserves w:rPr
              For w:p (regular): wraps content in <w:del>, converts <w:t> to <w:delText>
              For w:p (numbered list): adds <w:del/> to w:rPr in w:pPr, wraps content in <w:del>
      
              Args:
                  elem: A w:r or w:p DOM element without existing tracked changes
      
              Returns:
                  Element: The modified element
      
              Raises:
                  ValueError: If element has existing tracked changes or invalid structure
              """
              if elem.nodeName == "w:r":
                  # Check for existing w:delText
                  if elem.getElementsByTagName("w:delText"):
                      raise ValueError("w:r element already contains w:delText")
      
                  # Convert w:t → w:delText
                  for t_elem in list(elem.getElementsByTagName("w:t")):
                      del_text = self.dom.createElement("w:delText")
                      # Copy ALL child nodes (not just firstChild) to handle entities
                      while t_elem.firstChild:
                          del_text.appendChild(t_elem.firstChild)
                      # Preserve attributes like xml:space
                      for i in range(t_elem.attributes.length):
                          attr = t_elem.attributes.item(i)
                          del_text.setAttribute(attr.name, attr.value)
                      t_elem.parentNode.replaceChild(del_text, t_elem)
      
                  # Update run attributes: w:rsidR → w:rsidDel
                  if elem.hasAttribute("w:rsidR"):
                      elem.setAttribute("w:rsidDel", elem.getAttribute("w:rsidR"))
                      elem.removeAttribute("w:rsidR")
                  elif not elem.hasAttribute("w:rsidDel"):
                      elem.setAttribute("w:rsidDel", self.rsid)
      
                  # Wrap in w:del
                  del_wrapper = self.dom.createElement("w:del")
                  parent = elem.parentNode
                  parent.insertBefore(del_wrapper, elem)
                  parent.removeChild(elem)
                  del_wrapper.appendChild(elem)
      
                  # Inject attributes to the deletion wrapper
                  self._inject_attributes_to_nodes([del_wrapper])
      
                  return del_wrapper
      
              elif elem.nodeName == "w:p":
                  # Check for existing tracked changes
                  if elem.getElementsByTagName("w:ins") or elem.getElementsByTagName("w:del"):
                      raise ValueError("w:p element already contains tracked changes")
      
                  # Check if it's a numbered list item
                  pPr_list = elem.getElementsByTagName("w:pPr")
                  is_numbered = pPr_list and pPr_list[0].getElementsByTagName("w:numPr")
      
                  if is_numbered:
                      # Add <w:del/> to w:rPr in w:pPr
                      pPr = pPr_list[0]
                      rPr_list = pPr.getElementsByTagName("w:rPr")
      
                      if not rPr_list:
                          rPr = self.dom.createElement("w:rPr")
                          pPr.appendChild(rPr)
                      else:
                          rPr = rPr_list[0]
      
                      # Add <w:del/> marker
                      del_marker = self.dom.createElement("w:del")
                      rPr.insertBefore(
                          del_marker, rPr.firstChild
                      ) if rPr.firstChild else rPr.appendChild(del_marker)
      
                  # Convert w:t → w:delText in all runs
                  for t_elem in list(elem.getElementsByTagName("w:t")):
                      del_text = self.dom.createElement("w:delText")
                      # Copy ALL child nodes (not just firstChild) to handle entities
                      while t_elem.firstChild:
                          del_text.appendChild(t_elem.firstChild)
                      # Preserve attributes like xml:space
                      for i in range(t_elem.attributes.length):
                          attr = t_elem.attributes.item(i)
                          del_text.setAttribute(attr.name, attr.value)
                      t_elem.parentNode.replaceChild(del_text, t_elem)
      
                  # Update run attributes: w:rsidR → w:rsidDel
                  for run in elem.getElementsByTagName("w:r"):
                      if run.hasAttribute("w:rsidR"):
                          run.setAttribute("w:rsidDel", run.getAttribute("w:rsidR"))
                          run.removeAttribute("w:rsidR")
                      elif not run.hasAttribute("w:rsidDel"):
                          run.setAttribute("w:rsidDel", self.rsid)
      
                  # Wrap all non-pPr children in <w:del>
                  del_wrapper = self.dom.createElement("w:del")
                  for child in [c for c in elem.childNodes if c.nodeName != "w:pPr"]:
                      elem.removeChild(child)
                      del_wrapper.appendChild(child)
                  elem.appendChild(del_wrapper)
      
                  # Inject attributes to the deletion wrapper
                  self._inject_attributes_to_nodes([del_wrapper])
      
                  return elem
      
              else:
                  raise ValueError(f"Element must be w:r or w:p, got {elem.nodeName}")
      
      
      def _generate_hex_id() -> str:
          """Generate random 8-character hex ID for para/durable IDs.
      
          Values are constrained to be less than 0x7FFFFFFF per OOXML spec:
          - paraId must be < 0x80000000
          - durableId must be < 0x7FFFFFFF
          We use the stricter constraint (0x7FFFFFFF) for both.
          """
          return f"{random.randint(1, 0x7FFFFFFE):08X}"
      
      
      def _generate_rsid() -> str:
          """Generate random 8-character hex RSID."""
          return "".join(random.choices("0123456789ABCDEF", k=8))
      
      
      class Document:
          """Manages comments in unpacked Word documents."""
      
          def __init__(
              self,
              unpacked_dir,
              rsid=None,
              track_revisions=False,
              author="Claude",
              initials="C",
          ):
              """
              Initialize with path to unpacked Word document directory.
              Automatically sets up comment infrastructure (people.xml, RSIDs).
      
              Args:
                  unpacked_dir: Path to unpacked DOCX directory (must contain word/ subdirectory)
                  rsid: Optional RSID to use for all comment elements. If not provided, one will be generated.
                  track_revisions: If True, enables track revisions in settings.xml (default: False)
                  author: Default author name for comments (default: "Claude")
                  initials: Default author initials for comments (default: "C")
              """
              self.original_path = Path(unpacked_dir)
      
              if not self.original_path.exists() or not self.original_path.is_dir():
                  raise ValueError(f"Directory not found: {unpacked_dir}")
      
              # Create temporary directory with subdirectories for unpacked content and baseline
              self.temp_dir = tempfile.mkdtemp(prefix="docx_")
              self.unpacked_path = Path(self.temp_dir) / "unpacked"
              shutil.copytree(self.original_path, self.unpacked_path)
      
              # Pack original directory into temporary .docx for validation baseline (outside unpacked dir)
              self.original_docx = Path(self.temp_dir) / "original.docx"
              pack_document(self.original_path, self.original_docx, validate=False)
      
              self.word_path = self.unpacked_path / "word"
      
              # Generate RSID if not provided
              self.rsid = rsid if rsid else _generate_rsid()
              print(f"Using RSID: {self.rsid}")
      
              # Set default author and initials
              self.author = author
              self.initials = initials
      
              # Cache for lazy-loaded editors
              self._editors = {}
      
              # Comment file paths
              self.comments_path = self.word_path / "comments.xml"
              self.comments_extended_path = self.word_path / "commentsExtended.xml"
              self.comments_ids_path = self.word_path / "commentsIds.xml"
              self.comments_extensible_path = self.word_path / "commentsExtensible.xml"
      
              # Load existing comments and determine next ID (before setup modifies files)
              self.existing_comments = self._load_existing_comments()
              self.next_comment_id = self._get_next_comment_id()
      
              # Convenient access to document.xml editor (semi-private)
              self._document = self["word/document.xml"]
      
              # Setup tracked changes infrastructure
              self._setup_tracking(track_revisions=track_revisions)
      
              # Add author to people.xml
              self._add_author_to_people(author)
      
          def __getitem__(self, xml_path: str) -> DocxXMLEditor:
              """
              Get or create a DocxXMLEditor for the specified XML file.
      
              Enables lazy-loaded editors with bracket notation:
                  node = doc["word/document.xml"].get_node(tag="w:p", line_number=42)
      
              Args:
                  xml_path: Relative path to XML file (e.g., "word/document.xml", "word/comments.xml")
      
              Returns:
                  DocxXMLEditor instance for the specified file
      
              Raises:
                  ValueError: If the file does not exist
      
              Example:
                  # Get node from document.xml
                  node = doc["word/document.xml"].get_node(tag="w:del", attrs={"w:id": "1"})
      
                  # Get node from comments.xml
                  comment = doc["word/comments.xml"].get_node(tag="w:comment", attrs={"w:id": "0"})
              """
              if xml_path not in self._editors:
                  file_path = self.unpacked_path / xml_path
                  if not file_path.exists():
                      raise ValueError(f"XML file not found: {xml_path}")
                  # Use DocxXMLEditor with RSID, author, and initials for all editors
                  self._editors[xml_path] = DocxXMLEditor(
                      file_path, rsid=self.rsid, author=self.author, initials=self.initials
                  )
              return self._editors[xml_path]
      
          def add_comment(self, start, end, text: str) -> int:
              """
              Add a comment spanning from one element to another.
      
              Args:
                  start: DOM element for the starting point
                  end: DOM element for the ending point
                  text: Comment content
      
              Returns:
                  The comment ID that was created
      
              Example:
                  start_node = cm.get_document_node(tag="w:del", id="1")
                  end_node = cm.get_document_node(tag="w:ins", id="2")
                  cm.add_comment(start=start_node, end=end_node, text="Explanation")
              """
              comment_id = self.next_comment_id
              para_id = _generate_hex_id()
              durable_id = _generate_hex_id()
              timestamp = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
      
              # Add comment ranges to document.xml immediately
              self._document.insert_before(start, self._comment_range_start_xml(comment_id))
      
              # If end node is a paragraph, append comment markup inside it
              # Otherwise insert after it (for run-level anchors)
              if end.tagName == "w:p":
                  self._document.append_to(end, self._comment_range_end_xml(comment_id))
              else:
                  self._document.insert_after(end, self._comment_range_end_xml(comment_id))
      
              # Add to comments.xml immediately
              self._add_to_comments_xml(
                  comment_id, para_id, text, self.author, self.initials, timestamp
              )
      
              # Add to commentsExtended.xml immediately
              self._add_to_comments_extended_xml(para_id, parent_para_id=None)
      
              # Add to commentsIds.xml immediately
              self._add_to_comments_ids_xml(para_id, durable_id)
      
              # Add to commentsExtensible.xml immediately
              self._add_to_comments_extensible_xml(durable_id)
      
              # Update existing_comments so replies work
              self.existing_comments[comment_id] = {"para_id": para_id}
      
              self.next_comment_id += 1
              return comment_id
      
          def reply_to_comment(
              self,
              parent_comment_id: int,
              text: str,
          ) -> int:
              """
              Add a reply to an existing comment.
      
              Args:
                  parent_comment_id: The w:id of the parent comment to reply to
                  text: Reply text
      
              Returns:
                  The comment ID that was created for the reply
      
              Example:
                  cm.reply_to_comment(parent_comment_id=0, text="I agree with this change")
              """
              if parent_comment_id not in self.existing_comments:
                  raise ValueError(f"Parent comment with id={parent_comment_id} not found")
      
              parent_info = self.existing_comments[parent_comment_id]
              comment_id = self.next_comment_id
              para_id = _generate_hex_id()
              durable_id = _generate_hex_id()
              timestamp = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
      
              # Add comment ranges to document.xml immediately
              parent_start_elem = self._document.get_node(
                  tag="w:commentRangeStart", attrs={"w:id": str(parent_comment_id)}
              )
              parent_ref_elem = self._document.get_node(
                  tag="w:commentReference", attrs={"w:id": str(parent_comment_id)}
              )
      
              self._document.insert_after(
                  parent_start_elem, self._comment_range_start_xml(comment_id)
              )
              parent_ref_run = parent_ref_elem.parentNode
              self._document.insert_after(
                  parent_ref_run, f'<w:commentRangeEnd w:id="{comment_id}"/>'
              )
              self._document.insert_after(
                  parent_ref_run, self._comment_ref_run_xml(comment_id)
              )
      
              # Add to comments.xml immediately
              self._add_to_comments_xml(
                  comment_id, para_id, text, self.author, self.initials, timestamp
              )
      
              # Add to commentsExtended.xml immediately (with parent)
              self._add_to_comments_extended_xml(
                  para_id, parent_para_id=parent_info["para_id"]
              )
      
              # Add to commentsIds.xml immediately
              self._add_to_comments_ids_xml(para_id, durable_id)
      
              # Add to commentsExtensible.xml immediately
              self._add_to_comments_extensible_xml(durable_id)
      
              # Update existing_comments so replies work
              self.existing_comments[comment_id] = {"para_id": para_id}
      
              self.next_comment_id += 1
              return comment_id
      
          def __del__(self):
              """Clean up temporary directory on deletion."""
              if hasattr(self, "temp_dir") and Path(self.temp_dir).exists():
                  shutil.rmtree(self.temp_dir)
      
          def validate(self) -> None:
              """
              Validate the document against XSD schema and redlining rules.
      
              Raises:
                  ValueError: If validation fails.
              """
              # Create validators with current state
              schema_validator = DOCXSchemaValidator(
                  self.unpacked_path, self.original_docx, verbose=False
              )
              redlining_validator = RedliningValidator(
                  self.unpacked_path, self.original_docx, verbose=False
              )
      
              # Run validations
              if not schema_validator.validate():
                  raise ValueError("Schema validation failed")
              if not redlining_validator.validate():
                  raise ValueError("Redlining validation failed")
      
          def save(self, destination=None, validate=True) -> None:
              """
              Save all modified XML files to disk and copy to destination directory.
      
              This persists all changes made via add_comment() and reply_to_comment().
      
              Args:
                  destination: Optional path to save to. If None, saves back to original directory.
                  validate: If True, validates document before saving (default: True).
              """
              # Only ensure comment relationships and content types if comment files exist
              if self.comments_path.exists():
                  self._ensure_comment_relationships()
                  self._ensure_comment_content_types()
      
              # Save all modified XML files in temp directory
              for editor in self._editors.values():
                  editor.save()
      
              # Validate by default
              if validate:
                  self.validate()
      
              # Copy contents from temp directory to destination (or original directory)
              target_path = Path(destination) if destination else self.original_path
              shutil.copytree(self.unpacked_path, target_path, dirs_exist_ok=True)
      
          # ==================== Private: Initialization ====================
      
          def _get_next_comment_id(self):
              """Get the next available comment ID."""
              if not self.comments_path.exists():
                  return 0
      
              editor = self["word/comments.xml"]
              max_id = -1
              for comment_elem in editor.dom.getElementsByTagName("w:comment"):
                  comment_id = comment_elem.getAttribute("w:id")
                  if comment_id:
                      try:
                          max_id = max(max_id, int(comment_id))
                      except ValueError:
                          pass
              return max_id + 1
      
          def _load_existing_comments(self):
              """Load existing comments from files to enable replies."""
              if not self.comments_path.exists():
                  return {}
      
              editor = self["word/comments.xml"]
              existing = {}
      
              for comment_elem in editor.dom.getElementsByTagName("w:comment"):
                  comment_id = comment_elem.getAttribute("w:id")
                  if not comment_id:
                      continue
      
                  # Find para_id from the w:p element within the comment
                  para_id = None
                  for p_elem in comment_elem.getElementsByTagName("w:p"):
                      para_id = p_elem.getAttribute("w14:paraId")
                      if para_id:
                          break
      
                  if not para_id:
                      continue
      
                  existing[int(comment_id)] = {"para_id": para_id}
      
              return existing
      
          # ==================== Private: Setup Methods ====================
      
          def _setup_tracking(self, track_revisions=False):
              """Set up comment infrastructure in unpacked directory.
      
              Args:
                  track_revisions: If True, enables track revisions in settings.xml
              """
              # Create or update word/people.xml
              people_file = self.word_path / "people.xml"
              self._update_people_xml(people_file)
      
              # Update XML files
              self._add_content_type_for_people(self.unpacked_path / "[Content_Types].xml")
              self._add_relationship_for_people(
                  self.word_path / "_rels" / "document.xml.rels"
              )
      
              # Always add RSID to settings.xml, optionally enable trackRevisions
              self._update_settings(
                  self.word_path / "settings.xml", track_revisions=track_revisions
              )
      
          def _update_people_xml(self, path):
              """Create people.xml if it doesn't exist."""
              if not path.exists():
                  # Copy from template
                  shutil.copy(TEMPLATE_DIR / "people.xml", path)
      
          def _add_content_type_for_people(self, path):
              """Add people.xml content type to [Content_Types].xml if not already present."""
              editor = self["[Content_Types].xml"]
      
              if self._has_override(editor, "/word/people.xml"):
                  return
      
              # Add Override element
              root = editor.dom.documentElement
              override_xml = '<Override PartName="/word/people.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.people+xml"/>'
              editor.append_to(root, override_xml)
      
          def _add_relationship_for_people(self, path):
              """Add people.xml relationship to document.xml.rels if not already present."""
              editor = self["word/_rels/document.xml.rels"]
      
              if self._has_relationship(editor, "people.xml"):
                  return
      
              root = editor.dom.documentElement
              root_tag = root.tagName  # type: ignore
              prefix = root_tag.split(":")[0] + ":" if ":" in root_tag else ""
              next_rid = editor.get_next_rid()
      
              # Create the relationship entry
              rel_xml = f'<{prefix}Relationship Id="{next_rid}" Type="http://schemas.microsoft.com/office/2011/relationships/people" Target="people.xml"/>'
              editor.append_to(root, rel_xml)
      
          def _update_settings(self, path, track_revisions=False):
              """Add RSID and optionally enable track revisions in settings.xml.
      
              Args:
                  path: Path to settings.xml
                  track_revisions: If True, adds trackRevisions element
      
              Places elements per OOXML schema order:
              - trackRevisions: early (before defaultTabStop)
              - rsids: late (after compat)
              """
              editor = self["word/settings.xml"]
              root = editor.get_node(tag="w:settings")
              prefix = root.tagName.split(":")[0] if ":" in root.tagName else "w"
      
              # Conditionally add trackRevisions if requested
              if track_revisions:
                  track_revisions_exists = any(
                      elem.tagName == f"{prefix}:trackRevisions"
                      for elem in editor.dom.getElementsByTagName(f"{prefix}:trackRevisions")
                  )
      
                  if not track_revisions_exists:
                      track_rev_xml = f"<{prefix}:trackRevisions/>"
                      # Try to insert before documentProtection, defaultTabStop, or at start
                      inserted = False
                      for tag in [f"{prefix}:documentProtection", f"{prefix}:defaultTabStop"]:
                          elements = editor.dom.getElementsByTagName(tag)
                          if elements:
                              editor.insert_before(elements[0], track_rev_xml)
                              inserted = True
                              break
                      if not inserted:
                          # Insert as first child of settings
                          if root.firstChild:
                              editor.insert_before(root.firstChild, track_rev_xml)
                          else:
                              editor.append_to(root, track_rev_xml)
      
              # Always check if rsids section exists
              rsids_elements = editor.dom.getElementsByTagName(f"{prefix}:rsids")
      
              if not rsids_elements:
                  # Add new rsids section
                  rsids_xml = f'''<{prefix}:rsids>
        <{prefix}:rsidRoot {prefix}:val="{self.rsid}"/>
        <{prefix}:rsid {prefix}:val="{self.rsid}"/>
      </{prefix}:rsids>'''
      
                  # Try to insert after compat, before clrSchemeMapping, or before closing tag
                  inserted = False
                  compat_elements = editor.dom.getElementsByTagName(f"{prefix}:compat")
                  if compat_elements:
                      editor.insert_after(compat_elements[0], rsids_xml)
                      inserted = True
      
                  if not inserted:
                      clr_elements = editor.dom.getElementsByTagName(
                          f"{prefix}:clrSchemeMapping"
                      )
                      if clr_elements:
                          editor.insert_before(clr_elements[0], rsids_xml)
                          inserted = True
      
                  if not inserted:
                      editor.append_to(root, rsids_xml)
              else:
                  # Check if this rsid already exists
                  rsids_elem = rsids_elements[0]
                  rsid_exists = any(
                      elem.getAttribute(f"{prefix}:val") == self.rsid
                      for elem in rsids_elem.getElementsByTagName(f"{prefix}:rsid")
                  )
      
                  if not rsid_exists:
                      rsid_xml = f'<{prefix}:rsid {prefix}:val="{self.rsid}"/>'
                      editor.append_to(rsids_elem, rsid_xml)
      
          # ==================== Private: XML File Creation ====================
      
          def _add_to_comments_xml(
              self, comment_id, para_id, text, author, initials, timestamp
          ):
              """Add a single comment to comments.xml."""
              if not self.comments_path.exists():
                  shutil.copy(TEMPLATE_DIR / "comments.xml", self.comments_path)
      
              editor = self["word/comments.xml"]
              root = editor.get_node(tag="w:comments")
      
              escaped_text = (
                  text.replace("&", "&amp;").replace("<", "&lt;").replace(">", "&gt;")
              )
              # Note: w:rsidR, w:rsidRDefault, w:rsidP on w:p, w:rsidR on w:r,
              # and w:author, w:date, w:initials on w:comment are automatically added by DocxXMLEditor
              comment_xml = f'''<w:comment w:id="{comment_id}">
        <w:p w14:paraId="{para_id}" w14:textId="77777777">
          <w:r><w:rPr><w:rStyle w:val="CommentReference"/></w:rPr><w:annotationRef/></w:r>
          <w:r><w:rPr><w:color w:val="000000"/><w:sz w:val="20"/><w:szCs w:val="20"/></w:rPr><w:t>{escaped_text}</w:t></w:r>
        </w:p>
      </w:comment>'''
              editor.append_to(root, comment_xml)
      
          def _add_to_comments_extended_xml(self, para_id, parent_para_id):
              """Add a single comment to commentsExtended.xml."""
              if not self.comments_extended_path.exists():
                  shutil.copy(
                      TEMPLATE_DIR / "commentsExtended.xml", self.comments_extended_path
                  )
      
              editor = self["word/commentsExtended.xml"]
              root = editor.get_node(tag="w15:commentsEx")
      
              if parent_para_id:
                  xml = f'<w15:commentEx w15:paraId="{para_id}" w15:paraIdParent="{parent_para_id}" w15:done="0"/>'
              else:
                  xml = f'<w15:commentEx w15:paraId="{para_id}" w15:done="0"/>'
              editor.append_to(root, xml)
      
          def _add_to_comments_ids_xml(self, para_id, durable_id):
              """Add a single comment to commentsIds.xml."""
              if not self.comments_ids_path.exists():
                  shutil.copy(TEMPLATE_DIR / "commentsIds.xml", self.comments_ids_path)
      
              editor = self["word/commentsIds.xml"]
              root = editor.get_node(tag="w16cid:commentsIds")
      
              xml = f'<w16cid:commentId w16cid:paraId="{para_id}" w16cid:durableId="{durable_id}"/>'
              editor.append_to(root, xml)
      
          def _add_to_comments_extensible_xml(self, durable_id):
              """Add a single comment to commentsExtensible.xml."""
              if not self.comments_extensible_path.exists():
                  shutil.copy(
                      TEMPLATE_DIR / "commentsExtensible.xml", self.comments_extensible_path
                  )
      
              editor = self["word/commentsExtensible.xml"]
              root = editor.get_node(tag="w16cex:commentsExtensible")
      
              xml = f'<w16cex:commentExtensible w16cex:durableId="{durable_id}"/>'
              editor.append_to(root, xml)
      
          # ==================== Private: XML Fragments ====================
      
          def _comment_range_start_xml(self, comment_id):
              """Generate XML for comment range start."""
              return f'<w:commentRangeStart w:id="{comment_id}"/>'
      
          def _comment_range_end_xml(self, comment_id):
              """Generate XML for comment range end with reference run.
      
              Note: w:rsidR is automatically added by DocxXMLEditor.
              """
              return f'''<w:commentRangeEnd w:id="{comment_id}"/>
      <w:r>
        <w:rPr><w:rStyle w:val="CommentReference"/></w:rPr>
        <w:commentReference w:id="{comment_id}"/>
      </w:r>'''
      
          def _comment_ref_run_xml(self, comment_id):
              """Generate XML for comment reference run.
      
              Note: w:rsidR is automatically added by DocxXMLEditor.
              """
              return f'''<w:r>
        <w:rPr><w:rStyle w:val="CommentReference"/></w:rPr>
        <w:commentReference w:id="{comment_id}"/>
      </w:r>'''
      
          # ==================== Private: Metadata Updates ====================
      
          def _has_relationship(self, editor, target):
              """Check if a relationship with given target exists."""
              for rel_elem in editor.dom.getElementsByTagName("Relationship"):
                  if rel_elem.getAttribute("Target") == target:
                      return True
              return False
      
          def _has_override(self, editor, part_name):
              """Check if an override with given part name exists."""
              for override_elem in editor.dom.getElementsByTagName("Override"):
                  if override_elem.getAttribute("PartName") == part_name:
                      return True
              return False
      
          def _has_author(self, editor, author):
              """Check if an author already exists in people.xml."""
              for person_elem in editor.dom.getElementsByTagName("w15:person"):
                  if person_elem.getAttribute("w15:author") == author:
                      return True
              return False
      
          def _add_author_to_people(self, author):
              """Add author to people.xml (called during initialization)."""
              people_path = self.word_path / "people.xml"
      
              # people.xml should already exist from _setup_tracking
              if not people_path.exists():
                  raise ValueError("people.xml should exist after _setup_tracking")
      
              editor = self["word/people.xml"]
              root = editor.get_node(tag="w15:people")
      
              # Check if author already exists
              if self._has_author(editor, author):
                  return
      
              # Add author with proper XML escaping to prevent injection
              escaped_author = html.escape(author, quote=True)
              person_xml = f'''<w15:person w15:author="{escaped_author}">
        <w15:presenceInfo w15:providerId="None" w15:userId="{escaped_author}"/>
      </w15:person>'''
              editor.append_to(root, person_xml)
      
          def _ensure_comment_relationships(self):
              """Ensure word/_rels/document.xml.rels has comment relationships."""
              editor = self["word/_rels/document.xml.rels"]
      
              if self._has_relationship(editor, "comments.xml"):
                  return
      
              root = editor.dom.documentElement
              root_tag = root.tagName  # type: ignore
              prefix = root_tag.split(":")[0] + ":" if ":" in root_tag else ""
              next_rid_num = int(editor.get_next_rid()[3:])
      
              # Add relationship elements
              rels = [
                  (
                      next_rid_num,
                      "http://schemas.openxmlformats.org/officeDocument/2006/relationships/comments",
                      "comments.xml",
                  ),
                  (
                      next_rid_num + 1,
                      "http://schemas.microsoft.com/office/2011/relationships/commentsExtended",
                      "commentsExtended.xml",
                  ),
                  (
                      next_rid_num + 2,
                      "http://schemas.microsoft.com/office/2016/09/relationships/commentsIds",
                      "commentsIds.xml",
                  ),
                  (
                      next_rid_num + 3,
                      "http://schemas.microsoft.com/office/2018/08/relationships/commentsExtensible",
                      "commentsExtensible.xml",
                  ),
              ]
      
              for rel_id, rel_type, target in rels:
                  rel_xml = f'<{prefix}Relationship Id="rId{rel_id}" Type="{rel_type}" Target="{target}"/>'
                  editor.append_to(root, rel_xml)
      
          def _ensure_comment_content_types(self):
              """Ensure [Content_Types].xml has comment content types."""
              editor = self["[Content_Types].xml"]
      
              if self._has_override(editor, "/word/comments.xml"):
                  return
      
              root = editor.dom.documentElement
      
              # Add Override elements
              overrides = [
                  (
                      "/word/comments.xml",
                      "application/vnd.openxmlformats-officedocument.wordprocessingml.comments+xml",
                  ),
                  (
                      "/word/commentsExtended.xml",
                      "application/vnd.openxmlformats-officedocument.wordprocessingml.commentsExtended+xml",
                  ),
                  (
                      "/word/commentsIds.xml",
                      "application/vnd.openxmlformats-officedocument.wordprocessingml.commentsIds+xml",
                  ),
                  (
                      "/word/commentsExtensible.xml",
                      "application/vnd.openxmlformats-officedocument.wordprocessingml.commentsExtensible+xml",
                  ),
              ]
      
              for part_name, content_type in overrides:
                  override_xml = (
                      f'<Override PartName="{part_name}" ContentType="{content_type}"/>'
                  )
                  editor.append_to(root, override_xml)
      
    • utilities.py 13.4 KB
      #!/usr/bin/env python3
      """
      Utilities for editing OOXML documents.
      
      This module provides XMLEditor, a tool for manipulating XML files with support for
      line-number-based node finding and DOM manipulation. Each element is automatically
      annotated with its original line and column position during parsing.
      
      Example usage:
          editor = XMLEditor("document.xml")
      
          # Find node by line number or range
          elem = editor.get_node(tag="w:r", line_number=519)
          elem = editor.get_node(tag="w:p", line_number=range(100, 200))
      
          # Find node by text content
          elem = editor.get_node(tag="w:p", contains="specific text")
      
          # Find node by attributes
          elem = editor.get_node(tag="w:r", attrs={"w:id": "target"})
      
          # Combine filters
          elem = editor.get_node(tag="w:p", line_number=range(1, 50), contains="text")
      
          # Replace, insert, or manipulate
          new_elem = editor.replace_node(elem, "<w:r><w:t>new text</w:t></w:r>")
          editor.insert_after(new_elem, "<w:r><w:t>more</w:t></w:r>")
      
          # Save changes
          editor.save()
      """
      
      import html
      from pathlib import Path
      from typing import Optional, Union
      
      import defusedxml.minidom
      import defusedxml.sax
      
      
      class XMLEditor:
          """
          Editor for manipulating OOXML XML files with line-number-based node finding.
      
          This class parses XML files and tracks the original line and column position
          of each element. This enables finding nodes by their line number in the original
          file, which is useful when working with Read tool output.
      
          Attributes:
              xml_path: Path to the XML file being edited
              encoding: Detected encoding of the XML file ('ascii' or 'utf-8')
              dom: Parsed DOM tree with parse_position attributes on elements
          """
      
          def __init__(self, xml_path):
              """
              Initialize with path to XML file and parse with line number tracking.
      
              Args:
                  xml_path: Path to XML file to edit (str or Path)
      
              Raises:
                  ValueError: If the XML file does not exist
              """
              self.xml_path = Path(xml_path)
              if not self.xml_path.exists():
                  raise ValueError(f"XML file not found: {xml_path}")
      
              with open(self.xml_path, "rb") as f:
                  header = f.read(200).decode("utf-8", errors="ignore")
              self.encoding = "ascii" if 'encoding="ascii"' in header else "utf-8"
      
              parser = _create_line_tracking_parser()
              self.dom = defusedxml.minidom.parse(str(self.xml_path), parser)
      
          def get_node(
              self,
              tag: str,
              attrs: Optional[dict[str, str]] = None,
              line_number: Optional[Union[int, range]] = None,
              contains: Optional[str] = None,
          ):
              """
              Get a DOM element by tag and identifier.
      
              Finds an element by either its line number in the original file or by
              matching attribute values. Exactly one match must be found.
      
              Args:
                  tag: The XML tag name (e.g., "w:del", "w:ins", "w:r")
                  attrs: Dictionary of attribute name-value pairs to match (e.g., {"w:id": "1"})
                  line_number: Line number (int) or line range (range) in original XML file (1-indexed)
                  contains: Text string that must appear in any text node within the element.
                            Supports both entity notation (&#8220;) and Unicode characters (\u201c).
      
              Returns:
                  defusedxml.minidom.Element: The matching DOM element
      
              Raises:
                  ValueError: If node not found or multiple matches found
      
              Example:
                  elem = editor.get_node(tag="w:r", line_number=519)
                  elem = editor.get_node(tag="w:r", line_number=range(100, 200))
                  elem = editor.get_node(tag="w:del", attrs={"w:id": "1"})
                  elem = editor.get_node(tag="w:p", attrs={"w14:paraId": "12345678"})
                  elem = editor.get_node(tag="w:commentRangeStart", attrs={"w:id": "0"})
                  elem = editor.get_node(tag="w:p", contains="specific text")
                  elem = editor.get_node(tag="w:t", contains="&#8220;Agreement")  # Entity notation
                  elem = editor.get_node(tag="w:t", contains="\u201cAgreement")   # Unicode character
              """
              matches = []
              for elem in self.dom.getElementsByTagName(tag):
                  # Check line_number filter
                  if line_number is not None:
                      parse_pos = getattr(elem, "parse_position", (None,))
                      elem_line = parse_pos[0]
      
                      # Handle both single line number and range
                      if isinstance(line_number, range):
                          if elem_line not in line_number:
                              continue
                      else:
                          if elem_line != line_number:
                              continue
      
                  # Check attrs filter
                  if attrs is not None:
                      if not all(
                          elem.getAttribute(attr_name) == attr_value
                          for attr_name, attr_value in attrs.items()
                      ):
                          continue
      
                  # Check contains filter
                  if contains is not None:
                      elem_text = self._get_element_text(elem)
                      # Normalize the search string: convert HTML entities to Unicode characters
                      # This allows searching for both "&#8220;Rowan" and ""Rowan"
                      normalized_contains = html.unescape(contains)
                      if normalized_contains not in elem_text:
                          continue
      
                  # If all applicable filters passed, this is a match
                  matches.append(elem)
      
              if not matches:
                  # Build descriptive error message
                  filters = []
                  if line_number is not None:
                      line_str = (
                          f"lines {line_number.start}-{line_number.stop - 1}"
                          if isinstance(line_number, range)
                          else f"line {line_number}"
                      )
                      filters.append(f"at {line_str}")
                  if attrs is not None:
                      filters.append(f"with attributes {attrs}")
                  if contains is not None:
                      filters.append(f"containing '{contains}'")
      
                  filter_desc = " ".join(filters) if filters else ""
                  base_msg = f"Node not found: <{tag}> {filter_desc}".strip()
      
                  # Add helpful hint based on filters used
                  if contains:
                      hint = "Text may be split across elements or use different wording."
                  elif line_number:
                      hint = "Line numbers may have changed if document was modified."
                  elif attrs:
                      hint = "Verify attribute values are correct."
                  else:
                      hint = "Try adding filters (attrs, line_number, or contains)."
      
                  raise ValueError(f"{base_msg}. {hint}")
              if len(matches) > 1:
                  raise ValueError(
                      f"Multiple nodes found: <{tag}>. "
                      f"Add more filters (attrs, line_number, or contains) to narrow the search."
                  )
              return matches[0]
      
          def _get_element_text(self, elem):
              """
              Recursively extract all text content from an element.
      
              Skips text nodes that contain only whitespace (spaces, tabs, newlines),
              which typically represent XML formatting rather than document content.
      
              Args:
                  elem: defusedxml.minidom.Element to extract text from
      
              Returns:
                  str: Concatenated text from all non-whitespace text nodes within the element
              """
              text_parts = []
              for node in elem.childNodes:
                  if node.nodeType == node.TEXT_NODE:
                      # Skip whitespace-only text nodes (XML formatting)
                      if node.data.strip():
                          text_parts.append(node.data)
                  elif node.nodeType == node.ELEMENT_NODE:
                      text_parts.append(self._get_element_text(node))
              return "".join(text_parts)
      
          def replace_node(self, elem, new_content):
              """
              Replace a DOM element with new XML content.
      
              Args:
                  elem: defusedxml.minidom.Element to replace
                  new_content: String containing XML to replace the node with
      
              Returns:
                  List[defusedxml.minidom.Node]: All inserted nodes
      
              Example:
                  new_nodes = editor.replace_node(old_elem, "<w:r><w:t>text</w:t></w:r>")
              """
              parent = elem.parentNode
              nodes = self._parse_fragment(new_content)
              for node in nodes:
                  parent.insertBefore(node, elem)
              parent.removeChild(elem)
              return nodes
      
          def insert_after(self, elem, xml_content):
              """
              Insert XML content after a DOM element.
      
              Args:
                  elem: defusedxml.minidom.Element to insert after
                  xml_content: String containing XML to insert
      
              Returns:
                  List[defusedxml.minidom.Node]: All inserted nodes
      
              Example:
                  new_nodes = editor.insert_after(elem, "<w:r><w:t>text</w:t></w:r>")
              """
              parent = elem.parentNode
              next_sibling = elem.nextSibling
              nodes = self._parse_fragment(xml_content)
              for node in nodes:
                  if next_sibling:
                      parent.insertBefore(node, next_sibling)
                  else:
                      parent.appendChild(node)
              return nodes
      
          def insert_before(self, elem, xml_content):
              """
              Insert XML content before a DOM element.
      
              Args:
                  elem: defusedxml.minidom.Element to insert before
                  xml_content: String containing XML to insert
      
              Returns:
                  List[defusedxml.minidom.Node]: All inserted nodes
      
              Example:
                  new_nodes = editor.insert_before(elem, "<w:r><w:t>text</w:t></w:r>")
              """
              parent = elem.parentNode
              nodes = self._parse_fragment(xml_content)
              for node in nodes:
                  parent.insertBefore(node, elem)
              return nodes
      
          def append_to(self, elem, xml_content):
              """
              Append XML content as a child of a DOM element.
      
              Args:
                  elem: defusedxml.minidom.Element to append to
                  xml_content: String containing XML to append
      
              Returns:
                  List[defusedxml.minidom.Node]: All inserted nodes
      
              Example:
                  new_nodes = editor.append_to(elem, "<w:r><w:t>text</w:t></w:r>")
              """
              nodes = self._parse_fragment(xml_content)
              for node in nodes:
                  elem.appendChild(node)
              return nodes
      
          def get_next_rid(self):
              """Get the next available rId for relationships files."""
              max_id = 0
              for rel_elem in self.dom.getElementsByTagName("Relationship"):
                  rel_id = rel_elem.getAttribute("Id")
                  if rel_id.startswith("rId"):
                      try:
                          max_id = max(max_id, int(rel_id[3:]))
                      except ValueError:
                          pass
              return f"rId{max_id + 1}"
      
          def save(self):
              """
              Save the edited XML back to the file.
      
              Serializes the DOM tree and writes it back to the original file path,
              preserving the original encoding (ascii or utf-8).
              """
              content = self.dom.toxml(encoding=self.encoding)
              self.xml_path.write_bytes(content)
      
          def _parse_fragment(self, xml_content):
              """
              Parse XML fragment and return list of imported nodes.
      
              Args:
                  xml_content: String containing XML fragment
      
              Returns:
                  List of defusedxml.minidom.Node objects imported into this document
      
              Raises:
                  AssertionError: If fragment contains no element nodes
              """
              # Extract namespace declarations from the root document element
              root_elem = self.dom.documentElement
              namespaces = []
              if root_elem and root_elem.attributes:
                  for i in range(root_elem.attributes.length):
                      attr = root_elem.attributes.item(i)
                      if attr.name.startswith("xmlns"):  # type: ignore
                          namespaces.append(f'{attr.name}="{attr.value}"')  # type: ignore
      
              ns_decl = " ".join(namespaces)
              wrapper = f"<root {ns_decl}>{xml_content}</root>"
              fragment_doc = defusedxml.minidom.parseString(wrapper)
              nodes = [
                  self.dom.importNode(child, deep=True)
                  for child in fragment_doc.documentElement.childNodes  # type: ignore
              ]
              elements = [n for n in nodes if n.nodeType == n.ELEMENT_NODE]
              assert elements, "Fragment must contain at least one element"
              return nodes
      
      
      def _create_line_tracking_parser():
          """
          Create a SAX parser that tracks line and column numbers for each element.
      
          Monkey patches the SAX content handler to store the current line and column
          position from the underlying expat parser onto each element as a parse_position
          attribute (line, column) tuple.
      
          Returns:
              defusedxml.sax.xmlreader.XMLReader: Configured SAX parser
          """
      
          def set_content_handler(dom_handler):
              def startElementNS(name, tagName, attrs):
                  orig_start_cb(name, tagName, attrs)
                  cur_elem = dom_handler.elementStack[-1]
                  cur_elem.parse_position = (
                      parser._parser.CurrentLineNumber,  # type: ignore
                      parser._parser.CurrentColumnNumber,  # type: ignore
                  )
      
              orig_start_cb = dom_handler.startElementNS
              dom_handler.startElementNS = startElementNS
              orig_set_content_handler(dom_handler)
      
          parser = defusedxml.sax.make_parser()
          orig_set_content_handler = parser.setContentHandler
          parser.setContentHandler = set_content_handler  # type: ignore
          return parser
      
    • __init__.py 65 B
      # Make scripts directory a package for relative imports in tests
      
  • docx-js.md 16.1 KB
    # DOCX Library Tutorial
    
    Generate .docx files with JavaScript/TypeScript.
    
    **Important: Read this entire document before starting.** Critical formatting rules and common pitfalls are covered throughout - skipping sections may result in corrupted files or rendering issues.
    
    ## Setup
    Assumes docx is already installed globally
    If not installed: `npm install -g docx`
    
    ```javascript
    const { Document, Packer, Paragraph, TextRun, Table, TableRow, TableCell, ImageRun, Media, 
            Header, Footer, AlignmentType, PageOrientation, LevelFormat, ExternalHyperlink, 
            InternalHyperlink, TableOfContents, HeadingLevel, BorderStyle, WidthType, TabStopType, 
            TabStopPosition, UnderlineType, ShadingType, VerticalAlign, SymbolRun, PageNumber,
            FootnoteReferenceRun, Footnote, PageBreak } = require('docx');
    
    // Create & Save
    const doc = new Document({ sections: [{ children: [/* content */] }] });
    Packer.toBuffer(doc).then(buffer => fs.writeFileSync("doc.docx", buffer)); // Node.js
    Packer.toBlob(doc).then(blob => { /* download logic */ }); // Browser
    ```
    
    ## Text & Formatting
    ```javascript
    // IMPORTANT: Never use \n for line breaks - always use separate Paragraph elements
    // ❌ WRONG: new TextRun("Line 1\nLine 2")
    // ✅ CORRECT: new Paragraph({ children: [new TextRun("Line 1")] }), new Paragraph({ children: [new TextRun("Line 2")] })
    
    // Basic text with all formatting options
    new Paragraph({
      alignment: AlignmentType.CENTER,
      spacing: { before: 200, after: 200 },
      indent: { left: 720, right: 720 },
      children: [
        new TextRun({ text: "Bold", bold: true }),
        new TextRun({ text: "Italic", italics: true }),
        new TextRun({ text: "Underlined", underline: { type: UnderlineType.DOUBLE, color: "FF0000" } }),
        new TextRun({ text: "Colored", color: "FF0000", size: 28, font: "Arial" }), // Arial default
        new TextRun({ text: "Highlighted", highlight: "yellow" }),
        new TextRun({ text: "Strikethrough", strike: true }),
        new TextRun({ text: "x2", superScript: true }),
        new TextRun({ text: "H2O", subScript: true }),
        new TextRun({ text: "SMALL CAPS", smallCaps: true }),
        new SymbolRun({ char: "2022", font: "Symbol" }), // Bullet •
        new SymbolRun({ char: "00A9", font: "Arial" })   // Copyright © - Arial for symbols
      ]
    })
    ```
    
    ## Styles & Professional Formatting
    
    ```javascript
    const doc = new Document({
      styles: {
        default: { document: { run: { font: "Arial", size: 24 } } }, // 12pt default
        paragraphStyles: [
          // Document title style - override built-in Title style
          { id: "Title", name: "Title", basedOn: "Normal",
            run: { size: 56, bold: true, color: "000000", font: "Arial" },
            paragraph: { spacing: { before: 240, after: 120 }, alignment: AlignmentType.CENTER } },
          // IMPORTANT: Override built-in heading styles by using their exact IDs
          { id: "Heading1", name: "Heading 1", basedOn: "Normal", next: "Normal", quickFormat: true,
            run: { size: 32, bold: true, color: "000000", font: "Arial" }, // 16pt
            paragraph: { spacing: { before: 240, after: 240 }, outlineLevel: 0 } }, // Required for TOC
          { id: "Heading2", name: "Heading 2", basedOn: "Normal", next: "Normal", quickFormat: true,
            run: { size: 28, bold: true, color: "000000", font: "Arial" }, // 14pt
            paragraph: { spacing: { before: 180, after: 180 }, outlineLevel: 1 } },
          // Custom styles use your own IDs
          { id: "myStyle", name: "My Style", basedOn: "Normal",
            run: { size: 28, bold: true, color: "000000" },
            paragraph: { spacing: { after: 120 }, alignment: AlignmentType.CENTER } }
        ],
        characterStyles: [{ id: "myCharStyle", name: "My Char Style",
          run: { color: "FF0000", bold: true, underline: { type: UnderlineType.SINGLE } } }]
      },
      sections: [{
        properties: { page: { margin: { top: 1440, right: 1440, bottom: 1440, left: 1440 } } },
        children: [
          new Paragraph({ heading: HeadingLevel.TITLE, children: [new TextRun("Document Title")] }), // Uses overridden Title style
          new Paragraph({ heading: HeadingLevel.HEADING_1, children: [new TextRun("Heading 1")] }), // Uses overridden Heading1 style
          new Paragraph({ style: "myStyle", children: [new TextRun("Custom paragraph style")] }),
          new Paragraph({ children: [
            new TextRun("Normal with "),
            new TextRun({ text: "custom char style", style: "myCharStyle" })
          ]})
        ]
      }]
    });
    ```
    
    **Professional Font Combinations:**
    - **Arial (Headers) + Arial (Body)** - Most universally supported, clean and professional
    - **Times New Roman (Headers) + Arial (Body)** - Classic serif headers with modern sans-serif body
    - **Georgia (Headers) + Verdana (Body)** - Optimized for screen reading, elegant contrast
    
    **Key Styling Principles:**
    - **Override built-in styles**: Use exact IDs like "Heading1", "Heading2", "Heading3" to override Word's built-in heading styles
    - **HeadingLevel constants**: `HeadingLevel.HEADING_1` uses "Heading1" style, `HeadingLevel.HEADING_2` uses "Heading2" style, etc.
    - **Include outlineLevel**: Set `outlineLevel: 0` for H1, `outlineLevel: 1` for H2, etc. to ensure TOC works correctly
    - **Use custom styles** instead of inline formatting for consistency
    - **Set a default font** using `styles.default.document.run.font` - Arial is universally supported
    - **Establish visual hierarchy** with different font sizes (titles > headers > body)
    - **Add proper spacing** with `before` and `after` paragraph spacing
    - **Use colors sparingly**: Default to black (000000) and shades of gray for titles and headings (heading 1, heading 2, etc.)
    - **Set consistent margins** (1440 = 1 inch is standard)
    
    
    ## Lists (ALWAYS USE PROPER LISTS - NEVER USE UNICODE BULLETS)
    ```javascript
    // Bullets - ALWAYS use the numbering config, NOT unicode symbols
    // CRITICAL: Use LevelFormat.BULLET constant, NOT the string "bullet"
    const doc = new Document({
      numbering: {
        config: [
          { reference: "bullet-list",
            levels: [{ level: 0, format: LevelFormat.BULLET, text: "•", alignment: AlignmentType.LEFT,
              style: { paragraph: { indent: { left: 720, hanging: 360 } } } }] },
          { reference: "first-numbered-list",
            levels: [{ level: 0, format: LevelFormat.DECIMAL, text: "%1.", alignment: AlignmentType.LEFT,
              style: { paragraph: { indent: { left: 720, hanging: 360 } } } }] },
          { reference: "second-numbered-list", // Different reference = restarts at 1
            levels: [{ level: 0, format: LevelFormat.DECIMAL, text: "%1.", alignment: AlignmentType.LEFT,
              style: { paragraph: { indent: { left: 720, hanging: 360 } } } }] }
        ]
      },
      sections: [{
        children: [
          // Bullet list items
          new Paragraph({ numbering: { reference: "bullet-list", level: 0 },
            children: [new TextRun("First bullet point")] }),
          new Paragraph({ numbering: { reference: "bullet-list", level: 0 },
            children: [new TextRun("Second bullet point")] }),
          // Numbered list items
          new Paragraph({ numbering: { reference: "first-numbered-list", level: 0 },
            children: [new TextRun("First numbered item")] }),
          new Paragraph({ numbering: { reference: "first-numbered-list", level: 0 },
            children: [new TextRun("Second numbered item")] }),
          // ⚠️ CRITICAL: Different reference = INDEPENDENT list that restarts at 1
          // Same reference = CONTINUES previous numbering
          new Paragraph({ numbering: { reference: "second-numbered-list", level: 0 },
            children: [new TextRun("Starts at 1 again (because different reference)")] })
        ]
      }]
    });
    
    // ⚠️ CRITICAL NUMBERING RULE: Each reference creates an INDEPENDENT numbered list
    // - Same reference = continues numbering (1, 2, 3... then 4, 5, 6...)
    // - Different reference = restarts at 1 (1, 2, 3... then 1, 2, 3...)
    // Use unique reference names for each separate numbered section!
    
    // ⚠️ CRITICAL: NEVER use unicode bullets - they create fake lists that don't work properly
    // new TextRun("• Item")           // WRONG
    // new SymbolRun({ char: "2022" }) // WRONG
    // ✅ ALWAYS use numbering config with LevelFormat.BULLET for real Word lists
    ```
    
    ## Tables
    ```javascript
    // Complete table with margins, borders, headers, and bullet points
    const tableBorder = { style: BorderStyle.SINGLE, size: 1, color: "CCCCCC" };
    const cellBorders = { top: tableBorder, bottom: tableBorder, left: tableBorder, right: tableBorder };
    
    new Table({
      columnWidths: [4680, 4680], // ⚠️ CRITICAL: Set column widths at table level - values in DXA (twentieths of a point)
      margins: { top: 100, bottom: 100, left: 180, right: 180 }, // Set once for all cells
      rows: [
        new TableRow({
          tableHeader: true,
          children: [
            new TableCell({
              borders: cellBorders,
              width: { size: 4680, type: WidthType.DXA }, // ALSO set width on each cell
              // ⚠️ CRITICAL: Always use ShadingType.CLEAR to prevent black backgrounds in Word.
              shading: { fill: "D5E8F0", type: ShadingType.CLEAR }, 
              verticalAlign: VerticalAlign.CENTER,
              children: [new Paragraph({ 
                alignment: AlignmentType.CENTER,
                children: [new TextRun({ text: "Header", bold: true, size: 22 })]
              })]
            }),
            new TableCell({
              borders: cellBorders,
              width: { size: 4680, type: WidthType.DXA }, // ALSO set width on each cell
              shading: { fill: "D5E8F0", type: ShadingType.CLEAR },
              children: [new Paragraph({ 
                alignment: AlignmentType.CENTER,
                children: [new TextRun({ text: "Bullet Points", bold: true, size: 22 })]
              })]
            })
          ]
        }),
        new TableRow({
          children: [
            new TableCell({
              borders: cellBorders,
              width: { size: 4680, type: WidthType.DXA }, // ALSO set width on each cell
              children: [new Paragraph({ children: [new TextRun("Regular data")] })]
            }),
            new TableCell({
              borders: cellBorders,
              width: { size: 4680, type: WidthType.DXA }, // ALSO set width on each cell
              children: [
                new Paragraph({ 
                  numbering: { reference: "bullet-list", level: 0 },
                  children: [new TextRun("First bullet point")] 
                }),
                new Paragraph({ 
                  numbering: { reference: "bullet-list", level: 0 },
                  children: [new TextRun("Second bullet point")] 
                })
              ]
            })
          ]
        })
      ]
    })
    ```
    
    **IMPORTANT: Table Width & Borders**
    - Use BOTH `columnWidths: [width1, width2, ...]` array AND `width: { size: X, type: WidthType.DXA }` on each cell
    - Values in DXA (twentieths of a point): 1440 = 1 inch, Letter usable width = 9360 DXA (with 1" margins)
    - Apply borders to individual `TableCell` elements, NOT the `Table` itself
    
    **Precomputed Column Widths (Letter size with 1" margins = 9360 DXA total):**
    - **2 columns:** `columnWidths: [4680, 4680]` (equal width)
    - **3 columns:** `columnWidths: [3120, 3120, 3120]` (equal width)
    
    ## Links & Navigation
    ```javascript
    // TOC (requires headings) - CRITICAL: Use HeadingLevel only, NOT custom styles
    // ❌ WRONG: new Paragraph({ heading: HeadingLevel.HEADING_1, style: "customHeader", children: [new TextRun("Title")] })
    // ✅ CORRECT: new Paragraph({ heading: HeadingLevel.HEADING_1, children: [new TextRun("Title")] })
    new TableOfContents("Table of Contents", { hyperlink: true, headingStyleRange: "1-3" }),
    
    // External link
    new Paragraph({
      children: [new ExternalHyperlink({
        children: [new TextRun({ text: "Google", style: "Hyperlink" })],
        link: "https://www.google.com"
      })]
    }),
    
    // Internal link & bookmark
    new Paragraph({
      children: [new InternalHyperlink({
        children: [new TextRun({ text: "Go to Section", style: "Hyperlink" })],
        anchor: "section1"
      })]
    }),
    new Paragraph({
      children: [new TextRun("Section Content")],
      bookmark: { id: "section1", name: "section1" }
    }),
    ```
    
    ## Images & Media
    ```javascript
    // Basic image with sizing & positioning
    // CRITICAL: Always specify 'type' parameter - it's REQUIRED for ImageRun
    new Paragraph({
      alignment: AlignmentType.CENTER,
      children: [new ImageRun({
        type: "png", // NEW REQUIREMENT: Must specify image type (png, jpg, jpeg, gif, bmp, svg)
        data: fs.readFileSync("image.png"),
        transformation: { width: 200, height: 150, rotation: 0 }, // rotation in degrees
        altText: { title: "Logo", description: "Company logo", name: "Name" } // IMPORTANT: All three fields are required
      })]
    })
    ```
    
    ## Page Breaks
    ```javascript
    // Manual page break
    new Paragraph({ children: [new PageBreak()] }),
    
    // Page break before paragraph
    new Paragraph({
      pageBreakBefore: true,
      children: [new TextRun("This starts on a new page")]
    })
    
    // ⚠️ CRITICAL: NEVER use PageBreak standalone - it will create invalid XML that Word cannot open
    // ❌ WRONG: new PageBreak() 
    // ✅ CORRECT: new Paragraph({ children: [new PageBreak()] })
    ```
    
    ## Headers/Footers & Page Setup
    ```javascript
    const doc = new Document({
      sections: [{
        properties: {
          page: {
            margin: { top: 1440, right: 1440, bottom: 1440, left: 1440 }, // 1440 = 1 inch
            size: { orientation: PageOrientation.LANDSCAPE },
            pageNumbers: { start: 1, formatType: "decimal" } // "upperRoman", "lowerRoman", "upperLetter", "lowerLetter"
          }
        },
        headers: {
          default: new Header({ children: [new Paragraph({ 
            alignment: AlignmentType.RIGHT,
            children: [new TextRun("Header Text")]
          })] })
        },
        footers: {
          default: new Footer({ children: [new Paragraph({ 
            alignment: AlignmentType.CENTER,
            children: [new TextRun("Page "), new TextRun({ children: [PageNumber.CURRENT] }), new TextRun(" of "), new TextRun({ children: [PageNumber.TOTAL_PAGES] })]
          })] })
        },
        children: [/* content */]
      }]
    });
    ```
    
    ## Tabs
    ```javascript
    new Paragraph({
      tabStops: [
        { type: TabStopType.LEFT, position: TabStopPosition.MAX / 4 },
        { type: TabStopType.CENTER, position: TabStopPosition.MAX / 2 },
        { type: TabStopType.RIGHT, position: TabStopPosition.MAX * 3 / 4 }
      ],
      children: [new TextRun("Left\tCenter\tRight")]
    })
    ```
    
    ## Constants & Quick Reference
    - **Underlines:** `SINGLE`, `DOUBLE`, `WAVY`, `DASH`
    - **Borders:** `SINGLE`, `DOUBLE`, `DASHED`, `DOTTED`  
    - **Numbering:** `DECIMAL` (1,2,3), `UPPER_ROMAN` (I,II,III), `LOWER_LETTER` (a,b,c)
    - **Tabs:** `LEFT`, `CENTER`, `RIGHT`, `DECIMAL`
    - **Symbols:** `"2022"` (•), `"00A9"` (©), `"00AE"` (®), `"2122"` (™), `"00B0"` (°), `"F070"` (✓), `"F0FC"` (✗)
    
    ## Critical Issues & Common Mistakes
    - **CRITICAL: PageBreak must ALWAYS be inside a Paragraph** - standalone PageBreak creates invalid XML that Word cannot open
    - **ALWAYS use ShadingType.CLEAR for table cell shading** - Never use ShadingType.SOLID (causes black background).
    - Measurements in DXA (1440 = 1 inch) | Each table cell needs ≥1 Paragraph | TOC requires HeadingLevel styles only
    - **ALWAYS use custom styles** with Arial font for professional appearance and proper visual hierarchy
    - **ALWAYS set a default font** using `styles.default.document.run.font` - Arial recommended
    - **ALWAYS use columnWidths array for tables** + individual cell widths for compatibility
    - **NEVER use unicode symbols for bullets** - always use proper numbering configuration with `LevelFormat.BULLET` constant (NOT the string "bullet")
    - **NEVER use \n for line breaks anywhere** - always use separate Paragraph elements for each line
    - **ALWAYS use TextRun objects within Paragraph children** - never use text property directly on Paragraph
    - **CRITICAL for images**: ImageRun REQUIRES `type` parameter - always specify "png", "jpg", "jpeg", "gif", "bmp", or "svg"
    - **CRITICAL for bullets**: Must use `LevelFormat.BULLET` constant, not string "bullet", and include `text: "•"` for the bullet character
    - **CRITICAL for numbering**: Each numbering reference creates an INDEPENDENT list. Same reference = continues numbering (1,2,3 then 4,5,6). Different reference = restarts at 1 (1,2,3 then 1,2,3). Use unique reference names for each separate numbered section!
    - **CRITICAL for TOC**: When using TableOfContents, headings must use HeadingLevel ONLY - do NOT add custom styles to heading paragraphs or TOC will break
    - **Tables**: Set `columnWidths` array + individual cell widths, apply borders to cells not table
    - **Set table margins at TABLE level** for consistent cell padding (avoids repetition per cell)
  • LICENSE.txt 1.4 KB
    © 2025 Anthropic, PBC. All rights reserved.
    
    LICENSE: Use of these materials (including all code, prompts, assets, files,
    and other components of this Skill) is governed by your agreement with
    Anthropic regarding use of Anthropic's services. If no separate agreement
    exists, use is governed by Anthropic's Consumer Terms of Service or
    Commercial Terms of Service, as applicable:
    https://www.anthropic.com/legal/consumer-terms
    https://www.anthropic.com/legal/commercial-terms
    Your applicable agreement is referred to as the "Agreement." "Services" are
    as defined in the Agreement.
    
    ADDITIONAL RESTRICTIONS: Notwithstanding anything in the Agreement to the
    contrary, users may not:
    
    - Extract these materials from the Services or retain copies of these
      materials outside the Services
    - Reproduce or copy these materials, except for temporary copies created
      automatically during authorized use of the Services
    - Create derivative works based on these materials
    - Distribute, sublicense, or transfer these materials to any third party
    - Make, offer to sell, sell, or import any inventions embodied in these
      materials
    - Reverse engineer, decompile, or disassemble these materials
    
    The receipt, viewing, or possession of these materials does not convey or
    imply any license or right beyond those expressly granted above.
    
    Anthropic retains all right, title, and interest in these materials,
    including all copyrights, patents, and other intellectual property rights.
    
  • ooxml.md 23 KB
    # Office Open XML Technical Reference
    
    **Important: Read this entire document before starting.** This document covers:
    - [Technical Guidelines](#technical-guidelines) - Schema compliance rules and validation requirements
    - [Document Content Patterns](#document-content-patterns) - XML patterns for headings, lists, tables, formatting, etc.
    - [Document Library (Python)](#document-library-python) - Recommended approach for OOXML manipulation with automatic infrastructure setup
    - [Tracked Changes (Redlining)](#tracked-changes-redlining) - XML patterns for implementing tracked changes
    
    ## Technical Guidelines
    
    ### Schema Compliance
    - **Element ordering in `<w:pPr>`**: `<w:pStyle>`, `<w:numPr>`, `<w:spacing>`, `<w:ind>`, `<w:jc>`
    - **Whitespace**: Add `xml:space='preserve'` to `<w:t>` elements with leading/trailing spaces
    - **Unicode**: Escape characters in ASCII content: `"` becomes `&#8220;`
      - **Character encoding reference**: Curly quotes `""` become `&#8220;&#8221;`, apostrophe `'` becomes `&#8217;`, em-dash `—` becomes `&#8212;`
    - **Tracked changes**: Use `<w:del>` and `<w:ins>` tags with `w:author="Claude"` outside `<w:r>` elements
      - **Critical**: `<w:ins>` closes with `</w:ins>`, `<w:del>` closes with `</w:del>` - never mix
      - **RSIDs must be 8-digit hex**: Use values like `00AB1234` (only 0-9, A-F characters)
      - **trackRevisions placement**: Add `<w:trackRevisions/>` after `<w:proofState>` in settings.xml
    - **Images**: Add to `word/media/`, reference in `document.xml`, set dimensions to prevent overflow
    
    ## Document Content Patterns
    
    ### Basic Structure
    ```xml
    <w:p>
      <w:r><w:t>Text content</w:t></w:r>
    </w:p>
    ```
    
    ### Headings and Styles
    ```xml
    <w:p>
      <w:pPr>
        <w:pStyle w:val="Title"/>
        <w:jc w:val="center"/>
      </w:pPr>
      <w:r><w:t>Document Title</w:t></w:r>
    </w:p>
    
    <w:p>
      <w:pPr><w:pStyle w:val="Heading2"/></w:pPr>
      <w:r><w:t>Section Heading</w:t></w:r>
    </w:p>
    ```
    
    ### Text Formatting
    ```xml
    <!-- Bold -->
    <w:r><w:rPr><w:b/><w:bCs/></w:rPr><w:t>Bold</w:t></w:r>
    <!-- Italic -->
    <w:r><w:rPr><w:i/><w:iCs/></w:rPr><w:t>Italic</w:t></w:r>
    <!-- Underline -->
    <w:r><w:rPr><w:u w:val="single"/></w:rPr><w:t>Underlined</w:t></w:r>
    <!-- Highlight -->
    <w:r><w:rPr><w:highlight w:val="yellow"/></w:rPr><w:t>Highlighted</w:t></w:r>
    ```
    
    ### Lists
    ```xml
    <!-- Numbered list -->
    <w:p>
      <w:pPr>
        <w:pStyle w:val="ListParagraph"/>
        <w:numPr><w:ilvl w:val="0"/><w:numId w:val="1"/></w:numPr>
        <w:spacing w:before="240"/>
      </w:pPr>
      <w:r><w:t>First item</w:t></w:r>
    </w:p>
    
    <!-- Restart numbered list at 1 - use different numId -->
    <w:p>
      <w:pPr>
        <w:pStyle w:val="ListParagraph"/>
        <w:numPr><w:ilvl w:val="0"/><w:numId w:val="2"/></w:numPr>
        <w:spacing w:before="240"/>
      </w:pPr>
      <w:r><w:t>New list item 1</w:t></w:r>
    </w:p>
    
    <!-- Bullet list (level 2) -->
    <w:p>
      <w:pPr>
        <w:pStyle w:val="ListParagraph"/>
        <w:numPr><w:ilvl w:val="1"/><w:numId w:val="1"/></w:numPr>
        <w:spacing w:before="240"/>
        <w:ind w:left="900"/>
      </w:pPr>
      <w:r><w:t>Bullet item</w:t></w:r>
    </w:p>
    ```
    
    ### Tables
    ```xml
    <w:tbl>
      <w:tblPr>
        <w:tblStyle w:val="TableGrid"/>
        <w:tblW w:w="0" w:type="auto"/>
      </w:tblPr>
      <w:tblGrid>
        <w:gridCol w:w="4675"/><w:gridCol w:w="4675"/>
      </w:tblGrid>
      <w:tr>
        <w:tc>
          <w:tcPr><w:tcW w:w="4675" w:type="dxa"/></w:tcPr>
          <w:p><w:r><w:t>Cell 1</w:t></w:r></w:p>
        </w:tc>
        <w:tc>
          <w:tcPr><w:tcW w:w="4675" w:type="dxa"/></w:tcPr>
          <w:p><w:r><w:t>Cell 2</w:t></w:r></w:p>
        </w:tc>
      </w:tr>
    </w:tbl>
    ```
    
    ### Layout
    ```xml
    <!-- Page break before new section (common pattern) -->
    <w:p>
      <w:r>
        <w:br w:type="page"/>
      </w:r>
    </w:p>
    <w:p>
      <w:pPr>
        <w:pStyle w:val="Heading1"/>
      </w:pPr>
      <w:r>
        <w:t>New Section Title</w:t>
      </w:r>
    </w:p>
    
    <!-- Centered paragraph -->
    <w:p>
      <w:pPr>
        <w:spacing w:before="240" w:after="0"/>
        <w:jc w:val="center"/>
      </w:pPr>
      <w:r><w:t>Centered text</w:t></w:r>
    </w:p>
    
    <!-- Font change - paragraph level (applies to all runs) -->
    <w:p>
      <w:pPr>
        <w:rPr><w:rFonts w:ascii="Courier New" w:hAnsi="Courier New"/></w:rPr>
      </w:pPr>
      <w:r><w:t>Monospace text</w:t></w:r>
    </w:p>
    
    <!-- Font change - run level (specific to this text) -->
    <w:p>
      <w:r>
        <w:rPr><w:rFonts w:ascii="Courier New" w:hAnsi="Courier New"/></w:rPr>
        <w:t>This text is Courier New</w:t>
      </w:r>
      <w:r><w:t> and this text uses default font</w:t></w:r>
    </w:p>
    ```
    
    ## File Updates
    
    When adding content, update these files:
    
    **`word/_rels/document.xml.rels`:**
    ```xml
    <Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/numbering" Target="numbering.xml"/>
    <Relationship Id="rId5" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/image" Target="media/image1.png"/>
    ```
    
    **`[Content_Types].xml`:**
    ```xml
    <Default Extension="png" ContentType="image/png"/>
    <Override PartName="/word/numbering.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.numbering+xml"/>
    ```
    
    ### Images
    **CRITICAL**: Calculate dimensions to prevent page overflow and maintain aspect ratio.
    
    ```xml
    <!-- Minimal required structure -->
    <w:p>
      <w:r>
        <w:drawing>
          <wp:inline>
            <wp:extent cx="2743200" cy="1828800"/>
            <wp:docPr id="1" name="Picture 1"/>
            <a:graphic xmlns:a="http://schemas.openxmlformats.org/drawingml/2006/main">
              <a:graphicData uri="http://schemas.openxmlformats.org/drawingml/2006/picture">
                <pic:pic xmlns:pic="http://schemas.openxmlformats.org/drawingml/2006/picture">
                  <pic:nvPicPr>
                    <pic:cNvPr id="0" name="image1.png"/>
                    <pic:cNvPicPr/>
                  </pic:nvPicPr>
                  <pic:blipFill>
                    <a:blip r:embed="rId5"/>
                    <!-- Add for stretch fill with aspect ratio preservation -->
                    <a:stretch>
                      <a:fillRect/>
                    </a:stretch>
                  </pic:blipFill>
                  <pic:spPr>
                    <a:xfrm>
                      <a:ext cx="2743200" cy="1828800"/>
                    </a:xfrm>
                    <a:prstGeom prst="rect"/>
                  </pic:spPr>
                </pic:pic>
              </a:graphicData>
            </a:graphic>
          </wp:inline>
        </w:drawing>
      </w:r>
    </w:p>
    ```
    
    ### Links (Hyperlinks)
    
    **IMPORTANT**: All hyperlinks (both internal and external) require the Hyperlink style to be defined in styles.xml. Without this style, links will look like regular text instead of blue underlined clickable links.
    
    **External Links:**
    ```xml
    <!-- In document.xml -->
    <w:hyperlink r:id="rId5">
      <w:r>
        <w:rPr><w:rStyle w:val="Hyperlink"/></w:rPr>
        <w:t>Link Text</w:t>
      </w:r>
    </w:hyperlink>
    
    <!-- In word/_rels/document.xml.rels -->
    <Relationship Id="rId5" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/hyperlink" 
                  Target="https://www.example.com/" TargetMode="External"/>
    ```
    
    **Internal Links:**
    
    ```xml
    <!-- Link to bookmark -->
    <w:hyperlink w:anchor="myBookmark">
      <w:r>
        <w:rPr><w:rStyle w:val="Hyperlink"/></w:rPr>
        <w:t>Link Text</w:t>
      </w:r>
    </w:hyperlink>
    
    <!-- Bookmark target -->
    <w:bookmarkStart w:id="0" w:name="myBookmark"/>
    <w:r><w:t>Target content</w:t></w:r>
    <w:bookmarkEnd w:id="0"/>
    ```
    
    **Hyperlink Style (required in styles.xml):**
    ```xml
    <w:style w:type="character" w:styleId="Hyperlink">
      <w:name w:val="Hyperlink"/>
      <w:basedOn w:val="DefaultParagraphFont"/>
      <w:uiPriority w:val="99"/>
      <w:unhideWhenUsed/>
      <w:rPr>
        <w:color w:val="467886" w:themeColor="hyperlink"/>
        <w:u w:val="single"/>
      </w:rPr>
    </w:style>
    ```
    
    ## Document Library (Python)
    
    Use the Document class from `scripts/document.py` for all tracked changes and comments. It automatically handles infrastructure setup (people.xml, RSIDs, settings.xml, comment files, relationships, content types). Only use direct XML manipulation for complex scenarios not supported by the library.
    
    **Working with Unicode and Entities:**
    - **Searching**: Both entity notation and Unicode characters work - `contains="&#8220;Company"` and `contains="\u201cCompany"` find the same text
    - **Replacing**: Use either entities (`&#8220;`) or Unicode (`\u201c`) - both work and will be converted appropriately based on the file's encoding (ascii → entities, utf-8 → Unicode)
    
    ### Initialization
    
    **Find the docx skill root** (directory containing `scripts/` and `ooxml/`):
    ```bash
    # Search for document.py to locate the skill root
    # Note: /mnt/skills is used here as an example; check your context for the actual location
    find /mnt/skills -name "document.py" -path "*/docx/scripts/*" 2>/dev/null | head -1
    # Example output: /mnt/skills/docx/scripts/document.py
    # Skill root is: /mnt/skills/docx
    ```
    
    **Run your script with PYTHONPATH** set to the docx skill root:
    ```bash
    PYTHONPATH=/mnt/skills/docx python your_script.py
    ```
    
    **In your script**, import from the skill root:
    ```python
    from scripts.document import Document, DocxXMLEditor
    
    # Basic initialization (automatically creates temp copy and sets up infrastructure)
    doc = Document('unpacked')
    
    # Customize author and initials
    doc = Document('unpacked', author="John Doe", initials="JD")
    
    # Enable track revisions mode
    doc = Document('unpacked', track_revisions=True)
    
    # Specify custom RSID (auto-generated if not provided)
    doc = Document('unpacked', rsid="07DC5ECB")
    ```
    
    ### Creating Tracked Changes
    
    **CRITICAL**: Only mark text that actually changes. Keep ALL unchanged text outside `<w:del>`/`<w:ins>` tags. Marking unchanged text makes edits unprofessional and harder to review.
    
    **Attribute Handling**: The Document class auto-injects attributes (w:id, w:date, w:rsidR, w:rsidDel, w16du:dateUtc, xml:space) into new elements. When preserving unchanged text from the original document, copy the original `<w:r>` element with its existing attributes to maintain document integrity.
    
    **Method Selection Guide**:
    - **Adding your own changes to regular text**: Use `replace_node()` with `<w:del>`/`<w:ins>` tags, or `suggest_deletion()` for removing entire `<w:r>` or `<w:p>` elements
    - **Partially modifying another author's tracked change**: Use `replace_node()` to nest your changes inside their `<w:ins>`/`<w:del>`
    - **Completely rejecting another author's insertion**: Use `revert_insertion()` on the `<w:ins>` element (NOT `suggest_deletion()`)
    - **Completely rejecting another author's deletion**: Use `revert_deletion()` on the `<w:del>` element to restore deleted content using tracked changes
    
    ```python
    # Minimal edit - change one word: "The report is monthly" → "The report is quarterly"
    # Original: <w:r w:rsidR="00AB12CD"><w:rPr><w:rFonts w:ascii="Calibri"/></w:rPr><w:t>The report is monthly</w:t></w:r>
    node = doc["word/document.xml"].get_node(tag="w:r", contains="The report is monthly")
    rpr = tags[0].toxml() if (tags := node.getElementsByTagName("w:rPr")) else ""
    replacement = f'<w:r w:rsidR="00AB12CD">{rpr}<w:t>The report is </w:t></w:r><w:del><w:r>{rpr}<w:delText>monthly</w:delText></w:r></w:del><w:ins><w:r>{rpr}<w:t>quarterly</w:t></w:r></w:ins>'
    doc["word/document.xml"].replace_node(node, replacement)
    
    # Minimal edit - change number: "within 30 days" → "within 45 days"
    # Original: <w:r w:rsidR="00XYZ789"><w:rPr><w:rFonts w:ascii="Calibri"/></w:rPr><w:t>within 30 days</w:t></w:r>
    node = doc["word/document.xml"].get_node(tag="w:r", contains="within 30 days")
    rpr = tags[0].toxml() if (tags := node.getElementsByTagName("w:rPr")) else ""
    replacement = f'<w:r w:rsidR="00XYZ789">{rpr}<w:t>within </w:t></w:r><w:del><w:r>{rpr}<w:delText>30</w:delText></w:r></w:del><w:ins><w:r>{rpr}<w:t>45</w:t></w:r></w:ins><w:r w:rsidR="00XYZ789">{rpr}<w:t> days</w:t></w:r>'
    doc["word/document.xml"].replace_node(node, replacement)
    
    # Complete replacement - preserve formatting even when replacing all text
    node = doc["word/document.xml"].get_node(tag="w:r", contains="apple")
    rpr = tags[0].toxml() if (tags := node.getElementsByTagName("w:rPr")) else ""
    replacement = f'<w:del><w:r>{rpr}<w:delText>apple</w:delText></w:r></w:del><w:ins><w:r>{rpr}<w:t>banana orange</w:t></w:r></w:ins>'
    doc["word/document.xml"].replace_node(node, replacement)
    
    # Insert new content (no attributes needed - auto-injected)
    node = doc["word/document.xml"].get_node(tag="w:r", contains="existing text")
    doc["word/document.xml"].insert_after(node, '<w:ins><w:r><w:t>new text</w:t></w:r></w:ins>')
    
    # Partially delete another author's insertion
    # Original: <w:ins w:author="Jane Smith" w:date="..."><w:r><w:t>quarterly financial report</w:t></w:r></w:ins>
    # Goal: Delete only "financial" to make it "quarterly report"
    node = doc["word/document.xml"].get_node(tag="w:ins", attrs={"w:id": "5"})
    # IMPORTANT: Preserve w:author="Jane Smith" on the outer <w:ins> to maintain authorship
    replacement = '''<w:ins w:author="Jane Smith" w:date="2025-01-15T10:00:00Z">
      <w:r><w:t>quarterly </w:t></w:r>
      <w:del><w:r><w:delText>financial </w:delText></w:r></w:del>
      <w:r><w:t>report</w:t></w:r>
    </w:ins>'''
    doc["word/document.xml"].replace_node(node, replacement)
    
    # Change part of another author's insertion
    # Original: <w:ins w:author="Jane Smith"><w:r><w:t>in silence, safe and sound</w:t></w:r></w:ins>
    # Goal: Change "safe and sound" to "soft and unbound"
    node = doc["word/document.xml"].get_node(tag="w:ins", attrs={"w:id": "8"})
    replacement = f'''<w:ins w:author="Jane Smith" w:date="2025-01-15T10:00:00Z">
      <w:r><w:t>in silence, </w:t></w:r>
    </w:ins>
    <w:ins>
      <w:r><w:t>soft and unbound</w:t></w:r>
    </w:ins>
    <w:ins w:author="Jane Smith" w:date="2025-01-15T10:00:00Z">
      <w:del><w:r><w:delText>safe and sound</w:delText></w:r></w:del>
    </w:ins>'''
    doc["word/document.xml"].replace_node(node, replacement)
    
    # Delete entire run (use only when deleting all content; use replace_node for partial deletions)
    node = doc["word/document.xml"].get_node(tag="w:r", contains="text to delete")
    doc["word/document.xml"].suggest_deletion(node)
    
    # Delete entire paragraph (in-place, handles both regular and numbered list paragraphs)
    para = doc["word/document.xml"].get_node(tag="w:p", contains="paragraph to delete")
    doc["word/document.xml"].suggest_deletion(para)
    
    # Add new numbered list item
    target_para = doc["word/document.xml"].get_node(tag="w:p", contains="existing list item")
    pPr = tags[0].toxml() if (tags := target_para.getElementsByTagName("w:pPr")) else ""
    new_item = f'<w:p>{pPr}<w:r><w:t>New item</w:t></w:r></w:p>'
    tracked_para = DocxXMLEditor.suggest_paragraph(new_item)
    doc["word/document.xml"].insert_after(target_para, tracked_para)
    # Optional: add spacing paragraph before content for better visual separation
    # spacing = DocxXMLEditor.suggest_paragraph('<w:p><w:pPr><w:pStyle w:val="ListParagraph"/></w:pPr></w:p>')
    # doc["word/document.xml"].insert_after(target_para, spacing + tracked_para)
    ```
    
    ### Adding Comments
    
    ```python
    # Add comment spanning two existing tracked changes
    # Note: w:id is auto-generated. Only search by w:id if you know it from XML inspection
    start_node = doc["word/document.xml"].get_node(tag="w:del", attrs={"w:id": "1"})
    end_node = doc["word/document.xml"].get_node(tag="w:ins", attrs={"w:id": "2"})
    doc.add_comment(start=start_node, end=end_node, text="Explanation of this change")
    
    # Add comment on a paragraph
    para = doc["word/document.xml"].get_node(tag="w:p", contains="paragraph text")
    doc.add_comment(start=para, end=para, text="Comment on this paragraph")
    
    # Add comment on newly created tracked change
    # First create the tracked change
    node = doc["word/document.xml"].get_node(tag="w:r", contains="old")
    new_nodes = doc["word/document.xml"].replace_node(
        node,
        '<w:del><w:r><w:delText>old</w:delText></w:r></w:del><w:ins><w:r><w:t>new</w:t></w:r></w:ins>'
    )
    # Then add comment on the newly created elements
    # new_nodes[0] is the <w:del>, new_nodes[1] is the <w:ins>
    doc.add_comment(start=new_nodes[0], end=new_nodes[1], text="Changed old to new per requirements")
    
    # Reply to existing comment
    doc.reply_to_comment(parent_comment_id=0, text="I agree with this change")
    ```
    
    ### Rejecting Tracked Changes
    
    **IMPORTANT**: Use `revert_insertion()` to reject insertions and `revert_deletion()` to restore deletions using tracked changes. Use `suggest_deletion()` only for regular unmarked content.
    
    ```python
    # Reject insertion (wraps it in deletion)
    # Use this when another author inserted text that you want to delete
    ins = doc["word/document.xml"].get_node(tag="w:ins", attrs={"w:id": "5"})
    nodes = doc["word/document.xml"].revert_insertion(ins)  # Returns [ins]
    
    # Reject deletion (creates insertion to restore deleted content)
    # Use this when another author deleted text that you want to restore
    del_elem = doc["word/document.xml"].get_node(tag="w:del", attrs={"w:id": "3"})
    nodes = doc["word/document.xml"].revert_deletion(del_elem)  # Returns [del_elem, new_ins]
    
    # Reject all insertions in a paragraph
    para = doc["word/document.xml"].get_node(tag="w:p", contains="paragraph text")
    nodes = doc["word/document.xml"].revert_insertion(para)  # Returns [para]
    
    # Reject all deletions in a paragraph
    para = doc["word/document.xml"].get_node(tag="w:p", contains="paragraph text")
    nodes = doc["word/document.xml"].revert_deletion(para)  # Returns [para]
    ```
    
    ### Inserting Images
    
    **CRITICAL**: The Document class works with a temporary copy at `doc.unpacked_path`. Always copy images to this temp directory, not the original unpacked folder.
    
    ```python
    from PIL import Image
    import shutil, os
    
    # Initialize document first
    doc = Document('unpacked')
    
    # Copy image and calculate full-width dimensions with aspect ratio
    media_dir = os.path.join(doc.unpacked_path, 'word/media')
    os.makedirs(media_dir, exist_ok=True)
    shutil.copy('image.png', os.path.join(media_dir, 'image1.png'))
    img = Image.open(os.path.join(media_dir, 'image1.png'))
    width_emus = int(6.5 * 914400)  # 6.5" usable width, 914400 EMUs/inch
    height_emus = int(width_emus * img.size[1] / img.size[0])
    
    # Add relationship and content type
    rels_editor = doc['word/_rels/document.xml.rels']
    next_rid = rels_editor.get_next_rid()
    rels_editor.append_to(rels_editor.dom.documentElement,
        f'<Relationship Id="{next_rid}" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/image" Target="media/image1.png"/>')
    doc['[Content_Types].xml'].append_to(doc['[Content_Types].xml'].dom.documentElement,
        '<Default Extension="png" ContentType="image/png"/>')
    
    # Insert image
    node = doc["word/document.xml"].get_node(tag="w:p", line_number=100)
    doc["word/document.xml"].insert_after(node, f'''<w:p>
      <w:r>
        <w:drawing>
          <wp:inline distT="0" distB="0" distL="0" distR="0">
            <wp:extent cx="{width_emus}" cy="{height_emus}"/>
            <wp:docPr id="1" name="Picture 1"/>
            <a:graphic xmlns:a="http://schemas.openxmlformats.org/drawingml/2006/main">
              <a:graphicData uri="http://schemas.openxmlformats.org/drawingml/2006/picture">
                <pic:pic xmlns:pic="http://schemas.openxmlformats.org/drawingml/2006/picture">
                  <pic:nvPicPr><pic:cNvPr id="1" name="image1.png"/><pic:cNvPicPr/></pic:nvPicPr>
                  <pic:blipFill><a:blip r:embed="{next_rid}"/><a:stretch><a:fillRect/></a:stretch></pic:blipFill>
                  <pic:spPr><a:xfrm><a:ext cx="{width_emus}" cy="{height_emus}"/></a:xfrm><a:prstGeom prst="rect"><a:avLst/></a:prstGeom></pic:spPr>
                </pic:pic>
              </a:graphicData>
            </a:graphic>
          </wp:inline>
        </w:drawing>
      </w:r>
    </w:p>''')
    ```
    
    ### Getting Nodes
    
    ```python
    # By text content
    node = doc["word/document.xml"].get_node(tag="w:p", contains="specific text")
    
    # By line range
    para = doc["word/document.xml"].get_node(tag="w:p", line_number=range(100, 150))
    
    # By attributes
    node = doc["word/document.xml"].get_node(tag="w:del", attrs={"w:id": "1"})
    
    # By exact line number (must be line number where tag opens)
    para = doc["word/document.xml"].get_node(tag="w:p", line_number=42)
    
    # Combine filters
    node = doc["word/document.xml"].get_node(tag="w:r", line_number=range(40, 60), contains="text")
    
    # Disambiguate when text appears multiple times - add line_number range
    node = doc["word/document.xml"].get_node(tag="w:r", contains="Section", line_number=range(2400, 2500))
    ```
    
    ### Saving
    
    ```python
    # Save with automatic validation (copies back to original directory)
    doc.save()  # Validates by default, raises error if validation fails
    
    # Save to different location
    doc.save('modified-unpacked')
    
    # Skip validation (debugging only - needing this in production indicates XML issues)
    doc.save(validate=False)
    ```
    
    ### Direct DOM Manipulation
    
    For complex scenarios not covered by the library:
    
    ```python
    # Access any XML file
    editor = doc["word/document.xml"]
    editor = doc["word/comments.xml"]
    
    # Direct DOM access (defusedxml.minidom.Document)
    node = doc["word/document.xml"].get_node(tag="w:p", line_number=5)
    parent = node.parentNode
    parent.removeChild(node)
    parent.appendChild(node)  # Move to end
    
    # General document manipulation (without tracked changes)
    old_node = doc["word/document.xml"].get_node(tag="w:p", contains="original text")
    doc["word/document.xml"].replace_node(old_node, "<w:p><w:r><w:t>replacement text</w:t></w:r></w:p>")
    
    # Multiple insertions - use return value to maintain order
    node = doc["word/document.xml"].get_node(tag="w:r", line_number=100)
    nodes = doc["word/document.xml"].insert_after(node, "<w:r><w:t>A</w:t></w:r>")
    nodes = doc["word/document.xml"].insert_after(nodes[-1], "<w:r><w:t>B</w:t></w:r>")
    nodes = doc["word/document.xml"].insert_after(nodes[-1], "<w:r><w:t>C</w:t></w:r>")
    # Results in: original_node, A, B, C
    ```
    
    ## Tracked Changes (Redlining)
    
    **Use the Document class above for all tracked changes.** The patterns below are for reference when constructing replacement XML strings.
    
    ### Validation Rules
    The validator checks that the document text matches the original after reverting Claude's changes. This means:
    - **NEVER modify text inside another author's `<w:ins>` or `<w:del>` tags**
    - **ALWAYS use nested deletions** to remove another author's insertions
    - **Every edit must be properly tracked** with `<w:ins>` or `<w:del>` tags
    
    ### Tracked Change Patterns
    
    **CRITICAL RULES**:
    1. Never modify the content inside another author's tracked changes. Always use nested deletions.
    2. **XML Structure**: Always place `<w:del>` and `<w:ins>` at paragraph level containing complete `<w:r>` elements. Never nest inside `<w:r>` elements - this creates invalid XML that breaks document processing.
    
    **Text Insertion:**
    ```xml
    <w:ins w:id="1" w:author="Claude" w:date="2025-07-30T23:05:00Z" w16du:dateUtc="2025-07-31T06:05:00Z">
      <w:r w:rsidR="00792858">
        <w:t>inserted text</w:t>
      </w:r>
    </w:ins>
    ```
    
    **Text Deletion:**
    ```xml
    <w:del w:id="2" w:author="Claude" w:date="2025-07-30T23:05:00Z" w16du:dateUtc="2025-07-31T06:05:00Z">
      <w:r w:rsidDel="00792858">
        <w:delText>deleted text</w:delText>
      </w:r>
    </w:del>
    ```
    
    **Deleting Another Author's Insertion (MUST use nested structure):**
    ```xml
    <!-- Nest deletion inside the original insertion -->
    <w:ins w:author="Jane Smith" w:id="16">
      <w:del w:author="Claude" w:id="40">
        <w:r><w:delText>monthly</w:delText></w:r>
      </w:del>
    </w:ins>
    <w:ins w:author="Claude" w:id="41">
      <w:r><w:t>weekly</w:t></w:r>
    </w:ins>
    ```
    
    **Restoring Another Author's Deletion:**
    ```xml
    <!-- Leave their deletion unchanged, add new insertion after it -->
    <w:del w:author="Jane Smith" w:id="50">
      <w:r><w:delText>within 30 days</w:delText></w:r>
    </w:del>
    <w:ins w:author="Claude" w:id="51">
      <w:r><w:t>within 30 days</w:t></w:r>
    </w:ins>
    ```
  • SKILL.md 9.9 KB
    ---
    name: docx
    description: "Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction. When Claude needs to work with professional documents (.docx files) for: (1) Creating new documents, (2) Modifying or editing content, (3) Working with tracked changes, (4) Adding comments, or any other document tasks"
    license: Proprietary. LICENSE.txt has complete terms
    ---
    
    # DOCX creation, editing, and analysis
    
    ## Overview
    
    A user may ask you to create, edit, or analyze the contents of a .docx file. A .docx file is essentially a ZIP archive containing XML files and other resources that you can read or edit. You have different tools and workflows available for different tasks.
    
    ## Workflow Decision Tree
    
    ### Reading/Analyzing Content
    Use "Text extraction" or "Raw XML access" sections below
    
    ### Creating New Document
    Use "Creating a new Word document" workflow
    
    ### Editing Existing Document
    - **Your own document + simple changes**
      Use "Basic OOXML editing" workflow
    
    - **Someone else's document**
      Use **"Redlining workflow"** (recommended default)
    
    - **Legal, academic, business, or government docs**
      Use **"Redlining workflow"** (required)
    
    ## Reading and analyzing content
    
    ### Text extraction
    If you just need to read the text contents of a document, you should convert the document to markdown using pandoc. Pandoc provides excellent support for preserving document structure and can show tracked changes:
    
    ```bash
    # Convert document to markdown with tracked changes
    pandoc --track-changes=all path-to-file.docx -o output.md
    # Options: --track-changes=accept/reject/all
    ```
    
    ### Raw XML access
    You need raw XML access for: comments, complex formatting, document structure, embedded media, and metadata. For any of these features, you'll need to unpack a document and read its raw XML contents.
    
    #### Unpacking a file
    `python ooxml/scripts/unpack.py <office_file> <output_directory>`
    
    #### Key file structures
    * `word/document.xml` - Main document contents
    * `word/comments.xml` - Comments referenced in document.xml
    * `word/media/` - Embedded images and media files
    * Tracked changes use `<w:ins>` (insertions) and `<w:del>` (deletions) tags
    
    ## Creating a new Word document
    
    When creating a new Word document from scratch, use **docx-js**, which allows you to create Word documents using JavaScript/TypeScript.
    
    ### Workflow
    1. **MANDATORY - READ ENTIRE FILE**: Read [`docx-js.md`](docx-js.md) (~500 lines) completely from start to finish. **NEVER set any range limits when reading this file.** Read the full file content for detailed syntax, critical formatting rules, and best practices before proceeding with document creation.
    2. Create a JavaScript/TypeScript file using Document, Paragraph, TextRun components (You can assume all dependencies are installed, but if not, refer to the dependencies section below)
    3. Export as .docx using Packer.toBuffer()
    
    ## Editing an existing Word document
    
    When editing an existing Word document, use the **Document library** (a Python library for OOXML manipulation). The library automatically handles infrastructure setup and provides methods for document manipulation. For complex scenarios, you can access the underlying DOM directly through the library.
    
    ### Workflow
    1. **MANDATORY - READ ENTIRE FILE**: Read [`ooxml.md`](ooxml.md) (~600 lines) completely from start to finish. **NEVER set any range limits when reading this file.** Read the full file content for the Document library API and XML patterns for directly editing document files.
    2. Unpack the document: `python ooxml/scripts/unpack.py <office_file> <output_directory>`
    3. Create and run a Python script using the Document library (see "Document Library" section in ooxml.md)
    4. Pack the final document: `python ooxml/scripts/pack.py <input_directory> <office_file>`
    
    The Document library provides both high-level methods for common operations and direct DOM access for complex scenarios.
    
    ## Redlining workflow for document review
    
    This workflow allows you to plan comprehensive tracked changes using markdown before implementing them in OOXML. **CRITICAL**: For complete tracked changes, you must implement ALL changes systematically.
    
    **Batching Strategy**: Group related changes into batches of 3-10 changes. This makes debugging manageable while maintaining efficiency. Test each batch before moving to the next.
    
    **Principle: Minimal, Precise Edits**
    When implementing tracked changes, only mark text that actually changes. Repeating unchanged text makes edits harder to review and appears unprofessional. Break replacements into: [unchanged text] + [deletion] + [insertion] + [unchanged text]. Preserve the original run's RSID for unchanged text by extracting the `<w:r>` element from the original and reusing it.
    
    Example - Changing "30 days" to "60 days" in a sentence:
    ```python
    # BAD - Replaces entire sentence
    '<w:del><w:r><w:delText>The term is 30 days.</w:delText></w:r></w:del><w:ins><w:r><w:t>The term is 60 days.</w:t></w:r></w:ins>'
    
    # GOOD - Only marks what changed, preserves original <w:r> for unchanged text
    '<w:r w:rsidR="00AB12CD"><w:t>The term is </w:t></w:r><w:del><w:r><w:delText>30</w:delText></w:r></w:del><w:ins><w:r><w:t>60</w:t></w:r></w:ins><w:r w:rsidR="00AB12CD"><w:t> days.</w:t></w:r>'
    ```
    
    ### Tracked changes workflow
    
    1. **Get markdown representation**: Convert document to markdown with tracked changes preserved:
       ```bash
       pandoc --track-changes=all path-to-file.docx -o current.md
       ```
    
    2. **Identify and group changes**: Review the document and identify ALL changes needed, organizing them into logical batches:
    
       **Location methods** (for finding changes in XML):
       - Section/heading numbers (e.g., "Section 3.2", "Article IV")
       - Paragraph identifiers if numbered
       - Grep patterns with unique surrounding text
       - Document structure (e.g., "first paragraph", "signature block")
       - **DO NOT use markdown line numbers** - they don't map to XML structure
    
       **Batch organization** (group 3-10 related changes per batch):
       - By section: "Batch 1: Section 2 amendments", "Batch 2: Section 5 updates"
       - By type: "Batch 1: Date corrections", "Batch 2: Party name changes"
       - By complexity: Start with simple text replacements, then tackle complex structural changes
       - Sequential: "Batch 1: Pages 1-3", "Batch 2: Pages 4-6"
    
    3. **Read documentation and unpack**:
       - **MANDATORY - READ ENTIRE FILE**: Read [`ooxml.md`](ooxml.md) (~600 lines) completely from start to finish. **NEVER set any range limits when reading this file.** Pay special attention to the "Document Library" and "Tracked Change Patterns" sections.
       - **Unpack the document**: `python ooxml/scripts/unpack.py <file.docx> <dir>`
       - **Note the suggested RSID**: The unpack script will suggest an RSID to use for your tracked changes. Copy this RSID for use in step 4b.
    
    4. **Implement changes in batches**: Group changes logically (by section, by type, or by proximity) and implement them together in a single script. This approach:
       - Makes debugging easier (smaller batch = easier to isolate errors)
       - Allows incremental progress
       - Maintains efficiency (batch size of 3-10 changes works well)
    
       **Suggested batch groupings:**
       - By document section (e.g., "Section 3 changes", "Definitions", "Termination clause")
       - By change type (e.g., "Date changes", "Party name updates", "Legal term replacements")
       - By proximity (e.g., "Changes on pages 1-3", "Changes in first half of document")
    
       For each batch of related changes:
    
       **a. Map text to XML**: Grep for text in `word/document.xml` to verify how text is split across `<w:r>` elements.
    
       **b. Create and run script**: Use `get_node` to find nodes, implement changes, then `doc.save()`. See **"Document Library"** section in ooxml.md for patterns.
    
       **Note**: Always grep `word/document.xml` immediately before writing a script to get current line numbers and verify text content. Line numbers change after each script run.
    
    5. **Pack the document**: After all batches are complete, convert the unpacked directory back to .docx:
       ```bash
       python ooxml/scripts/pack.py unpacked reviewed-document.docx
       ```
    
    6. **Final verification**: Do a comprehensive check of the complete document:
       - Convert final document to markdown:
         ```bash
         pandoc --track-changes=all reviewed-document.docx -o verification.md
         ```
       - Verify ALL changes were applied correctly:
         ```bash
         grep "original phrase" verification.md  # Should NOT find it
         grep "replacement phrase" verification.md  # Should find it
         ```
       - Check that no unintended changes were introduced
    
    
    ## Converting Documents to Images
    
    To visually analyze Word documents, convert them to images using a two-step process:
    
    1. **Convert DOCX to PDF**:
       ```bash
       soffice --headless --convert-to pdf document.docx
       ```
    
    2. **Convert PDF pages to JPEG images**:
       ```bash
       pdftoppm -jpeg -r 150 document.pdf page
       ```
       This creates files like `page-1.jpg`, `page-2.jpg`, etc.
    
    Options:
    - `-r 150`: Sets resolution to 150 DPI (adjust for quality/size balance)
    - `-jpeg`: Output JPEG format (use `-png` for PNG if preferred)
    - `-f N`: First page to convert (e.g., `-f 2` starts from page 2)
    - `-l N`: Last page to convert (e.g., `-l 5` stops at page 5)
    - `page`: Prefix for output files
    
    Example for specific range:
    ```bash
    pdftoppm -jpeg -r 150 -f 2 -l 5 document.pdf page  # Converts only pages 2-5
    ```
    
    ## Code Style Guidelines
    **IMPORTANT**: When generating code for DOCX operations:
    - Write concise code
    - Avoid verbose variable names and redundant operations
    - Avoid unnecessary print statements
    
    ## Dependencies
    
    Required dependencies (install if not available):
    
    - **pandoc**: `sudo apt-get install pandoc` (for text extraction)
    - **docx**: `npm install -g docx` (for creating new documents)
    - **LibreOffice**: `sudo apt-get install libreoffice` (for PDF conversion)
    - **Poppler**: `sudo apt-get install poppler-utils` (for pdftoppm to convert PDF to images)
    - **defusedxml**: `pip install defusedxml` (for secure XML parsing)

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related