Claude Skill

pdf

Imported from paulrberg/agent-skills/skills/pdf.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download paulrberg-agent-skills-skills_pdf-913232a.zip · 12 KB
Part of paulrberg/agent-skills — 42 skills

Install

skills CLI npx skills add https://github.com/PaulRBerg/agent-skills/tree/main/skills/pdf
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install paulrberg-agent-skills@llmmart
Git git clone https://github.com/PaulRBerg/agent-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole paulrberg/agent-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

PDF

Process PDFs locally on macOS with exact extraction, source preservation, deliberate tool routing, and structural plus semantic validation.

Invariants

  1. Run extraction and transformations locally. Task-relevant document evidence in tool output and internal agent reports may be processed by the configured model provider. Require explicit user authorization and an external-disclosure review before uploading or sending document contents outside that agent workflow. Package and language-data downloads do not authorize document disclosure.
  2. Preserve every original PDF byte-for-byte. Write a sibling output, copy, or explicitly named destination unless the user authorizes destructive replacement.
  3. Preserve monetary values, identifiers, dates, signs, and displayed precision as strings. Use decimal.Decimal for arithmetic; never infer missing rows or silently discard headers, footnotes, continuation lines, or boundary pages.
  4. Inspect structure and representative renders before choosing a transformation. Use the smallest tool that preserves the required layout, forms, annotations, and image quality.
  5. Validate every written PDF structurally and against task semantics. A command exiting successfully is not evidence that extracted rows, totals, page boundaries, form appearances, or visual layout are correct.
  6. Keep reports concise for private financial, tax, legal, and health documents. Prefer counts, reconciliations, and file references over raw sensitive rows unless the rows materially support the task or the user asks for them.

Profile First

Resolve the skill directory from this SKILL.md, then profile every unknown input:

uv run "<skill-dir>/scripts/profile.py" "<input.pdf>"

The helper emits schema-versioned JSON with integrity, encryption, page geometry/rotation, image counts, and per-page text coverage without document text. Stop on password_required; password handling is outside this skill.

When layout, cropping, OCR quality, signatures, or form placement matters, render the first and last page, every structural boundary, and any page behind a discrepancy. For dense charts, tables, or technical drawings, render at higher resolution and crop or zoom the relevant region before reading values.

Route by Evidence

Need Preferred route
Quick reading or page-aware extraction Host PDF reader when available, then pdftotext -layout
Coordinates, columns, or difficult tables Poppler bounding boxes, then pdfplumber through uv run
Image-only or materially incomplete text OCRmyPDF with Tesseract; default languages eng+ron
Merge, split, rotate, or integrity checks qpdf
Render pages or extract embedded images pdftocairo or pdfimages
Convert ordered images into a PDF img2pdf
Reduce size qpdf lossless rewrite first; Ghostscript only for an accepted lossy pass
Inspect, fill, flatten, or overlay forms Read references/forms.md first

Read references/recipes.md only when exact commands for the selected extraction, transformation, OCR, image, comparison, or compression branch are needed.

Execute and Reconcile

  1. Profile inputs and identify whether each page is digital, scanned, mixed, rotated, or image-heavy.
  2. Extract or transform into a new path. For tabular documents, retain page provenance and parse continuations across page breaks before assigning rows.
  3. Reconcile financial and evidentiary output with every available invariant: page and row counts, opening/closing balances, inflows/outflows, subtotals, displayed totals, date coverage, and source hashes when provenance matters.
  4. For comparisons, extract both sources independently, enumerate overlapping and unique facts, and render the pages behind every material disagreement. Distinguish a real discrepancy from an extraction failure.
  5. For split or rename work, establish an old-to-new map from stable content identifiers. Copy by default, preserve contextual boundary pages when needed, and verify the first and last page of every result.
  6. Validate outputs with qpdf, expected page count/dimensions, text coverage, representative renders, and the task's semantic invariants. Retain OCR sidecars or extraction intermediates only when they are requested or useful evidence.

Completion requires preserved originals, intentional outputs, successful structural checks, semantic reconciliation, and a concise report of paths and evidence. Lead read-only reports with ### 📄 PDF — 🔎 inspected, no files written; use ### 📄 PDF — ✅ updated only after all required validation passes, and ### 📄 PDF — ⛔ not deliverable when a required check fails.

Files (agent-skills)
  • agents
    • openai.yaml 42 B
      policy:
        allow_implicit_invocation: true
      
  • references
    • forms.md 2.6 KB
      # PDF Forms
      
      Choose the branch from the document structure; do not guess from appearance.
      
      ## Inspect
      
      ```sh
      uv run "<skill-dir>/scripts/form.py" inspect "input.pdf"
      ```
      
      The JSON reports whether an AcroForm exists, whether XFA is present, and each field's fully qualified name, type,
      current and default values, options, flags, page, and rectangle.
      
      Route by result:
      
      - AcroForm without XFA or signatures: fill known fields.
      - Flat page with no AcroForm: use coordinate overlays after rendering and calibration.
      - XFA or signature fields: stop. This helper intentionally does not preserve or generate those workflows.
      
      ## Fill an AcroForm
      
      Create a JSON object keyed by the exact field names returned by `inspect`:
      
      ```json
      {
        "person.name": "Ada Lovelace",
        "preferences.language": "Romanian",
        "topics": ["tax", "finance"]
      }
      ```
      
      Fill to a new output:
      
      ```sh
      uv run "<skill-dir>/scripts/form.py" fill "input.pdf" "values.json" "filled.pdf"
      ```
      
      Use `--flatten` only for a final, non-editable delivery. Flattening draws field appearances, removes widget annotations,
      and removes the AcroForm dictionary. Verify a rendered output because stored values alone do not prove visibility.
      
      The helper rejects unknown fields, XFA, signatures, input/output aliasing, and an existing output unless `--force` is
      explicitly supplied.
      
      ## Overlay a flat form
      
      Render the page first and calibrate in PDF points from the lower-left corner. Page numbers are one-based. Create a JSON
      array:
      
      ```json
      [
        { "page": 1, "x": 126, "y": 618, "text": "Ada Lovelace", "font_size": 10 },
        { "page": 2, "x": 90, "y": 144, "text": "București" }
      ]
      ```
      
      Apply it to a new output:
      
      ```sh
      uv run "<skill-dir>/scripts/form.py" overlay "input.pdf" "placements.json" "filled-flat.pdf"
      ```
      
      The default font is `/System/Library/Fonts/Supplemental/Arial.ttf`, which supports Romanian text. `font_size` defaults
      to 10. Placements are single-line text; add separate entries instead of relying on wrapping.
      
      Render every affected page after overlaying. Check baseline, clipping, diacritics, rotation, and whether the visual
      position still matches at ordinary and high zoom. The helper preserves page count and dimensions but cannot determine
      whether coordinates are semantically correct.
      
      ## Validate
      
      For either branch:
      
      1. Require the helper's success JSON and qpdf integrity.
      2. Re-run `inspect` for editable AcroForms and compare intended values.
      3. Render every affected page and visually inspect field appearances or overlay placement.
      4. Confirm page count and dimensions match the input.
      5. Keep the untouched input alongside the validated output.
      
    • recipes.md 4.7 KB
      # PDF Recipes
      
      Load only the branch needed for the current task. Keep every output distinct from its input and quote paths.
      
      ## Inspect and extract
      
      Start with the bundled factual profile, then inspect raw tool output only as needed:
      
      ```sh
      uv run "<skill-dir>/scripts/profile.py" "input.pdf"
      pdfinfo -box "input.pdf"
      qpdf --check "input.pdf"
      pdfimages -list "input.pdf"
      ```
      
      Preserve reading order where possible:
      
      ```sh
      pdftotext -layout "input.pdf" "output.txt"
      pdftotext -f 3 -l 7 -layout "input.pdf" "pages-3-7.txt"
      pdftotext -f 1 -l 1 -x 36 -y 72 -W 540 -H 648 -layout "input.pdf" "crop.txt"
      pdftotext -bbox-layout "input.pdf" "layout.html"
      ```
      
      Use bounding boxes to diagnose column interleaving or logical half-pages. Escalate to `pdfplumber` only when the Poppler
      output cannot preserve the required table structure:
      
      ```sh
      uv run --with 'pdfplumber>=0.11.10,<0.12' python - "input.pdf" <<'PY'
      import json
      import sys
      
      import pdfplumber
      
      with pdfplumber.open(sys.argv[1]) as document:
          print(json.dumps([page.extract_tables() for page in document.pages], ensure_ascii=False))
      PY
      ```
      
      Treat extracted tables as candidates, not truth. Rebuild wrapped descriptions and continuations, retain page numbers,
      and reconcile exact row counts and totals against the PDF.
      
      ## Compare documents
      
      Extract each input independently with the same appropriate route. Compare normalized facts while retaining the original
      strings and page provenance. Report:
      
      - facts present in both documents;
      - facts unique to each document;
      - materially different values, dates, identifiers, qualifications, or footnotes;
      - pages rendered to distinguish source differences from extraction errors.
      
      Do not infer that missing extracted text means missing source content. Inspect the relevant render or OCR coverage
      first.
      
      ## OCR scans
      
      Use OCR only for image-only or materially incomplete pages. The default covers English and Romanian:
      
      ```sh
      ocrmypdf --output-type pdf --skip-text -l eng+ron --sidecar "ocr.txt" "input.pdf" "ocr.pdf"
      qpdf --check "ocr.pdf"
      pdftotext -layout "ocr.pdf" "ocr-check.txt"
      ```
      
      Add `--rotate-pages` or `--deskew` only when profiling or rendered pages show the need. Add `--clean` only after
      accepting that image processing may alter visual evidence. Re-render affected pages after any of these options.
      
      ## Render and extract images
      
      Render pages for visual comparison:
      
      ```sh
      pdftocairo -png -r 200 "input.pdf" "page"
      pdftocairo -f 4 -l 4 -png -r 300 "input.pdf" "page-4"
      ```
      
      Extract embedded images without rasterizing whole pages:
      
      ```sh
      pdfimages -list "input.pdf"
      pdfimages -all "input.pdf" "image"
      ```
      
      Combine ordered images without recompressing them unnecessarily:
      
      ```sh
      img2pdf "page-01.png" "page-02.jpg" --output "combined.pdf"
      qpdf --check "combined.pdf"
      ```
      
      Confirm image order, orientation, page dimensions, and representative renders.
      
      ## Merge, split, and rotate
      
      Merge in explicit order:
      
      ```sh
      qpdf --empty --pages "part-1.pdf" "part-2.pdf" -- "merged.pdf"
      ```
      
      Extract a range or retain a boundary page for context:
      
      ```sh
      qpdf "input.pdf" --pages . 1-5 -- "part-1.pdf"
      qpdf "input.pdf" --pages . 5-10 -- "part-2-with-boundary.pdf"
      ```
      
      Rotate selected pages clockwise:
      
      ```sh
      qpdf "input.pdf" "rotated.pdf" --rotate=+90:1,3
      ```
      
      Check every result and inspect its first and last page:
      
      ```sh
      qpdf --check "output.pdf"
      pdfinfo "output.pdf"
      pdftotext -f 1 -l 1 -layout "output.pdf" -
      ```
      
      ## Rename a corpus
      
      Build a complete old-to-new map before writing. Derive names from stable content such as issuer, document type, account
      suffix, and covered date range. Detect collisions and ambiguous documents first. Copy to the new names unless the user
      explicitly authorizes renaming originals, then profile both sides and compare hashes when a byte-identical copy is
      expected.
      
      ## Compress
      
      Try a lossless structural rewrite first:
      
      ```sh
      qpdf --object-streams=generate --recompress-flate --compression-level=9 "input.pdf" "lossless.pdf"
      ```
      
      Use Ghostscript only when a smaller lossy output is acceptable:
      
      ```sh
      gs -sDEVICE=pdfwrite -dCompatibilityLevel=1.7 -dPDFSETTINGS=/ebook -dNOPAUSE -dQUIET -dBATCH \
        -sOutputFile="compressed.pdf" "input.pdf"
      ```
      
      Compare byte size only after qpdf integrity, page count/dimensions, text extraction, forms/annotations where relevant,
      and representative renders pass. Keep the original and the smallest acceptable validated output.
      
      ## Final validation
      
      At minimum, require:
      
      ```sh
      qpdf --check "output.pdf"
      pdfinfo -box "output.pdf"
      pdftotext -layout "output.pdf" "output.txt"
      ```
      
      Add domain checks: exact totals and balances for statements, field values and appearances for forms, first/last pages
      for splits, order and dimensions for image conversions, and visual comparison for OCR or compression.
      
  • scripts
    • form.py 14.1 KB
      #!/usr/bin/env -S uv run --script
      # /// script
      # requires-python = ">=3.12"
      # dependencies = ["pypdf>=6.15,<7", "reportlab>=5,<6"]
      # ///
      """Inspect, fill, flatten, and overlay local PDF forms."""
      
      from __future__ import annotations
      
      import argparse
      import io
      import json
      import os
      import shutil
      import subprocess
      import tempfile
      from collections import defaultdict
      from pathlib import Path
      from typing import Any
      
      from pypdf import PdfReader, PdfWriter
      from pypdf.errors import PdfReadError
      from pypdf.generic import ArrayObject, NameObject
      from reportlab.pdfbase import pdfmetrics
      from reportlab.pdfbase.ttfonts import TTFont
      from reportlab.pdfgen import canvas
      
      SCHEMA_VERSION = 1
      DEFAULT_FONT = Path("/System/Library/Fonts/Supplemental/Arial.ttf")
      FONT_NAME = "PDFSkillArial"
      
      
      class UserError(Exception):
          def __init__(self, code: str, message: str) -> None:
              super().__init__(message)
              self.code = code
      
      
      def emit(payload: dict[str, Any], returncode: int = 0) -> int:
          print(json.dumps(payload, ensure_ascii=False, indent=2, sort_keys=True))
          return returncode
      
      
      def pdf_object(value: Any) -> Any:
          if hasattr(value, "get_object"):
              value = value.get_object()
          if value is None or isinstance(value, (bool, int, float, str)):
              return value
          if isinstance(value, dict):
              return {str(key): pdf_object(item) for key, item in value.items()}
          if isinstance(value, (list, tuple)):
              return [pdf_object(item) for item in value]
          return str(value)
      
      
      def resolved_dictionary(value: Any) -> Any:
          return value.get_object() if hasattr(value, "get_object") else value
      
      
      def acroform(reader: PdfReader) -> Any | None:
          root = resolved_dictionary(reader.trailer["/Root"])
          value = root.get("/AcroForm")
          return resolved_dictionary(value) if value is not None else None
      
      
      def ensure_readable(reader: PdfReader) -> None:
          if reader.is_encrypted:
              raise UserError("password_required", "encrypted PDFs are not supported")
      
      
      def qualified_field_name(annotation: Any) -> str | None:
          parts: list[str] = []
          node = resolved_dictionary(annotation)
          seen: set[int] = set()
          while node is not None and id(node) not in seen:
              seen.add(id(node))
              if node.get("/T") is not None:
                  parts.append(str(node["/T"]))
              parent = node.get("/Parent")
              node = resolved_dictionary(parent) if parent is not None else None
          return ".".join(reversed(parts)) or None
      
      
      def widget_locations(reader: PdfReader) -> dict[str, list[dict[str, Any]]]:
          locations: dict[str, list[dict[str, Any]]] = defaultdict(list)
          for page_number, page in enumerate(reader.pages, 1):
              for reference in page.get("/Annots", []):
                  annotation = resolved_dictionary(reference)
                  if str(annotation.get("/Subtype")) != "/Widget":
                      continue
                  name = qualified_field_name(annotation)
                  if name is None:
                      continue
                  rectangle = annotation.get("/Rect")
                  locations[name].append(
                      {
                          "page": page_number,
                          "rectangle": [float(value) for value in rectangle] if rectangle else None,
                      }
                  )
          return locations
      
      
      def inspect_payload(input_path: Path) -> dict[str, Any]:
          reader = PdfReader(str(input_path))
          ensure_readable(reader)
          form = acroform(reader)
          fields = reader.get_fields() or {}
          locations = widget_locations(reader)
          items: list[dict[str, Any]] = []
          for name in sorted(fields):
              field = resolved_dictionary(fields[name])
              location = locations.get(name, [{}])[0]
              items.append(
                  {
                      "name": name,
                      "type": str(field.get("/FT")) if field.get("/FT") is not None else None,
                      "value": pdf_object(field.get("/V")),
                      "default_value": pdf_object(field.get("/DV")),
                      "options": pdf_object(field.get("/Opt")),
                      "flags": int(field.get("/Ff", 0)),
                      "page": location.get("page"),
                      "rectangle": location.get("rectangle"),
                  }
              )
          return {
              "schema_version": SCHEMA_VERSION,
              "status": "ok",
              "operation": "inspect",
              "input": str(input_path),
              "page_count": len(reader.pages),
              "acroform": form is not None,
              "xfa": bool(form is not None and form.get("/XFA") is not None),
              "fields": items,
          }
      
      
      def load_json(path: Path) -> Any:
          try:
              return json.loads(path.read_text(encoding="utf-8"))
          except FileNotFoundError as error:
              raise UserError("missing_json", "JSON input does not exist") from error
          except json.JSONDecodeError as error:
              raise UserError("invalid_json", f"invalid JSON at line {error.lineno}") from error
      
      
      def prepare_paths(input_value: str, output_value: str, force: bool) -> tuple[Path, Path]:
          input_path = Path(input_value).expanduser().resolve()
          output_path = Path(output_value).expanduser().resolve()
          if not input_path.is_file():
              raise UserError("missing_input", "input PDF does not exist")
          if input_path == output_path:
              raise UserError("input_output_alias", "input and output must be different paths")
          if output_path.exists() and not force:
              raise UserError("output_exists", "output exists; pass --force to replace it")
          if not output_path.parent.is_dir():
              raise UserError("missing_output_directory", "output directory does not exist")
          return input_path, output_path
      
      
      def validate_writable_form(reader: PdfReader) -> dict[str, Any]:
          ensure_readable(reader)
          form = acroform(reader)
          fields = reader.get_fields() or {}
          if form is None or not fields:
              raise UserError("no_acroform", "PDF has no fillable AcroForm fields")
          if form.get("/XFA") is not None:
              raise UserError("unsupported_xfa", "XFA forms are not supported")
          if any(str(resolved_dictionary(field).get("/FT")) == "/Sig" for field in fields.values()):
              raise UserError("unsupported_signature", "signature fields are not supported")
          return fields
      
      
      def atomic_write(writer: PdfWriter, output_path: Path) -> None:
          if shutil.which("qpdf") is None:
              raise UserError("missing_qpdf", "qpdf is required to validate written PDFs")
          descriptor, temporary_name = tempfile.mkstemp(
              prefix=f".{output_path.name}.", suffix=".tmp.pdf", dir=output_path.parent
          )
          os.close(descriptor)
          temporary_path = Path(temporary_name)
          try:
              writer.write(str(temporary_path))
              check = subprocess.run(
                  ["qpdf", "--check", str(temporary_path)],
                  capture_output=True,
                  check=False,
                  text=True,
              )
              if check.returncode not in (0, 3):
                  raise UserError("integrity_failed", "qpdf rejected the generated PDF")
              os.replace(temporary_path, output_path)
          finally:
              temporary_path.unlink(missing_ok=True)
      
      
      def command_inspect(args: argparse.Namespace) -> int:
          input_path = Path(args.input).expanduser().resolve()
          if not input_path.is_file():
              raise UserError("missing_input", "input PDF does not exist")
          return emit(inspect_payload(input_path))
      
      
      def command_fill(args: argparse.Namespace) -> int:
          input_path, output_path = prepare_paths(args.input, args.output, args.force)
          values_path = Path(args.values).expanduser().resolve()
          values = load_json(values_path)
          if not isinstance(values, dict) or not all(isinstance(key, str) for key in values):
              raise UserError("invalid_values", "values JSON must be an object keyed by field name")
          for value in values.values():
              if isinstance(value, str):
                  continue
              if isinstance(value, list) and all(isinstance(item, str) for item in value):
                  continue
              raise UserError("invalid_values", "field values must be strings or lists of strings")
      
          reader = PdfReader(str(input_path))
          fields = validate_writable_form(reader)
          unknown = sorted(set(values) - set(fields))
          if unknown:
              raise UserError("unknown_fields", f"unknown field names: {', '.join(unknown)}")
      
          writer = PdfWriter(clone_from=str(input_path))
          for page in writer.pages:
              writer.update_page_form_field_values(
                  page,
                  values,
                  auto_regenerate=False,
                  flatten=args.flatten,
              )
          if args.flatten:
              for page in writer.pages:
                  annotations = page.get("/Annots")
                  if annotations is None:
                      continue
                  retained = ArrayObject(
                      [
                          reference
                          for reference in annotations
                          if str(resolved_dictionary(reference).get("/Subtype")) != "/Widget"
                      ]
                  )
                  if retained:
                      page[NameObject("/Annots")] = retained
                  else:
                      page.pop(NameObject("/Annots"), None)
              writer.root_object.pop(NameObject("/AcroForm"), None)
          atomic_write(writer, output_path)
          return emit(
              {
                  "schema_version": SCHEMA_VERSION,
                  "status": "ok",
                  "operation": "fill",
                  "output": str(output_path),
                  "fields_written": sorted(values),
                  "flattened": args.flatten,
              }
          )
      
      
      def positive_number(value: Any, name: str) -> float:
          if isinstance(value, bool) or not isinstance(value, (int, float)):
              raise UserError("invalid_placements", f"{name} must be a number")
          result = float(value)
          if name == "font_size" and result <= 0:
              raise UserError("invalid_placements", "font_size must be greater than zero")
          return result
      
      
      def command_overlay(args: argparse.Namespace) -> int:
          input_path, output_path = prepare_paths(args.input, args.output, args.force)
          placements = load_json(Path(args.placements).expanduser().resolve())
          if not isinstance(placements, list):
              raise UserError("invalid_placements", "placements JSON must be an array")
          reader = PdfReader(str(input_path))
          ensure_readable(reader)
          grouped: dict[int, list[dict[str, Any]]] = defaultdict(list)
          for index, placement in enumerate(placements):
              if not isinstance(placement, dict):
                  raise UserError("invalid_placements", f"placement {index} must be an object")
              page = placement.get("page")
              text = placement.get("text")
              if isinstance(page, bool) or not isinstance(page, int) or not 1 <= page <= len(reader.pages):
                  raise UserError("invalid_placements", f"placement {index} has an invalid page")
              if not isinstance(text, str) or "\n" in text or "\r" in text:
                  raise UserError("invalid_placements", f"placement {index} text must be one line")
              grouped[page].append(
                  {
                      "x": positive_number(placement.get("x"), "x"),
                      "y": positive_number(placement.get("y"), "y"),
                      "text": text,
                      "font_size": positive_number(placement.get("font_size", 10), "font_size"),
                  }
              )
      
          if not DEFAULT_FONT.is_file():
              raise UserError("missing_font", f"default font not found: {DEFAULT_FONT}")
          pdfmetrics.registerFont(TTFont(FONT_NAME, str(DEFAULT_FONT)))
          writer = PdfWriter(clone_from=str(input_path))
          input_dimensions = [
              (float(page.mediabox.width), float(page.mediabox.height)) for page in reader.pages
          ]
          for page_number, page_placements in grouped.items():
              width, height = input_dimensions[page_number - 1]
              buffer = io.BytesIO()
              overlay_canvas = canvas.Canvas(buffer, pagesize=(width, height))
              for placement in page_placements:
                  overlay_canvas.setFont(FONT_NAME, placement["font_size"])
                  overlay_canvas.drawString(placement["x"], placement["y"], placement["text"])
              overlay_canvas.save()
              buffer.seek(0)
              overlay_page = PdfReader(buffer).pages[0]
              writer.pages[page_number - 1].merge_page(overlay_page)
      
          output_dimensions = [
              (float(page.mediabox.width), float(page.mediabox.height)) for page in writer.pages
          ]
          if output_dimensions != input_dimensions:
              raise UserError("geometry_changed", "overlay changed page dimensions")
          atomic_write(writer, output_path)
          return emit(
              {
                  "schema_version": SCHEMA_VERSION,
                  "status": "ok",
                  "operation": "overlay",
                  "output": str(output_path),
                  "placements_written": len(placements),
                  "pages_affected": sorted(grouped),
              }
          )
      
      
      def parser() -> argparse.ArgumentParser:
          result = argparse.ArgumentParser(description=__doc__)
          commands = result.add_subparsers(dest="command", required=True)
      
          inspect_parser = commands.add_parser("inspect")
          inspect_parser.add_argument("input")
          inspect_parser.set_defaults(handler=command_inspect)
      
          fill_parser = commands.add_parser("fill")
          fill_parser.add_argument("input")
          fill_parser.add_argument("values")
          fill_parser.add_argument("output")
          fill_parser.add_argument("--flatten", action="store_true")
          fill_parser.add_argument("--force", action="store_true")
          fill_parser.set_defaults(handler=command_fill)
      
          overlay_parser = commands.add_parser("overlay")
          overlay_parser.add_argument("input")
          overlay_parser.add_argument("placements")
          overlay_parser.add_argument("output")
          overlay_parser.add_argument("--force", action="store_true")
          overlay_parser.set_defaults(handler=command_overlay)
          return result
      
      
      def main() -> int:
          args = parser().parse_args()
          try:
              return args.handler(args)
          except UserError as error:
              return emit(
                  {
                      "schema_version": SCHEMA_VERSION,
                      "status": "error",
                      "error": {"code": error.code, "message": str(error)},
                  },
                  1,
              )
          except (KeyError, OSError, PdfReadError, TypeError, ValueError) as error:
              return emit(
                  {
                      "schema_version": SCHEMA_VERSION,
                      "status": "error",
                      "error": {"code": "unexpected_error", "message": str(error)},
                  },
                  1,
              )
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
    • profile.py 6.5 KB
      #!/usr/bin/env -S uv run --script
      # /// script
      # requires-python = ">=3.12"
      # dependencies = []
      # ///
      """Emit a compact factual profile for a local PDF."""
      
      from __future__ import annotations
      
      import json
      import re
      import shutil
      import subprocess
      import sys
      from collections import Counter
      from pathlib import Path
      from typing import Any
      
      REQUIRED_TOOLS = ("pdfimages", "pdfinfo", "pdftotext", "qpdf")
      
      
      def emit(payload: dict[str, Any], returncode: int) -> int:
          print(json.dumps(payload, ensure_ascii=False, indent=2, sort_keys=True))
          return returncode
      
      
      def fail(code: str, message: str, pdf_path: Path | None = None) -> int:
          payload: dict[str, Any] = {
              "schema_version": 1,
              "status": "error",
              "error": {"code": code, "message": message},
          }
          if pdf_path is not None:
              payload["path"] = str(pdf_path)
          return emit(payload, 1)
      
      
      def run(args: list[str]) -> subprocess.CompletedProcess[str]:
          return subprocess.run(
              args,
              capture_output=True,
              check=False,
              encoding="utf-8",
              errors="replace",
              text=True,
          )
      
      
      def parse_pdfinfo(output: str) -> dict[str, str]:
          result: dict[str, str] = {}
          for line in output.splitlines():
              if ":" not in line:
                  continue
              key, value = line.split(":", 1)
              result[key.strip()] = value.strip()
          return result
      
      
      def parse_geometry(output: str, page_count: int, fallback: dict[str, str]) -> list[dict[str, Any]]:
          pages: dict[int, dict[str, Any]] = {
              page: {"page": page, "width_points": None, "height_points": None, "rotation": 0}
              for page in range(1, page_count + 1)
          }
          size_pattern = re.compile(
              r"^Page\s+(\d+)\s+size:\s+([0-9.]+)\s+x\s+([0-9.]+)\s+pts",
              re.MULTILINE,
          )
          rotation_pattern = re.compile(r"^Page\s+(\d+)\s+rot:\s+(-?\d+)", re.MULTILINE)
          for match in size_pattern.finditer(output):
              page = int(match.group(1))
              if page in pages:
                  pages[page]["width_points"] = float(match.group(2))
                  pages[page]["height_points"] = float(match.group(3))
          for match in rotation_pattern.finditer(output):
              page = int(match.group(1))
              if page in pages:
                  pages[page]["rotation"] = int(match.group(2))
      
          fallback_size = re.search(r"([0-9.]+)\s+x\s+([0-9.]+)\s+pts", fallback.get("Page size", ""))
          fallback_rotation = int(fallback.get("Page rot", "0") or 0)
          if fallback_size:
              for page in pages.values():
                  if page["width_points"] is None:
                      page["width_points"] = float(fallback_size.group(1))
                      page["height_points"] = float(fallback_size.group(2))
                      page["rotation"] = fallback_rotation
          return list(pages.values())
      
      
      def text_coverage(output: str, page_count: int) -> dict[str, Any]:
          chunks = output.split("\f")
          if len(chunks) > page_count and not chunks[-1].strip():
              chunks.pop()
          chunks.extend([""] * max(0, page_count - len(chunks)))
          counts = [sum(not character.isspace() for character in chunk) for chunk in chunks[:page_count]]
          return {
              "characters": sum(counts),
              "pages_with_text": [index for index, count in enumerate(counts, 1) if count > 0],
              "pages_without_text": [index for index, count in enumerate(counts, 1) if count == 0],
              "per_page": [
                  {"page": index, "characters": count} for index, count in enumerate(counts, 1)
              ],
          }
      
      
      def image_coverage(output: str, page_count: int) -> dict[str, Any]:
          counts: Counter[int] = Counter()
          for line in output.splitlines():
              match = re.match(r"^\s*(\d+)\s+\d+\s+", line)
              if match:
                  counts[int(match.group(1))] += 1
          return {
              "total": sum(counts.values()),
              "per_page": [
                  {"page": page, "count": counts.get(page, 0)} for page in range(1, page_count + 1)
              ],
          }
      
      
      def main() -> int:
          if len(sys.argv) != 2:
              return fail("usage", "usage: profile.py <input.pdf>")
      
          pdf_path = Path(sys.argv[1]).expanduser().resolve()
          if not pdf_path.is_file():
              return fail("missing_file", "input is not a readable file", pdf_path)
      
          missing = [tool for tool in REQUIRED_TOOLS if shutil.which(tool) is None]
          if missing:
              return fail("missing_tools", f"required tools not found: {', '.join(missing)}", pdf_path)
      
          basic = run(["pdfinfo", str(pdf_path)])
          if basic.returncode != 0:
              return fail("unreadable_pdf", "pdfinfo could not read the input", pdf_path)
          info = parse_pdfinfo(basic.stdout)
          try:
              page_count = int(info.get("Pages", "0"))
          except ValueError:
              return fail("invalid_metadata", "pdfinfo returned an invalid page count", pdf_path)
          if page_count < 1:
              return fail("invalid_metadata", "PDF has no readable pages", pdf_path)
      
          encrypted = info.get("Encrypted", "unknown").lower().startswith("yes")
          base: dict[str, Any] = {
              "schema_version": 1,
              "path": str(pdf_path),
              "size_bytes": pdf_path.stat().st_size,
              "pdf_version": info.get("PDF version"),
              "encrypted": encrypted,
              "page_count": page_count,
          }
          if encrypted:
              base.update({"status": "password_required", "integrity": {"status": "not_checked"}})
              return emit(base, 1)
      
          integrity = run(["qpdf", "--check", str(pdf_path)])
          if integrity.returncode not in (0, 3):
              return fail("integrity_failed", "qpdf rejected the input", pdf_path)
      
          detailed = run(["pdfinfo", "-f", "1", "-l", str(page_count), "-box", str(pdf_path)])
          if detailed.returncode != 0:
              return fail("metadata_failed", "pdfinfo could not inspect page geometry", pdf_path)
          text_result = run(["pdftotext", "-layout", str(pdf_path), "-"])
          if text_result.returncode != 0:
              return fail("text_failed", "pdftotext could not inspect text coverage", pdf_path)
          images_result = run(["pdfimages", "-list", str(pdf_path)])
          if images_result.returncode != 0:
              return fail("images_failed", "pdfimages could not inspect image coverage", pdf_path)
      
          base.update(
              {
                  "status": "warnings" if integrity.returncode == 3 else "ok",
                  "integrity": {
                      "status": "warnings" if integrity.returncode == 3 else "ok",
                  },
                  "pages": parse_geometry(detailed.stdout, page_count, info),
                  "images": image_coverage(images_result.stdout, page_count),
                  "text": text_coverage(text_result.stdout, page_count),
              }
          )
          return emit(base, 0)
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
  • SKILL.md 5.6 KB
    ---
    argument-hint: "[file ...]"
    compatibility: Requires macOS, uv, Poppler, qpdf, Ghostscript, OCRmyPDF with Tesseract language data, and img2pdf.
    name: pdf
    description:
      "Use when PDF files are the primary input or output: read, compare, reconcile, extract text/tables/images, OCR scans,
      fill forms, split, merge, rotate, rename, compress, or convert between PDF and images. Optimized for private
      financial, tax, legal, and health documents on macOS."
    ---
    
    # PDF
    
    Process PDFs locally on macOS with exact extraction, source preservation, deliberate tool routing, and structural plus
    semantic validation.
    
    ## Invariants
    
    1. Run extraction and transformations locally. Task-relevant document evidence in tool output and internal agent reports
       may be processed by the configured model provider. Require explicit user authorization and an external-disclosure
       review before uploading or sending document contents outside that agent workflow. Package and language-data downloads
       do not authorize document disclosure.
    2. Preserve every original PDF byte-for-byte. Write a sibling output, copy, or explicitly named destination unless the
       user authorizes destructive replacement.
    3. Preserve monetary values, identifiers, dates, signs, and displayed precision as strings. Use `decimal.Decimal` for
       arithmetic; never infer missing rows or silently discard headers, footnotes, continuation lines, or boundary pages.
    4. Inspect structure and representative renders before choosing a transformation. Use the smallest tool that preserves
       the required layout, forms, annotations, and image quality.
    5. Validate every written PDF structurally and against task semantics. A command exiting successfully is not evidence
       that extracted rows, totals, page boundaries, form appearances, or visual layout are correct.
    6. Keep reports concise for private financial, tax, legal, and health documents. Prefer counts, reconciliations, and
       file references over raw sensitive rows unless the rows materially support the task or the user asks for them.
    
    ## Profile First
    
    Resolve the skill directory from this `SKILL.md`, then profile every unknown input:
    
    ```sh
    uv run "<skill-dir>/scripts/profile.py" "<input.pdf>"
    ```
    
    The helper emits schema-versioned JSON with integrity, encryption, page geometry/rotation, image counts, and per-page
    text coverage without document text. Stop on `password_required`; password handling is outside this skill.
    
    When layout, cropping, OCR quality, signatures, or form placement matters, render the first and last page, every
    structural boundary, and any page behind a discrepancy. For dense charts, tables, or technical drawings, render at
    higher resolution and crop or zoom the relevant region before reading values.
    
    ## Route by Evidence
    
    | Need                                      | Preferred route                                                          |
    | ----------------------------------------- | ------------------------------------------------------------------------ |
    | Quick reading or page-aware extraction    | Host PDF reader when available, then `pdftotext -layout`                 |
    | Coordinates, columns, or difficult tables | Poppler bounding boxes, then `pdfplumber` through `uv run`               |
    | Image-only or materially incomplete text  | OCRmyPDF with Tesseract; default languages `eng+ron`                     |
    | Merge, split, rotate, or integrity checks | qpdf                                                                     |
    | Render pages or extract embedded images   | `pdftocairo` or `pdfimages`                                              |
    | Convert ordered images into a PDF         | img2pdf                                                                  |
    | Reduce size                               | qpdf lossless rewrite first; Ghostscript only for an accepted lossy pass |
    | Inspect, fill, flatten, or overlay forms  | Read [references/forms.md](references/forms.md) first                    |
    
    Read [references/recipes.md](references/recipes.md) only when exact commands for the selected extraction,
    transformation, OCR, image, comparison, or compression branch are needed.
    
    ## Execute and Reconcile
    
    1. Profile inputs and identify whether each page is digital, scanned, mixed, rotated, or image-heavy.
    2. Extract or transform into a new path. For tabular documents, retain page provenance and parse continuations across
       page breaks before assigning rows.
    3. Reconcile financial and evidentiary output with every available invariant: page and row counts, opening/closing
       balances, inflows/outflows, subtotals, displayed totals, date coverage, and source hashes when provenance matters.
    4. For comparisons, extract both sources independently, enumerate overlapping and unique facts, and render the pages
       behind every material disagreement. Distinguish a real discrepancy from an extraction failure.
    5. For split or rename work, establish an old-to-new map from stable content identifiers. Copy by default, preserve
       contextual boundary pages when needed, and verify the first and last page of every result.
    6. Validate outputs with qpdf, expected page count/dimensions, text coverage, representative renders, and the task's
       semantic invariants. Retain OCR sidecars or extraction intermediates only when they are requested or useful evidence.
    
    Completion requires preserved originals, intentional outputs, successful structural checks, semantic reconciliation, and
    a concise report of paths and evidence. Lead read-only reports with `### 📄 PDF — 🔎 inspected, no files written`; use
    `### 📄 PDF — ✅ updated` only after all required validation passes, and `### 📄 PDF — ⛔ not deliverable` when a
    required check fails.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related