Imported from paulrberg/agent-skills/skills/pdf.
Install
npx skills add https://github.com/PaulRBerg/agent-skills/tree/main/skills/pdf
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install paulrberg-agent-skills@llmmart
git clone https://github.com/PaulRBerg/agent-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole paulrberg/agent-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Process PDFs locally on macOS with exact extraction, source preservation, deliberate tool routing, and structural plus semantic validation.
Invariants
- Run extraction and transformations locally. Task-relevant document evidence in tool output and internal agent reports may be processed by the configured model provider. Require explicit user authorization and an external-disclosure review before uploading or sending document contents outside that agent workflow. Package and language-data downloads do not authorize document disclosure.
- Preserve every original PDF byte-for-byte. Write a sibling output, copy, or explicitly named destination unless the user authorizes destructive replacement.
- Preserve monetary values, identifiers, dates, signs, and displayed precision as strings. Use
decimal.Decimalfor arithmetic; never infer missing rows or silently discard headers, footnotes, continuation lines, or boundary pages. - Inspect structure and representative renders before choosing a transformation. Use the smallest tool that preserves the required layout, forms, annotations, and image quality.
- Validate every written PDF structurally and against task semantics. A command exiting successfully is not evidence that extracted rows, totals, page boundaries, form appearances, or visual layout are correct.
- Keep reports concise for private financial, tax, legal, and health documents. Prefer counts, reconciliations, and file references over raw sensitive rows unless the rows materially support the task or the user asks for them.
Profile First
Resolve the skill directory from this SKILL.md, then profile every unknown input:
uv run "<skill-dir>/scripts/profile.py" "<input.pdf>"
The helper emits schema-versioned JSON with integrity, encryption, page geometry/rotation, image counts, and per-page
text coverage without document text. Stop on password_required; password handling is outside this skill.
When layout, cropping, OCR quality, signatures, or form placement matters, render the first and last page, every structural boundary, and any page behind a discrepancy. For dense charts, tables, or technical drawings, render at higher resolution and crop or zoom the relevant region before reading values.
Route by Evidence
| Need | Preferred route |
|---|---|
| Quick reading or page-aware extraction | Host PDF reader when available, then pdftotext -layout |
| Coordinates, columns, or difficult tables | Poppler bounding boxes, then pdfplumber through uv run |
| Image-only or materially incomplete text | OCRmyPDF with Tesseract; default languages eng+ron |
| Merge, split, rotate, or integrity checks | qpdf |
| Render pages or extract embedded images | pdftocairo or pdfimages |
| Convert ordered images into a PDF | img2pdf |
| Reduce size | qpdf lossless rewrite first; Ghostscript only for an accepted lossy pass |
| Inspect, fill, flatten, or overlay forms | Read references/forms.md first |
Read references/recipes.md only when exact commands for the selected extraction, transformation, OCR, image, comparison, or compression branch are needed.
Execute and Reconcile
- Profile inputs and identify whether each page is digital, scanned, mixed, rotated, or image-heavy.
- Extract or transform into a new path. For tabular documents, retain page provenance and parse continuations across page breaks before assigning rows.
- Reconcile financial and evidentiary output with every available invariant: page and row counts, opening/closing balances, inflows/outflows, subtotals, displayed totals, date coverage, and source hashes when provenance matters.
- For comparisons, extract both sources independently, enumerate overlapping and unique facts, and render the pages behind every material disagreement. Distinguish a real discrepancy from an extraction failure.
- For split or rename work, establish an old-to-new map from stable content identifiers. Copy by default, preserve contextual boundary pages when needed, and verify the first and last page of every result.
- Validate outputs with qpdf, expected page count/dimensions, text coverage, representative renders, and the task's semantic invariants. Retain OCR sidecars or extraction intermediates only when they are requested or useful evidence.
Completion requires preserved originals, intentional outputs, successful structural checks, semantic reconciliation, and
a concise report of paths and evidence. Lead read-only reports with ### 📄 PDF — 🔎 inspected, no files written; use
### 📄 PDF — ✅ updated only after all required validation passes, and ### 📄 PDF — ⛔ not deliverable when a
required check fails.
Files (agent-skills)
-
agents
-
openai.yaml 42 B
policy: allow_implicit_invocation: true
-
-
references
-
forms.md 2.6 KB
# PDF Forms Choose the branch from the document structure; do not guess from appearance. ## Inspect ```sh uv run "<skill-dir>/scripts/form.py" inspect "input.pdf" ``` The JSON reports whether an AcroForm exists, whether XFA is present, and each field's fully qualified name, type, current and default values, options, flags, page, and rectangle. Route by result: - AcroForm without XFA or signatures: fill known fields. - Flat page with no AcroForm: use coordinate overlays after rendering and calibration. - XFA or signature fields: stop. This helper intentionally does not preserve or generate those workflows. ## Fill an AcroForm Create a JSON object keyed by the exact field names returned by `inspect`: ```json { "person.name": "Ada Lovelace", "preferences.language": "Romanian", "topics": ["tax", "finance"] } ``` Fill to a new output: ```sh uv run "<skill-dir>/scripts/form.py" fill "input.pdf" "values.json" "filled.pdf" ``` Use `--flatten` only for a final, non-editable delivery. Flattening draws field appearances, removes widget annotations, and removes the AcroForm dictionary. Verify a rendered output because stored values alone do not prove visibility. The helper rejects unknown fields, XFA, signatures, input/output aliasing, and an existing output unless `--force` is explicitly supplied. ## Overlay a flat form Render the page first and calibrate in PDF points from the lower-left corner. Page numbers are one-based. Create a JSON array: ```json [ { "page": 1, "x": 126, "y": 618, "text": "Ada Lovelace", "font_size": 10 }, { "page": 2, "x": 90, "y": 144, "text": "București" } ] ``` Apply it to a new output: ```sh uv run "<skill-dir>/scripts/form.py" overlay "input.pdf" "placements.json" "filled-flat.pdf" ``` The default font is `/System/Library/Fonts/Supplemental/Arial.ttf`, which supports Romanian text. `font_size` defaults to 10. Placements are single-line text; add separate entries instead of relying on wrapping. Render every affected page after overlaying. Check baseline, clipping, diacritics, rotation, and whether the visual position still matches at ordinary and high zoom. The helper preserves page count and dimensions but cannot determine whether coordinates are semantically correct. ## Validate For either branch: 1. Require the helper's success JSON and qpdf integrity. 2. Re-run `inspect` for editable AcroForms and compare intended values. 3. Render every affected page and visually inspect field appearances or overlay placement. 4. Confirm page count and dimensions match the input. 5. Keep the untouched input alongside the validated output. -
recipes.md 4.7 KB
# PDF Recipes Load only the branch needed for the current task. Keep every output distinct from its input and quote paths. ## Inspect and extract Start with the bundled factual profile, then inspect raw tool output only as needed: ```sh uv run "<skill-dir>/scripts/profile.py" "input.pdf" pdfinfo -box "input.pdf" qpdf --check "input.pdf" pdfimages -list "input.pdf" ``` Preserve reading order where possible: ```sh pdftotext -layout "input.pdf" "output.txt" pdftotext -f 3 -l 7 -layout "input.pdf" "pages-3-7.txt" pdftotext -f 1 -l 1 -x 36 -y 72 -W 540 -H 648 -layout "input.pdf" "crop.txt" pdftotext -bbox-layout "input.pdf" "layout.html" ``` Use bounding boxes to diagnose column interleaving or logical half-pages. Escalate to `pdfplumber` only when the Poppler output cannot preserve the required table structure: ```sh uv run --with 'pdfplumber>=0.11.10,<0.12' python - "input.pdf" <<'PY' import json import sys import pdfplumber with pdfplumber.open(sys.argv[1]) as document: print(json.dumps([page.extract_tables() for page in document.pages], ensure_ascii=False)) PY ``` Treat extracted tables as candidates, not truth. Rebuild wrapped descriptions and continuations, retain page numbers, and reconcile exact row counts and totals against the PDF. ## Compare documents Extract each input independently with the same appropriate route. Compare normalized facts while retaining the original strings and page provenance. Report: - facts present in both documents; - facts unique to each document; - materially different values, dates, identifiers, qualifications, or footnotes; - pages rendered to distinguish source differences from extraction errors. Do not infer that missing extracted text means missing source content. Inspect the relevant render or OCR coverage first. ## OCR scans Use OCR only for image-only or materially incomplete pages. The default covers English and Romanian: ```sh ocrmypdf --output-type pdf --skip-text -l eng+ron --sidecar "ocr.txt" "input.pdf" "ocr.pdf" qpdf --check "ocr.pdf" pdftotext -layout "ocr.pdf" "ocr-check.txt" ``` Add `--rotate-pages` or `--deskew` only when profiling or rendered pages show the need. Add `--clean` only after accepting that image processing may alter visual evidence. Re-render affected pages after any of these options. ## Render and extract images Render pages for visual comparison: ```sh pdftocairo -png -r 200 "input.pdf" "page" pdftocairo -f 4 -l 4 -png -r 300 "input.pdf" "page-4" ``` Extract embedded images without rasterizing whole pages: ```sh pdfimages -list "input.pdf" pdfimages -all "input.pdf" "image" ``` Combine ordered images without recompressing them unnecessarily: ```sh img2pdf "page-01.png" "page-02.jpg" --output "combined.pdf" qpdf --check "combined.pdf" ``` Confirm image order, orientation, page dimensions, and representative renders. ## Merge, split, and rotate Merge in explicit order: ```sh qpdf --empty --pages "part-1.pdf" "part-2.pdf" -- "merged.pdf" ``` Extract a range or retain a boundary page for context: ```sh qpdf "input.pdf" --pages . 1-5 -- "part-1.pdf" qpdf "input.pdf" --pages . 5-10 -- "part-2-with-boundary.pdf" ``` Rotate selected pages clockwise: ```sh qpdf "input.pdf" "rotated.pdf" --rotate=+90:1,3 ``` Check every result and inspect its first and last page: ```sh qpdf --check "output.pdf" pdfinfo "output.pdf" pdftotext -f 1 -l 1 -layout "output.pdf" - ``` ## Rename a corpus Build a complete old-to-new map before writing. Derive names from stable content such as issuer, document type, account suffix, and covered date range. Detect collisions and ambiguous documents first. Copy to the new names unless the user explicitly authorizes renaming originals, then profile both sides and compare hashes when a byte-identical copy is expected. ## Compress Try a lossless structural rewrite first: ```sh qpdf --object-streams=generate --recompress-flate --compression-level=9 "input.pdf" "lossless.pdf" ``` Use Ghostscript only when a smaller lossy output is acceptable: ```sh gs -sDEVICE=pdfwrite -dCompatibilityLevel=1.7 -dPDFSETTINGS=/ebook -dNOPAUSE -dQUIET -dBATCH \ -sOutputFile="compressed.pdf" "input.pdf" ``` Compare byte size only after qpdf integrity, page count/dimensions, text extraction, forms/annotations where relevant, and representative renders pass. Keep the original and the smallest acceptable validated output. ## Final validation At minimum, require: ```sh qpdf --check "output.pdf" pdfinfo -box "output.pdf" pdftotext -layout "output.pdf" "output.txt" ``` Add domain checks: exact totals and balances for statements, field values and appearances for forms, first/last pages for splits, order and dimensions for image conversions, and visual comparison for OCR or compression.
-
-
scripts
-
form.py 14.1 KB
#!/usr/bin/env -S uv run --script # /// script # requires-python = ">=3.12" # dependencies = ["pypdf>=6.15,<7", "reportlab>=5,<6"] # /// """Inspect, fill, flatten, and overlay local PDF forms.""" from __future__ import annotations import argparse import io import json import os import shutil import subprocess import tempfile from collections import defaultdict from pathlib import Path from typing import Any from pypdf import PdfReader, PdfWriter from pypdf.errors import PdfReadError from pypdf.generic import ArrayObject, NameObject from reportlab.pdfbase import pdfmetrics from reportlab.pdfbase.ttfonts import TTFont from reportlab.pdfgen import canvas SCHEMA_VERSION = 1 DEFAULT_FONT = Path("/System/Library/Fonts/Supplemental/Arial.ttf") FONT_NAME = "PDFSkillArial" class UserError(Exception): def __init__(self, code: str, message: str) -> None: super().__init__(message) self.code = code def emit(payload: dict[str, Any], returncode: int = 0) -> int: print(json.dumps(payload, ensure_ascii=False, indent=2, sort_keys=True)) return returncode def pdf_object(value: Any) -> Any: if hasattr(value, "get_object"): value = value.get_object() if value is None or isinstance(value, (bool, int, float, str)): return value if isinstance(value, dict): return {str(key): pdf_object(item) for key, item in value.items()} if isinstance(value, (list, tuple)): return [pdf_object(item) for item in value] return str(value) def resolved_dictionary(value: Any) -> Any: return value.get_object() if hasattr(value, "get_object") else value def acroform(reader: PdfReader) -> Any | None: root = resolved_dictionary(reader.trailer["/Root"]) value = root.get("/AcroForm") return resolved_dictionary(value) if value is not None else None def ensure_readable(reader: PdfReader) -> None: if reader.is_encrypted: raise UserError("password_required", "encrypted PDFs are not supported") def qualified_field_name(annotation: Any) -> str | None: parts: list[str] = [] node = resolved_dictionary(annotation) seen: set[int] = set() while node is not None and id(node) not in seen: seen.add(id(node)) if node.get("/T") is not None: parts.append(str(node["/T"])) parent = node.get("/Parent") node = resolved_dictionary(parent) if parent is not None else None return ".".join(reversed(parts)) or None def widget_locations(reader: PdfReader) -> dict[str, list[dict[str, Any]]]: locations: dict[str, list[dict[str, Any]]] = defaultdict(list) for page_number, page in enumerate(reader.pages, 1): for reference in page.get("/Annots", []): annotation = resolved_dictionary(reference) if str(annotation.get("/Subtype")) != "/Widget": continue name = qualified_field_name(annotation) if name is None: continue rectangle = annotation.get("/Rect") locations[name].append( { "page": page_number, "rectangle": [float(value) for value in rectangle] if rectangle else None, } ) return locations def inspect_payload(input_path: Path) -> dict[str, Any]: reader = PdfReader(str(input_path)) ensure_readable(reader) form = acroform(reader) fields = reader.get_fields() or {} locations = widget_locations(reader) items: list[dict[str, Any]] = [] for name in sorted(fields): field = resolved_dictionary(fields[name]) location = locations.get(name, [{}])[0] items.append( { "name": name, "type": str(field.get("/FT")) if field.get("/FT") is not None else None, "value": pdf_object(field.get("/V")), "default_value": pdf_object(field.get("/DV")), "options": pdf_object(field.get("/Opt")), "flags": int(field.get("/Ff", 0)), "page": location.get("page"), "rectangle": location.get("rectangle"), } ) return { "schema_version": SCHEMA_VERSION, "status": "ok", "operation": "inspect", "input": str(input_path), "page_count": len(reader.pages), "acroform": form is not None, "xfa": bool(form is not None and form.get("/XFA") is not None), "fields": items, } def load_json(path: Path) -> Any: try: return json.loads(path.read_text(encoding="utf-8")) except FileNotFoundError as error: raise UserError("missing_json", "JSON input does not exist") from error except json.JSONDecodeError as error: raise UserError("invalid_json", f"invalid JSON at line {error.lineno}") from error def prepare_paths(input_value: str, output_value: str, force: bool) -> tuple[Path, Path]: input_path = Path(input_value).expanduser().resolve() output_path = Path(output_value).expanduser().resolve() if not input_path.is_file(): raise UserError("missing_input", "input PDF does not exist") if input_path == output_path: raise UserError("input_output_alias", "input and output must be different paths") if output_path.exists() and not force: raise UserError("output_exists", "output exists; pass --force to replace it") if not output_path.parent.is_dir(): raise UserError("missing_output_directory", "output directory does not exist") return input_path, output_path def validate_writable_form(reader: PdfReader) -> dict[str, Any]: ensure_readable(reader) form = acroform(reader) fields = reader.get_fields() or {} if form is None or not fields: raise UserError("no_acroform", "PDF has no fillable AcroForm fields") if form.get("/XFA") is not None: raise UserError("unsupported_xfa", "XFA forms are not supported") if any(str(resolved_dictionary(field).get("/FT")) == "/Sig" for field in fields.values()): raise UserError("unsupported_signature", "signature fields are not supported") return fields def atomic_write(writer: PdfWriter, output_path: Path) -> None: if shutil.which("qpdf") is None: raise UserError("missing_qpdf", "qpdf is required to validate written PDFs") descriptor, temporary_name = tempfile.mkstemp( prefix=f".{output_path.name}.", suffix=".tmp.pdf", dir=output_path.parent ) os.close(descriptor) temporary_path = Path(temporary_name) try: writer.write(str(temporary_path)) check = subprocess.run( ["qpdf", "--check", str(temporary_path)], capture_output=True, check=False, text=True, ) if check.returncode not in (0, 3): raise UserError("integrity_failed", "qpdf rejected the generated PDF") os.replace(temporary_path, output_path) finally: temporary_path.unlink(missing_ok=True) def command_inspect(args: argparse.Namespace) -> int: input_path = Path(args.input).expanduser().resolve() if not input_path.is_file(): raise UserError("missing_input", "input PDF does not exist") return emit(inspect_payload(input_path)) def command_fill(args: argparse.Namespace) -> int: input_path, output_path = prepare_paths(args.input, args.output, args.force) values_path = Path(args.values).expanduser().resolve() values = load_json(values_path) if not isinstance(values, dict) or not all(isinstance(key, str) for key in values): raise UserError("invalid_values", "values JSON must be an object keyed by field name") for value in values.values(): if isinstance(value, str): continue if isinstance(value, list) and all(isinstance(item, str) for item in value): continue raise UserError("invalid_values", "field values must be strings or lists of strings") reader = PdfReader(str(input_path)) fields = validate_writable_form(reader) unknown = sorted(set(values) - set(fields)) if unknown: raise UserError("unknown_fields", f"unknown field names: {', '.join(unknown)}") writer = PdfWriter(clone_from=str(input_path)) for page in writer.pages: writer.update_page_form_field_values( page, values, auto_regenerate=False, flatten=args.flatten, ) if args.flatten: for page in writer.pages: annotations = page.get("/Annots") if annotations is None: continue retained = ArrayObject( [ reference for reference in annotations if str(resolved_dictionary(reference).get("/Subtype")) != "/Widget" ] ) if retained: page[NameObject("/Annots")] = retained else: page.pop(NameObject("/Annots"), None) writer.root_object.pop(NameObject("/AcroForm"), None) atomic_write(writer, output_path) return emit( { "schema_version": SCHEMA_VERSION, "status": "ok", "operation": "fill", "output": str(output_path), "fields_written": sorted(values), "flattened": args.flatten, } ) def positive_number(value: Any, name: str) -> float: if isinstance(value, bool) or not isinstance(value, (int, float)): raise UserError("invalid_placements", f"{name} must be a number") result = float(value) if name == "font_size" and result <= 0: raise UserError("invalid_placements", "font_size must be greater than zero") return result def command_overlay(args: argparse.Namespace) -> int: input_path, output_path = prepare_paths(args.input, args.output, args.force) placements = load_json(Path(args.placements).expanduser().resolve()) if not isinstance(placements, list): raise UserError("invalid_placements", "placements JSON must be an array") reader = PdfReader(str(input_path)) ensure_readable(reader) grouped: dict[int, list[dict[str, Any]]] = defaultdict(list) for index, placement in enumerate(placements): if not isinstance(placement, dict): raise UserError("invalid_placements", f"placement {index} must be an object") page = placement.get("page") text = placement.get("text") if isinstance(page, bool) or not isinstance(page, int) or not 1 <= page <= len(reader.pages): raise UserError("invalid_placements", f"placement {index} has an invalid page") if not isinstance(text, str) or "\n" in text or "\r" in text: raise UserError("invalid_placements", f"placement {index} text must be one line") grouped[page].append( { "x": positive_number(placement.get("x"), "x"), "y": positive_number(placement.get("y"), "y"), "text": text, "font_size": positive_number(placement.get("font_size", 10), "font_size"), } ) if not DEFAULT_FONT.is_file(): raise UserError("missing_font", f"default font not found: {DEFAULT_FONT}") pdfmetrics.registerFont(TTFont(FONT_NAME, str(DEFAULT_FONT))) writer = PdfWriter(clone_from=str(input_path)) input_dimensions = [ (float(page.mediabox.width), float(page.mediabox.height)) for page in reader.pages ] for page_number, page_placements in grouped.items(): width, height = input_dimensions[page_number - 1] buffer = io.BytesIO() overlay_canvas = canvas.Canvas(buffer, pagesize=(width, height)) for placement in page_placements: overlay_canvas.setFont(FONT_NAME, placement["font_size"]) overlay_canvas.drawString(placement["x"], placement["y"], placement["text"]) overlay_canvas.save() buffer.seek(0) overlay_page = PdfReader(buffer).pages[0] writer.pages[page_number - 1].merge_page(overlay_page) output_dimensions = [ (float(page.mediabox.width), float(page.mediabox.height)) for page in writer.pages ] if output_dimensions != input_dimensions: raise UserError("geometry_changed", "overlay changed page dimensions") atomic_write(writer, output_path) return emit( { "schema_version": SCHEMA_VERSION, "status": "ok", "operation": "overlay", "output": str(output_path), "placements_written": len(placements), "pages_affected": sorted(grouped), } ) def parser() -> argparse.ArgumentParser: result = argparse.ArgumentParser(description=__doc__) commands = result.add_subparsers(dest="command", required=True) inspect_parser = commands.add_parser("inspect") inspect_parser.add_argument("input") inspect_parser.set_defaults(handler=command_inspect) fill_parser = commands.add_parser("fill") fill_parser.add_argument("input") fill_parser.add_argument("values") fill_parser.add_argument("output") fill_parser.add_argument("--flatten", action="store_true") fill_parser.add_argument("--force", action="store_true") fill_parser.set_defaults(handler=command_fill) overlay_parser = commands.add_parser("overlay") overlay_parser.add_argument("input") overlay_parser.add_argument("placements") overlay_parser.add_argument("output") overlay_parser.add_argument("--force", action="store_true") overlay_parser.set_defaults(handler=command_overlay) return result def main() -> int: args = parser().parse_args() try: return args.handler(args) except UserError as error: return emit( { "schema_version": SCHEMA_VERSION, "status": "error", "error": {"code": error.code, "message": str(error)}, }, 1, ) except (KeyError, OSError, PdfReadError, TypeError, ValueError) as error: return emit( { "schema_version": SCHEMA_VERSION, "status": "error", "error": {"code": "unexpected_error", "message": str(error)}, }, 1, ) if __name__ == "__main__": raise SystemExit(main()) -
profile.py 6.5 KB
#!/usr/bin/env -S uv run --script # /// script # requires-python = ">=3.12" # dependencies = [] # /// """Emit a compact factual profile for a local PDF.""" from __future__ import annotations import json import re import shutil import subprocess import sys from collections import Counter from pathlib import Path from typing import Any REQUIRED_TOOLS = ("pdfimages", "pdfinfo", "pdftotext", "qpdf") def emit(payload: dict[str, Any], returncode: int) -> int: print(json.dumps(payload, ensure_ascii=False, indent=2, sort_keys=True)) return returncode def fail(code: str, message: str, pdf_path: Path | None = None) -> int: payload: dict[str, Any] = { "schema_version": 1, "status": "error", "error": {"code": code, "message": message}, } if pdf_path is not None: payload["path"] = str(pdf_path) return emit(payload, 1) def run(args: list[str]) -> subprocess.CompletedProcess[str]: return subprocess.run( args, capture_output=True, check=False, encoding="utf-8", errors="replace", text=True, ) def parse_pdfinfo(output: str) -> dict[str, str]: result: dict[str, str] = {} for line in output.splitlines(): if ":" not in line: continue key, value = line.split(":", 1) result[key.strip()] = value.strip() return result def parse_geometry(output: str, page_count: int, fallback: dict[str, str]) -> list[dict[str, Any]]: pages: dict[int, dict[str, Any]] = { page: {"page": page, "width_points": None, "height_points": None, "rotation": 0} for page in range(1, page_count + 1) } size_pattern = re.compile( r"^Page\s+(\d+)\s+size:\s+([0-9.]+)\s+x\s+([0-9.]+)\s+pts", re.MULTILINE, ) rotation_pattern = re.compile(r"^Page\s+(\d+)\s+rot:\s+(-?\d+)", re.MULTILINE) for match in size_pattern.finditer(output): page = int(match.group(1)) if page in pages: pages[page]["width_points"] = float(match.group(2)) pages[page]["height_points"] = float(match.group(3)) for match in rotation_pattern.finditer(output): page = int(match.group(1)) if page in pages: pages[page]["rotation"] = int(match.group(2)) fallback_size = re.search(r"([0-9.]+)\s+x\s+([0-9.]+)\s+pts", fallback.get("Page size", "")) fallback_rotation = int(fallback.get("Page rot", "0") or 0) if fallback_size: for page in pages.values(): if page["width_points"] is None: page["width_points"] = float(fallback_size.group(1)) page["height_points"] = float(fallback_size.group(2)) page["rotation"] = fallback_rotation return list(pages.values()) def text_coverage(output: str, page_count: int) -> dict[str, Any]: chunks = output.split("\f") if len(chunks) > page_count and not chunks[-1].strip(): chunks.pop() chunks.extend([""] * max(0, page_count - len(chunks))) counts = [sum(not character.isspace() for character in chunk) for chunk in chunks[:page_count]] return { "characters": sum(counts), "pages_with_text": [index for index, count in enumerate(counts, 1) if count > 0], "pages_without_text": [index for index, count in enumerate(counts, 1) if count == 0], "per_page": [ {"page": index, "characters": count} for index, count in enumerate(counts, 1) ], } def image_coverage(output: str, page_count: int) -> dict[str, Any]: counts: Counter[int] = Counter() for line in output.splitlines(): match = re.match(r"^\s*(\d+)\s+\d+\s+", line) if match: counts[int(match.group(1))] += 1 return { "total": sum(counts.values()), "per_page": [ {"page": page, "count": counts.get(page, 0)} for page in range(1, page_count + 1) ], } def main() -> int: if len(sys.argv) != 2: return fail("usage", "usage: profile.py <input.pdf>") pdf_path = Path(sys.argv[1]).expanduser().resolve() if not pdf_path.is_file(): return fail("missing_file", "input is not a readable file", pdf_path) missing = [tool for tool in REQUIRED_TOOLS if shutil.which(tool) is None] if missing: return fail("missing_tools", f"required tools not found: {', '.join(missing)}", pdf_path) basic = run(["pdfinfo", str(pdf_path)]) if basic.returncode != 0: return fail("unreadable_pdf", "pdfinfo could not read the input", pdf_path) info = parse_pdfinfo(basic.stdout) try: page_count = int(info.get("Pages", "0")) except ValueError: return fail("invalid_metadata", "pdfinfo returned an invalid page count", pdf_path) if page_count < 1: return fail("invalid_metadata", "PDF has no readable pages", pdf_path) encrypted = info.get("Encrypted", "unknown").lower().startswith("yes") base: dict[str, Any] = { "schema_version": 1, "path": str(pdf_path), "size_bytes": pdf_path.stat().st_size, "pdf_version": info.get("PDF version"), "encrypted": encrypted, "page_count": page_count, } if encrypted: base.update({"status": "password_required", "integrity": {"status": "not_checked"}}) return emit(base, 1) integrity = run(["qpdf", "--check", str(pdf_path)]) if integrity.returncode not in (0, 3): return fail("integrity_failed", "qpdf rejected the input", pdf_path) detailed = run(["pdfinfo", "-f", "1", "-l", str(page_count), "-box", str(pdf_path)]) if detailed.returncode != 0: return fail("metadata_failed", "pdfinfo could not inspect page geometry", pdf_path) text_result = run(["pdftotext", "-layout", str(pdf_path), "-"]) if text_result.returncode != 0: return fail("text_failed", "pdftotext could not inspect text coverage", pdf_path) images_result = run(["pdfimages", "-list", str(pdf_path)]) if images_result.returncode != 0: return fail("images_failed", "pdfimages could not inspect image coverage", pdf_path) base.update( { "status": "warnings" if integrity.returncode == 3 else "ok", "integrity": { "status": "warnings" if integrity.returncode == 3 else "ok", }, "pages": parse_geometry(detailed.stdout, page_count, info), "images": image_coverage(images_result.stdout, page_count), "text": text_coverage(text_result.stdout, page_count), } ) return emit(base, 0) if __name__ == "__main__": raise SystemExit(main())
-
-
SKILL.md 5.6 KB
--- argument-hint: "[file ...]" compatibility: Requires macOS, uv, Poppler, qpdf, Ghostscript, OCRmyPDF with Tesseract language data, and img2pdf. name: pdf description: "Use when PDF files are the primary input or output: read, compare, reconcile, extract text/tables/images, OCR scans, fill forms, split, merge, rotate, rename, compress, or convert between PDF and images. Optimized for private financial, tax, legal, and health documents on macOS." --- # PDF Process PDFs locally on macOS with exact extraction, source preservation, deliberate tool routing, and structural plus semantic validation. ## Invariants 1. Run extraction and transformations locally. Task-relevant document evidence in tool output and internal agent reports may be processed by the configured model provider. Require explicit user authorization and an external-disclosure review before uploading or sending document contents outside that agent workflow. Package and language-data downloads do not authorize document disclosure. 2. Preserve every original PDF byte-for-byte. Write a sibling output, copy, or explicitly named destination unless the user authorizes destructive replacement. 3. Preserve monetary values, identifiers, dates, signs, and displayed precision as strings. Use `decimal.Decimal` for arithmetic; never infer missing rows or silently discard headers, footnotes, continuation lines, or boundary pages. 4. Inspect structure and representative renders before choosing a transformation. Use the smallest tool that preserves the required layout, forms, annotations, and image quality. 5. Validate every written PDF structurally and against task semantics. A command exiting successfully is not evidence that extracted rows, totals, page boundaries, form appearances, or visual layout are correct. 6. Keep reports concise for private financial, tax, legal, and health documents. Prefer counts, reconciliations, and file references over raw sensitive rows unless the rows materially support the task or the user asks for them. ## Profile First Resolve the skill directory from this `SKILL.md`, then profile every unknown input: ```sh uv run "<skill-dir>/scripts/profile.py" "<input.pdf>" ``` The helper emits schema-versioned JSON with integrity, encryption, page geometry/rotation, image counts, and per-page text coverage without document text. Stop on `password_required`; password handling is outside this skill. When layout, cropping, OCR quality, signatures, or form placement matters, render the first and last page, every structural boundary, and any page behind a discrepancy. For dense charts, tables, or technical drawings, render at higher resolution and crop or zoom the relevant region before reading values. ## Route by Evidence | Need | Preferred route | | ----------------------------------------- | ------------------------------------------------------------------------ | | Quick reading or page-aware extraction | Host PDF reader when available, then `pdftotext -layout` | | Coordinates, columns, or difficult tables | Poppler bounding boxes, then `pdfplumber` through `uv run` | | Image-only or materially incomplete text | OCRmyPDF with Tesseract; default languages `eng+ron` | | Merge, split, rotate, or integrity checks | qpdf | | Render pages or extract embedded images | `pdftocairo` or `pdfimages` | | Convert ordered images into a PDF | img2pdf | | Reduce size | qpdf lossless rewrite first; Ghostscript only for an accepted lossy pass | | Inspect, fill, flatten, or overlay forms | Read [references/forms.md](references/forms.md) first | Read [references/recipes.md](references/recipes.md) only when exact commands for the selected extraction, transformation, OCR, image, comparison, or compression branch are needed. ## Execute and Reconcile 1. Profile inputs and identify whether each page is digital, scanned, mixed, rotated, or image-heavy. 2. Extract or transform into a new path. For tabular documents, retain page provenance and parse continuations across page breaks before assigning rows. 3. Reconcile financial and evidentiary output with every available invariant: page and row counts, opening/closing balances, inflows/outflows, subtotals, displayed totals, date coverage, and source hashes when provenance matters. 4. For comparisons, extract both sources independently, enumerate overlapping and unique facts, and render the pages behind every material disagreement. Distinguish a real discrepancy from an extraction failure. 5. For split or rename work, establish an old-to-new map from stable content identifiers. Copy by default, preserve contextual boundary pages when needed, and verify the first and last page of every result. 6. Validate outputs with qpdf, expected page count/dimensions, text coverage, representative renders, and the task's semantic invariants. Retain OCR sidecars or extraction intermediates only when they are requested or useful evidence. Completion requires preserved originals, intentional outputs, successful structural checks, semantic reconciliation, and a concise report of paths and evidence. Lead read-only reports with `### 📄 PDF — 🔎 inspected, no files written`; use `### 📄 PDF — ✅ updated` only after all required validation passes, and `### 📄 PDF — ⛔ not deliverable` when a required check fails.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.