{"slug":"pdf-10","title":"pdf","summary":"Imported from paulrberg/agent-skills/skills/pdf.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-10-01T14:18:55.498507Z","repo":{"url":"https://github.com/PaulRBerg/agent-skills","stars":94,"forks":7,"license":"MIT","updatedAt":"2026-10-01T13:58:51Z"},"bodyHtml":"<hr>\n<h2>argument-hint: \"[file ...]\"\ncompatibility: Requires macOS, uv, Poppler, qpdf, Ghostscript, OCRmyPDF with Tesseract language data, and img2pdf.\nname: pdf\ndescription:\n\"Use when PDF files are the primary input or output: read, compare, reconcile, extract text/tables/images, OCR scans,\nfill forms, split, merge, rotate, rename, compress, or convert between PDF and images. Optimized for private\nfinancial, tax, legal, and health documents on macOS.\"</h2>\n<h1>PDF</h1>\n<p>Process PDFs locally on macOS with exact extraction, source preservation, deliberate tool routing, and structural plus\nsemantic validation.</p>\n<h2>Invariants</h2>\n<ol>\n<li>Run extraction and transformations locally. Task-relevant document evidence in tool output and internal agent reports\nmay be processed by the configured model provider. Require explicit user authorization and an external-disclosure\nreview before uploading or sending document contents outside that agent workflow. Package and language-data downloads\ndo not authorize document disclosure.</li>\n<li>Preserve every original PDF byte-for-byte. Write a sibling output, copy, or explicitly named destination unless the\nuser authorizes destructive replacement.</li>\n<li>Preserve monetary values, identifiers, dates, signs, and displayed precision as strings. Use <code>decimal.Decimal</code> for\narithmetic; never infer missing rows or silently discard headers, footnotes, continuation lines, or boundary pages.</li>\n<li>Inspect structure and representative renders before choosing a transformation. Use the smallest tool that preserves\nthe required layout, forms, annotations, and image quality.</li>\n<li>Validate every written PDF structurally and against task semantics. A command exiting successfully is not evidence\nthat extracted rows, totals, page boundaries, form appearances, or visual layout are correct.</li>\n<li>Keep reports concise for private financial, tax, legal, and health documents. Prefer counts, reconciliations, and\nfile references over raw sensitive rows unless the rows materially support the task or the user asks for them.</li>\n</ol>\n<h2>Profile First</h2>\n<p>Resolve the skill directory from this <code>SKILL.md</code>, then profile every unknown input:</p>\n<pre><code>uv run \"&lt;skill-dir&gt;/scripts/profile.py\" \"&lt;input.pdf&gt;\"\n</code></pre>\n<p>The helper emits schema-versioned JSON with integrity, encryption, page geometry/rotation, image counts, and per-page\ntext coverage without document text. Stop on <code>password_required</code>; password handling is outside this skill.</p>\n<p>When layout, cropping, OCR quality, signatures, or form placement matters, render the first and last page, every\nstructural boundary, and any page behind a discrepancy. For dense charts, tables, or technical drawings, render at\nhigher resolution and crop or zoom the relevant region before reading values.</p>\n<h2>Route by Evidence</h2>\n<table>\n<thead>\n<tr>\n<th>Need</th>\n<th>Preferred route</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Quick reading or page-aware extraction</td>\n<td>Host PDF reader when available, then <code>pdftotext -layout</code></td>\n</tr>\n<tr>\n<td>Coordinates, columns, or difficult tables</td>\n<td>Poppler bounding boxes, then <code>pdfplumber</code> through <code>uv run</code></td>\n</tr>\n<tr>\n<td>Image-only or materially incomplete text</td>\n<td>OCRmyPDF with Tesseract; default languages <code>eng+ron</code></td>\n</tr>\n<tr>\n<td>Merge, split, rotate, or integrity checks</td>\n<td>qpdf</td>\n</tr>\n<tr>\n<td>Render pages or extract embedded images</td>\n<td><code>pdftocairo</code> or <code>pdfimages</code></td>\n</tr>\n<tr>\n<td>Convert ordered images into a PDF</td>\n<td>img2pdf</td>\n</tr>\n<tr>\n<td>Reduce size</td>\n<td>qpdf lossless rewrite first; Ghostscript only for an accepted lossy pass</td>\n</tr>\n<tr>\n<td>Inspect, fill, flatten, or overlay forms</td>\n<td>Read <a href=\"references/forms.md\">references/forms.md</a> first</td>\n</tr>\n</tbody>\n</table>\n<p>Read <a href=\"references/recipes.md\">references/recipes.md</a> only when exact commands for the selected extraction,\ntransformation, OCR, image, comparison, or compression branch are needed.</p>\n<h2>Execute and Reconcile</h2>\n<ol>\n<li>Profile inputs and identify whether each page is digital, scanned, mixed, rotated, or image-heavy.</li>\n<li>Extract or transform into a new path. For tabular documents, retain page provenance and parse continuations across\npage breaks before assigning rows.</li>\n<li>Reconcile financial and evidentiary output with every available invariant: page and row counts, opening/closing\nbalances, inflows/outflows, subtotals, displayed totals, date coverage, and source hashes when provenance matters.</li>\n<li>For comparisons, extract both sources independently, enumerate overlapping and unique facts, and render the pages\nbehind every material disagreement. Distinguish a real discrepancy from an extraction failure.</li>\n<li>For split or rename work, establish an old-to-new map from stable content identifiers. Copy by default, preserve\ncontextual boundary pages when needed, and verify the first and last page of every result.</li>\n<li>Validate outputs with qpdf, expected page count/dimensions, text coverage, representative renders, and the task's\nsemantic invariants. Retain OCR sidecars or extraction intermediates only when they are requested or useful evidence.</li>\n</ol>\n<p>Completion requires preserved originals, intentional outputs, successful structural checks, semantic reconciliation, and\na concise report of paths and evidence. Lead read-only reports with <code>### \uD83D\uDCC4 PDF — \uD83D\uDD0E inspected, no files written</code>; use\n<code>### \uD83D\uDCC4 PDF — ✅ updated</code> only after all required validation passes, and <code>### \uD83D\uDCC4 PDF — ⛔ not deliverable</code> when a\nrequired check fails.</p>\n","files":[{"path":"agents/openai.yaml","sizeBytes":42,"isText":true},{"path":"references/forms.md","sizeBytes":2624,"isText":true},{"path":"references/recipes.md","sizeBytes":4772,"isText":true},{"path":"scripts/form.py","sizeBytes":14410,"isText":true},{"path":"scripts/profile.py","sizeBytes":6647,"isText":true},{"path":"SKILL.md","sizeBytes":5743,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-10-01T14:20:17.422016Z","sha256":"6A6B7D8C883B5436807894D432903CBB9389E8DCEE1101CB447C7D3D4805630E","sizeBytes":12860},"review":null,"source":{"repositoryUrl":"https://github.com/PaulRBerg/agent-skills","path":"skills/pdf","license":"MIT","commit":"913232a0d608198e178e7adb9ff830ceef4776fc","subtreeSha":"9C449645ED82AC18609D2FB6D972282DA19D0015D1B32257B2A77298B2A6730A","lastSyncedAt":"2026-10-01T14:18:48.927149Z"},"reviewedAt":"2026-10-01T14:23:17.942292Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/PaulRBerg/agent-skills/tree/main/skills/pdf"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install paulrberg-agent-skills@llmmart"},{"target":"git","command":"git clone https://github.com/PaulRBerg/agent-skills.git"}]}