design-export-repair
Fixes broken decks/PDFs exported from Claude's "Design" feature (or similar AI deck generators) — cut-off or clipped text, wrong/substituted fonts, and corrupted .pptx package structure that shows up only once you export to PDF, not while looking at it in the app. Use this skill
Install
npx skills add https://github.com/OneWave-AI/claude-skills/tree/main/design-export-repair
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install onewave-ai-claude-skills@llmmart
git clone https://github.com/OneWave-AI/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole onewave-ai/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Design export repair
Fixes AI-generated deck/PDF exports (built for Claude Design's export
path, but the checks are generic OOXML/PDF hygiene, not Claude-specific)
that look fine in the app but come apart once exported — text missing
its last few characters, fonts that don't match the original design, or
a .pptx that some converters choke on even though PowerPoint opens it
fine.
Four defect classes were found and fixed against a real broken export;
references/defect-classes.md has the full story for each, including how
they were confirmed and what the fix trades off. Skim it before extending
this skill to a case it doesn't already handle — in particular, defect #4
there is a good example of why "the shape geometry looks right" is not
the same claim as "the rendered PDF looks right," and it's worth reading
before assuming a new complaint is a geometry problem.
End-to-end: from Claude Design to a clean file
This is written for the person running Claude Code, not just for Claude — skim it once so you know what to actually do. There are two different ways a Design project reaches Claude Code, and they behave differently.
Path A — plain Export (this is the one that matches "my deck/PDF looks
broken"): In Claude Design, use Export and pick PDF, PPTX, or HTML. Save
the file, then bring it to whichever Claude Code you're using — attach it
in the chat if your setup allows that, or save it to disk and give Claude
the path. Then tell Claude to use this skill: plain language like "this
deck looks broken when I export it to PDF, can you fix it" is enough on
its own (the skill's description is written to match that phrasing), or
name it directly (design-export-repair) if you want to be explicit.
Path B — the "Send to Claude Code" button. This is a related but different feature from plain Export — it's aimed at continuing design/ prototype work in code, not at fixing a deck export, but you may end up here anyway. It offers two options, and they behave differently:
- "Send to Claude Code Web" opens a brand-new claude.ai/code session with the design bundle already sitting in that session's working directory — there's no file to save or attach, it's just already there. In that new session, tell Claude to look at what's in the working directory and use this skill on it.
- "Send to local coding agent" gives you a prompt to paste into your local terminal Claude Code. Be aware: as of this writing, that default prompt depends on a "Claude Design connector" that local Claude Code doesn't ship, and reliably fails silently (see anthropics/claude-code#69246 if you want the details). The same dialog has a "Download zip instead" option — that one actually works: it downloads a real zip of the design files. Save that zip, point local Claude Code at it (or unzip it into your project folder first), and tell Claude to use this skill on it.
Whichever path you took, scripts/fix_export.py accepts a file, a zip, or
a whole directory — unpack.py figures out what it's looking at rather
than requiring one specific shape. One thing worth knowing about Path B
specifically: its bundle format is Design's own *.dc.html canvas files,
which is a different animal from the .pptx this skill's repair pipeline
was actually built and verified against (see
references/defect-classes.md) — a .dc.html bundle gets routed through
the best-effort HTML path (convert_html.py) instead of the fully-tested
one. If you're trying to fix a broken deck/PDF export specifically, Path A
is the one to reach for.
After either path, Claude runs scripts/fix_export.py and hands back a
repaired .pptx (Path A) or PDF, plus a report explaining exactly what
was wrong and what changed. Open the repaired .pptx if you want to keep
editing in PowerPoint/Keynote/Google Slides, or use the .pdf directly if
that's the deliverable you needed.
One honest caveat: buttons and flows in Claude products change over time,
and this description of Path B is based on Claude Code's own public issue
tracker rather than hands-on testing of that specific handoff. If what you
see doesn't match this, don't get stuck on it — just get the actual design
files to Claude Code by whatever route works, and tell Claude to use this
skill; unpack.py is written to figure out what it's been handed.
Quick summary of what gets fixed
| # | Defect | Symptom | Fix |
|---|---|---|---|
| 1 | [Content_Types].xml declares parts that don't exist in the zip |
Some converters fail, drop slides, or garble output; PowerPoint may silently "repair" and hide it | Load + re-save through python-pptx (rebuilds the manifest from what's actually there) |
| 2 | A shape's box extends past the slide edge | Can get clipped by renderers that treat the slide as a strict viewport | Translate the shape back in bounds (or scale down + rescale table grids, if the shape is bigger than the slide itself) |
| 3 | Custom webfonts (Inter, Plus Jakarta Sans, JetBrains Mono seen in practice) used but not embedded | Renderer without those fonts substitutes something else — wrong look, different text measurements | Install matching bundled/fetched font files where the renderer will find them, before conversion |
| 4 | LibreOffice clips the tail of any text run with letter-spacing (a:rPr spc) |
Tracked-out labels/kickers/badges lose their last 1-4 characters — this is usually the actual cause of "text is cut off" even when #2 is also present | Render a letter-spacing-neutralized copy through LibreOffice only; the returned .pptx keeps its real letter-spacing for editing |
How to run it
Everything is one script:
python3 scripts/fix_export.py <input> --out-dir <where to write results>
<input> can be:
- a
.pptxhanded directly (this is what Claude Design's export has produced in practice — don't assume it must be wrapped in a zip) - a
.zip— it gets extracted and searched for a.pptxfirst, then an HTML/CSS/JS bundle (index.html+ assets), then a bare.pdf - a bare
.pdf— repair is limited without the source; see below
Read the script's own docstring for the full pipeline description before running it if anything about the input is unusual — it's kept in sync with what the code actually does.
After it runs, look at <out-dir>/repair_report.md. It documents, in
order: what was wrong with the package structure, every shape that got
repositioned/resized (slide, before/after coordinates), the status of
every font the deck uses (installed / fetched / falling back), whether the
letter-spacing workaround applied, and a validation pass on the final PDF
(page count vs. slide count, any pages that look unexpectedly blank, any
text that still touches a page edge). Read this back to the user — it's
written to be shown, not just logged, because the whole point of this
skill is "nothing changes silently."
Outputs land in <out-dir>:
*.repaired.pptx— still fully editable, structurally clean, no off-canvas shapes, original letter-spacing intact*.repaired.pdf— the clean PDF, converted from a letter-spacing- neutralized render copy (see defect #4) so tracked-out labels don't cliprepair_report.md/repair_report.json
When the input is a bare PDF with no source file
Structural repair (package structure, shape geometry, font substitution)
needs the editable source — there's no way to move a shape or fix a font
reference in something that's already been flattened to fixed vector
paths and rasterized text. fix_export.py still runs the same PDF
validation pass (page count sanity, blank-page check, edge-overflow check)
so you at least get an honest read on whether the PDF itself looks broken,
but say so plainly to the user rather than implying a fix happened: ask if
they can re-export from Design as a .pptx/zip instead, since that's
where the actual repair leverage is.
When the input is an HTML/CSS/JS bundle
This path (scripts/convert_html.py) is best-effort, not backed by a real
defect investigation the way the .pptx path is — there was no sample of
this shape to test against. It applies the generic fixes that most
commonly break an HTML deck's print-to-PDF (forces background/color
printing, un-hides anything relying on overflow: hidden + scrolling,
avoids mid-element page breaks) and picks a page size from whatever the
markup hints at. If you hit a real broken HTML export, treat this as a
starting point and read references/defect-classes.md's closing section
on how to diagnose a new case (render, compare, isolate, don't guess from
markup alone) rather than assuming the existing generic fixes cover it.
Extending this skill
If a deck comes through with a defect not in the table above, the method
that found defect #4 generalizes: don't debug from the XML/CSS alone,
because "looks structurally fine" and "renders fine" are different claims
and only the second one is what the user is reporting. Render the
untouched original to PDF, render a copy with one specific, isolated
change, and compare the two — that's what confirmed the letter-spacing bug
after the geometry and font fixes alone didn't resolve the visible
clipping. scripts/validate_pdf.py already does the page-count/blank-
page/edge-overflow checks; add to it if you find another cheap,
objective signal worth checking automatically on every run.
Files (claude-skills)
-
assets
-
fonts
-
Inter.ttf 856 KB · in bundle
-
JetBrainsMono-Regular.ttf 182.8 KB · in bundle
-
NOTICE.md 733 B
# Bundled fonts These are shipped with the skill so a repaired export renders with the typefaces the original design used, rather than silently substituting. Each keeps its own upstream license — the repository's MIT license covers the skill, not the font files. | Font | License | Source | |------|---------|--------| | Inter | SIL Open Font License 1.1 | https://github.com/rsms/inter | | Plus Jakarta Sans | SIL Open Font License 1.1 | https://github.com/tokotype/PlusJakartaSans | | JetBrains Mono | SIL Open Font License 1.1 | https://github.com/JetBrains/JetBrainsMono | The OFL permits bundling and redistribution, including commercially, as long as the fonts are not sold on their own and this notice travels with them. -
PlusJakartaSans-Italic.ttf 178.9 KB · in bundle
-
PlusJakartaSans.ttf 172.2 KB · in bundle
-
-
-
references
-
defect-classes.md 11.6 KB
# Defect classes this skill fixes Everything below was found by actually breaking down a real Claude Design export (an 11-slide `.pptx`, generated under the hood by PptxGenJS) and comparing a LibreOffice-headless PDF render of the original against a render of a repaired copy, slide by slide. Three of these were hypotheses confirmed by inspection; the fourth (or in this case a lot of the actual visible damage) only showed up once real before/after PDF renders were compared side by side — worth remembering if you're extending this skill: "the shape's box looks wrong" and "the export looks wrong" are related but not the same claim, and only the second one is what the user actually sees. ## 1. Corrupt package structure ([Content_Types].xml) **What's wrong:** A `.pptx` is a zip of XML parts. `[Content_Types].xml` is the manifest — it declares which parts exist and what content type each one is. The sample export declared `Override` entries for ten `ppt/slideMasters/slideMasterN.xml` parts (N = 2..11) that were never actually written into the zip. Only `slideMaster1.xml` exists; every slide in the deck correctly points at it via its own `.rels`. The other ten declarations are pure manifest noise — nothing in the package references them, they just shouldn't be there. **Why it matters:** PowerPoint is forgiving about this and may quietly "repair" the file on open without telling you. Stricter OOXML consumers are not guaranteed to be — a strict XSD validator, a from-scratch parser in a cloud pptx→pdf API, or an OOXML library with less defensive coding than what ships in Office, can choke on a manifest that promises parts that aren't there. This is exactly the kind of thing that produces a failure with no obvious connection to its cause ("conversion failed" with no further detail, or slides silently dropped). **Detection:** `content_types.py::audit_content_types()` parses every `PartName="..."` in the manifest and checks each one against the zip's actual file list. **Fix:** Nothing bespoke needed. `python-pptx` builds its in-memory package graph strictly from parts it can load and the relationships that actually connect them; `Presentation.save()` serializes a fresh `[Content_Types].xml` from that graph. A plain load-then-save round trip already drops any dangling `Override` that pointed at nothing. Confirmed: running the sample through `Presentation(path).save(out)` with zero other changes reduced the declared-parts count from 42 to 32, and all 10 phantom `slideMasterN.xml` entries were gone. `fix_export.py` re-audits the saved file afterward and reports before/after counts so this isn't just assumed. ## 2. Off-canvas shape geometry **What's wrong:** Several shapes per slide — specifically the small tracked-out "kicker" labels (page footers, section eyebrows) — have a left offset + width whose sum exceeds the slide's own width, by a consistent ~731,565 EMU (~0.8in) in the sample. The box is defined wider than the canvas it's supposed to sit on. **Why it matters in general:** PowerPoint's on-screen renderer doesn't clip a shape at the slide edge; the overflow just draws into space that never gets shown. Renderers that treat the slide as a fixed, precisely- bounded viewport (which is closer to how a browser or a strict rasterizer thinks about a "page") are not guaranteed to be as forgiving. This is a real, worth-fixing structural defect independent of any specific renderer's quirks — it's just wrong for a shape to be wider than the canvas it's drawn on, and a different PDF pipeline than the one this skill happens to test against could clip it for real. **A caveat worth being honest about:** in the sample deck specifically, this defect turned out *not* to be the thing actually causing visible clipping in LibreOffice's render (see #4 below — a different bug was responsible for essentially all of the visible damage, since these kicker boxes were ~19.8in wide, so a 0.8in overflow past a 20in-wide slide left an enormous, unused margin — the short label text inside never got anywhere near either edge). This skill fixes it anyway, because (a) it's a genuine defect that could bite in a different rendering pipeline even though it didn't bite here, and (b) "shape exceeds its canvas" is the general form of the bug — the specific box that happened to overflow in the sample is incidental, not the point. **Detection & fix:** `geometry.py`. For every top-level shape on every slide, compute the effective right/bottom edge and compare against `prs.slide_width` / `prs.slide_height`. Translate the shape back inside the slide when its own size allows it (the safe, zero-distortion fix — this is what fired on every affected shape in the sample: same size, just slid left). Fall back to a proportional scale-down (with table columns/ rows scaled to match, so a table's grid stays consistent with its frame) only when the shape is genuinely larger than the slide itself. Shapes nested inside a group are flagged for manual review rather than auto-fixed — see the module docstring in `geometry.py` for why translating a group child's local coordinates safely is a harder problem than it looks, and getting it wrong is worse than leaving a rare case alone. ## 3. Missing font embedding **What's wrong:** The deck references three non-system webfonts — Inter, Plus Jakarta Sans, JetBrains Mono — via plain `typeface="..."` attributes, with zero embedded font data anywhere in the package (no `<p:embeddedFontLst>`). That's normal for a file authored somewhere those fonts are already installed (a browser, a design tool, a machine with the Office add-ins that ship them). It's a problem the moment something *else* — a server, a CI box, a cloud converter — tries to render the file without those fonts present: the renderer picks a fallback (Calibri / Liberation Sans / whatever it has) with different glyph widths, which changes text measurements throughout the deck. **Detection:** `fonts.py::scan_typefaces()` scans every part's XML for `<a:latin typeface="...">` — deliberately *only* `<a:latin>`, not every `typeface="..."` in the file. Every OOXML theme also carries a boilerplate list of ~10 per-script fallback fonts (`<a:font script="Jpan" typeface="..."/>`, `script="Hang"`, `"Thai"`, `"Arab"`, etc., inside `majorFont`/`minorFont`) that only matter if the deck contains text in that script. Matching every `typeface="..."` attribute (an earlier version of this scan did) pulled in ~40 CJK/Indic/Southeast-Asian fonts nothing on an English-language deck ever renders with — pure noise that would have sent this skill off trying to fetch fonts nobody needed. **Fix:** `fonts.py::ensure_fonts_available()`. Classifies every font found into: safe (assumed present everywhere — Calibri, Arial, etc.), bundled (this skill ships actual OFL-licensed TTFs for Inter / Plus Jakarta Sans / JetBrains Mono in `assets/fonts/`, since these are the ones observed in practice), or unknown (attempt a best-effort fetch from the open `google/fonts` OFL mirror; if that fails — no network, or the font isn't on Google Fonts — report it plainly as "will render with a fallback" rather than pretending it was handled). Fonts get installed to `/usr/local/share/fonts/design-export-repair/` (falling back to the user's XDG font dir if that's not writable) — a directory fontconfig already scans by default on essentially every Linux box — followed by a global `fc-cache -f`. An earlier version of this tried to point `FONTCONFIG_PATH` at an arbitrary temp directory, which does nothing useful: that variable controls where fontconfig looks for *configuration*, not where it looks for *font files*. Confirmed via `fc-list` before/after: the bundled fonts were invisible to the system until they landed in a real scanned font directory, at which point LibreOffice picked them up with zero other changes. ## 4. LibreOffice clips text runs with letter-spacing (`a:rPr spc`) **What's wrong, and how it was found:** After fixing #1–#3, the repaired PDF *still* showed the exact same visible symptom the user reported — short, tracked-out uppercase labels missing their last 1–4 characters ("HOW WE WORK" → "HOW WE WORI", "ONGOING" → "ONGOIN", "MOST POPULAR" → "MOST POPUL", "SCOPED PER ENGAGEMENT" → "SCOPED PER ENGAGEME"). Since the containing boxes were confirmed enormous (~19.8in wide, see #2) and the fonts were confirmed installed and in use (#3), neither of those could be the cause. A direct A/B test settled it: take an *otherwise completely untouched* copy of the original broken slide, strip only the `spc="NN"` attribute (character tracking/letter-spacing, in hundredths of a point) from every run, and reconvert. Every previously-clipped label rendered in full, with no other change made. `spc` is the trigger. **Why it matters:** this is a LibreOffice text-shaping/clip-region bug, not a document defect — the deck is authored correctly (the box has 15+ inches of unused room; there is no legitimate reason for the text to be cut). Something in LibreOffice's handling of `a:rPr spc` computes a clip region that doesn't match where it actually places the tracked-out glyphs, and the tail end of the run gets cut. It reproduces regardless of `normAutofit`, box width, anchor, or font, as long as `spc` is set on the run — meaning **this is the dominant cause of the "text gets cut off" complaint** in decks that use tracked-out labels (a very common style choice for eyebrows/kickers/badges in exactly the kind of generated deck this skill targets), and defect #2's shape-boundary fix does not touch it at all. **Fix:** There's no flag to tell LibreOffice to render `spc` correctly, so `render_workaround.py::make_render_copy()` builds a throwaway copy of the repaired deck with every `spc="NN"` attribute stripped from slide XML, and *that* copy — not the deliverable `.pptx` — is what gets fed to `soffice --convert-to pdf`. The actual repaired `.pptx` this skill hands back keeps its original letter-spacing untouched, because PowerPoint, Keynote, and Google Slides don't have this bug — the design's intended tracking is exactly right for anyone opening the file in one of those. The trade-off (tracked-out labels render at normal spacing in the PDF only, not in the editable deck) is real and is called out explicitly in the repair report rather than changed silently. If you're extending this skill and hit a case where this workaround doesn't fully resolve clipping, don't reach straight for "strip more attributes" — check whether the specific LibreOffice version in use has fixed this upstream (it's the kind of thing that gets patched), and consider re-running the A/B test (strip one attribute at a time from a known-broken slide, reconvert, compare) rather than assuming the same root cause. ## What this means for a deck you haven't seen before Defects #1 and #3 are cheap, general, and safe to always run — auditing a manifest and making sure referenced fonts exist can't make a correct file worse. Defect #2's shape-boundary fix is also always safe (translate-first never distorts anything) but, per the caveat above, don't assume fixing it is what makes an export look right — verify against an actual rendered PDF, not just the shape geometry, because #4 or something like it may be the real cause of what you're looking at. If you're diagnosing a new "broken export" complaint this skill's checks don't catch, the fastest path is the same one that found #4: convert an untouched copy to PDF, form a hypothesis about what's different between how it looks and how it should look, strip/change exactly that one thing in a scratch copy, reconvert, and compare. Don't guess from the XML alone — LibreOffice's actual rendering behavior is the ground truth for what an export-to-PDF user will see, and it doesn't always match what the OOXML looks like it should do.
-
-
scripts
-
content_types.py 3.2 KB
""" OOXML package structure repair (defect class 1 — see references/defect-classes.md). A .pptx is a zip of XML parts, and [Content_Types].xml is the manifest that declares which parts exist and what content type each one is. It is possible — and, empirically, something Claude Design's export path sometimes does — for that manifest to declare Override entries for parts that were never actually written into the zip (observed: ten slideMasterN.xml declarations with no corresponding files on disk). PowerPoint tolerates this and may quietly "repair" it on open. Stricter OOXML consumers (LibreOffice headless, many cloud pptx->pdf converters) are not guaranteed to be as forgiving, so this can turn into a failed conversion, dropped slides, or garbled output somewhere downstream — with no obvious error pointing back at the real cause. The reliable fix turns out to be simple: python-pptx builds its in-memory package graph strictly from parts it can actually load and the relationships that actually connect them, and Presentation.save() then serializes [Content_Types].xml fresh from that graph — so a normal load-then-save round trip through python-pptx already drops any dangling Override that pointed at a part which was never there to begin with. audit_content_types() exists so the repair report can say precisely what was wrong and confirm it's gone, rather than silently trusting that the round trip worked. """ import re import zipfile from dataclasses import dataclass from pathlib import Path @dataclass class ContentTypesAudit: declared_parts: int missing_parts: list # PartNames declared in [Content_Types].xml but absent from the zip undeclared_parts: list # real XML parts present in the zip with no Content_Types coverage at all _ALWAYS_COVERED_BY_DEFAULT = {".rels"} # covered by the <Default Extension="rels".../> entry, not an Override def audit_content_types(pptx_path: Path) -> ContentTypesAudit: with zipfile.ZipFile(pptx_path) as zf: names = set(zf.namelist()) try: ct_xml = zf.read("[Content_Types].xml").decode("utf-8", errors="ignore") except KeyError: # No content types part at all is a much more serious problem than # this skill tries to auto-fix; surface it plainly instead of guessing. return ContentTypesAudit(declared_parts=0, missing_parts=["[Content_Types].xml itself is missing"], undeclared_parts=[]) declared = set(re.findall(r'PartName="([^"]+)"', ct_xml)) default_exts = set(re.findall(r'Default Extension="([^"]+)"', ct_xml)) missing = sorted(p for p in declared if p.lstrip("/") not in names) undeclared = [] for n in names: if n in ("[Content_Types].xml",) or n.endswith("/"): continue part_name = "/" + n ext = n.rsplit(".", 1)[-1] if "." in n else "" if part_name in declared: continue if ext in default_exts: continue # covered by a Default Extension entry undeclared.append(part_name) return ContentTypesAudit(declared_parts=len(declared), missing_parts=missing, undeclared_parts=sorted(undeclared)) -
convert_html.py 3.1 KB
""" Best-effort fallback for a Claude Design export that comes as an HTML/CSS/JS bundle (index.html + assets) rather than a .pptx. This path is NOT backed by a real defect investigation the way the .pptx path is (repair_pptx.py + geometry.py + content_types.py were built against an actual broken export and verified to fix it) — it exists so a zip full of HTML doesn't just fail outright, applying the generic fixes that most commonly break an HTML-deck-to-PDF print: - `overflow: hidden` on a slide/page container silently clips content that would otherwise just scroll on screen — fine in a browser, fatal the moment you print, because print has no scrollbar. This is the HTML equivalent of defect class 2 (off-canvas shapes) in the pptx path. - Browsers only print backgrounds/colors when explicitly told to (print-color-adjust / -webkit-print-color-adjust: exact); without it a dark-themed deck can print with a white background and invisible light-colored text. - `@page` size defaults to the browser's page setup, not the deck's actual aspect ratio, unless the page explicitly sets one. If you hit a real broken HTML export, treat this script as a starting point to extend rather than a finished, battle-tested pipeline the way the .pptx path is. """ import re from pathlib import Path PRINT_FIX_CSS = """ <style id="design-export-repair-print-fix"> * { -webkit-print-color-adjust: exact !important; print-color-adjust: exact !important; color-adjust: exact !important; } html, body { overflow: visible !important; } [class*="slide" i], [class*="page" i], [id*="slide" i] { overflow: visible !important; page-break-inside: avoid; break-inside: avoid; } </style> """ def _detect_page_size(html_text: str) -> tuple: """Look for an explicit slide/page dimension hint in the markup (common in generated decks: a data attribute, inline width/height on the root slide container, or a CSS custom property). Falls back to a standard 16:9 slide size in inches if nothing is found.""" m = re.search(r'width["\']?\s*[:=]\s*["\']?(\d{3,5})px["\']?[^}]*height["\']?\s*[:=]\s*["\']?(\d{3,5})px', html_text) if m: w_px, h_px = int(m.group(1)), int(m.group(2)) return (w_px / 96, h_px / 96) # 96 CSS px/in return (13.333, 7.5) # standard 16:9 slide, in inches def convert_html_to_pdf(html_path: Path, out_pdf_path: Path, fonts_dir: Path | None = None) -> Path: from playwright.sync_api import sync_playwright html_text = html_path.read_text(encoding="utf-8", errors="ignore") width_in, height_in = _detect_page_size(html_text) with sync_playwright() as p: browser = p.chromium.launch() page = browser.new_page() page.goto(html_path.resolve().as_uri()) page.add_style_tag(content=PRINT_FIX_CSS) page.wait_for_timeout(300) # let webfonts/late layout settle page.pdf( path=str(out_pdf_path), width=f"{width_in}in", height=f"{height_in}in", print_background=True, margin={"top": "0in", "bottom": "0in", "left": "0in", "right": "0in"}, ) browser.close() return out_pdf_path -
fix_export.py 12.4 KB
#!/usr/bin/env python3 """ Main entry point for the design-export-repair skill. Usage: python3 fix_export.py <input> [--out-dir DIR] [--no-pdf] [--no-network] <input> can be: - a .pptx exported from Claude Design (this is what's been seen in practice — Design's export sometimes comes through as a bare .pptx, not always wrapped in a zip) - a .zip containing a .pptx, or an HTML/CSS/JS deck bundle, or a PDF - a bare .pdf (repair is limited without the source — see below) What it does for a .pptx (the fully-verified path — see references/defect-classes.md for how each of these was confirmed against a real broken export): 1. Audits [Content_Types].xml against what's actually in the zip. 2. Fixes any shape whose box extends past the slide edge. 3. Scans every typeface reference and makes sure the renderer has matching font files instead of silently substituting one. 4. Re-saves the repaired .pptx (this also naturally clears any dangling Content_Types entries — python-pptx rebuilds that manifest from its own part graph on save). 5. Converts the repaired deck to PDF with LibreOffice headless. 6. Validates the resulting PDF: page count, blank pages, anything still touching a page edge. 7. Writes a report (JSON + Markdown) describing exactly what changed and why, so nothing is fixed silently. Exits non-zero only if something genuinely could not be completed (e.g. LibreOffice conversion failed outright); font/geometry issues that were fixed, or fonts that fell back to a substitute, are reported but do not fail the run. """ import argparse import dataclasses import json import sys from pathlib import Path sys.path.insert(0, str(Path(__file__).resolve().parent)) import content_types import fonts as fonts_mod import geometry import render_workaround import soffice import unpack import validate_pdf as validate_pdf_mod def _asdict_list(items): return [dataclasses.asdict(i) for i in items] def repair_pptx(pptx_path: Path, out_dir: Path, allow_network_fetch: bool) -> dict: from pptx import Presentation report: dict = {"input_kind": "pptx", "source_path": str(pptx_path)} # --- defect 1: audit package structure (before) --- before_audit = content_types.audit_content_types(pptx_path) report["content_types_before"] = dataclasses.asdict(before_audit) # --- open, fix geometry (defect 2) --- prs = Presentation(str(pptx_path)) slide_count = len(prs.slides) changes, flags = geometry.fix_presentation_geometry(prs) report["geometry_changes"] = _asdict_list(changes) report["geometry_flags_for_manual_review"] = _asdict_list(flags) # --- save repaired pptx (this also clears dangling Content_Types entries) --- out_dir.mkdir(parents=True, exist_ok=True) repaired_pptx_path = out_dir / (pptx_path.stem.replace(".repaired", "") + ".repaired.pptx") prs.save(str(repaired_pptx_path)) report["repaired_pptx_path"] = str(repaired_pptx_path) after_audit = content_types.audit_content_types(repaired_pptx_path) report["content_types_after"] = dataclasses.asdict(after_audit) report["content_types_fixed"] = ( len(before_audit.missing_parts) > 0 and len(after_audit.missing_parts) == 0 ) # --- defect 3: fonts --- import tempfile import zipfile as zf_mod scan_dir = Path(tempfile.mkdtemp(prefix="der_fontscan_")) with zf_mod.ZipFile(repaired_pptx_path) as zf: zf.extractall(scan_dir) families = fonts_mod.scan_typefaces(scan_dir) work_fonts_dir = out_dir / "_fonts" font_status = fonts_mod.ensure_fonts_available(families, work_fonts_dir, allow_network_fetch=allow_network_fetch) report["fonts_used"] = sorted(families) report["font_status"] = font_status # --- convert to PDF --- # Render a letter-spacing-neutralized copy, not the repaired pptx # itself — see render_workaround.py for why. The returned .pptx keeps # its original spc values; only this throwaway copy is altered. render_copy_path, spc_removed = render_workaround.make_render_copy(repaired_pptx_path) report["letter_spacing_workaround"] = { "attributes_neutralized_for_pdf_only": spc_removed, "note": ( "LibreOffice clips the trailing characters of any text run with a:rPr spc set " "(letter-spacing/tracking), regardless of how much room the containing shape has. " "Neutralized for the PDF render only; the repaired .pptx below keeps the original " "letter-spacing intact for editing in PowerPoint/Keynote/Google Slides." ) if spc_removed else "No letter-spacing attributes found; workaround was a no-op.", } pdf_path = out_dir / (repaired_pptx_path.stem + ".pdf") try: produced = soffice.convert_to_pdf(str(render_copy_path), str(out_dir), fonts_dir=str(work_fonts_dir)) produced.rename(pdf_path) report["pdf_path"] = str(pdf_path) report["pdf_conversion_ok"] = True except RuntimeError as exc: report["pdf_conversion_ok"] = False report["pdf_conversion_error"] = str(exc) return report finally: render_copy_path.unlink(missing_ok=True) # --- validate output --- validation = validate_pdf_mod.validate_pdf(str(pdf_path), expected_page_count=slide_count) report["validation"] = dataclasses.asdict(validation) return report def repair_html(resolved: unpack.ResolvedInput, out_dir: Path) -> dict: import convert_html out_dir.mkdir(parents=True, exist_ok=True) report = {"input_kind": "html", "source_note": resolved.source_note, "structural_repair": "not_applicable_best_effort_path"} pdf_path = out_dir / (resolved.primary_path.stem + ".pdf") try: convert_html.convert_html_to_pdf(resolved.primary_path, pdf_path) report["pdf_path"] = str(pdf_path) report["pdf_conversion_ok"] = True validation = validate_pdf_mod.validate_pdf(str(pdf_path), expected_page_count=validate_pdf_mod.fitz.open(str(pdf_path)).page_count) report["validation"] = dataclasses.asdict(validation) except Exception as exc: # noqa: BLE001 - surface any failure in the report rather than crashing report["pdf_conversion_ok"] = False report["pdf_conversion_error"] = str(exc) return report def pass_through_pdf(resolved: unpack.ResolvedInput, out_dir: Path) -> dict: import shutil out_dir.mkdir(parents=True, exist_ok=True) dest = out_dir / resolved.primary_path.name shutil.copy2(resolved.primary_path, dest) validation = validate_pdf_mod.validate_pdf(str(dest), expected_page_count=validate_pdf_mod.fitz.open(str(dest)).page_count) return { "input_kind": "pdf", "source_note": resolved.source_note, "structural_repair": "not_possible_no_source_file", "pdf_path": str(dest), "validation": dataclasses.asdict(validation), } def render_markdown_report(report: dict) -> str: lines = ["# Design export repair report", ""] kind = report.get("input_kind", "unknown") lines.append(f"**Input type detected:** `{kind}`") if report.get("source_note"): lines.append(f"**Note:** {report['source_note']}") lines.append("") if kind == "pptx": ct_before = report.get("content_types_before", {}) ct_after = report.get("content_types_after", {}) lines.append("## 1. Package structure ([Content_Types].xml)") if ct_before.get("missing_parts"): lines.append(f"- Found {len(ct_before['missing_parts'])} declared part(s) with no matching file in the archive:") for p in ct_before["missing_parts"]: lines.append(f" - `{p}`") lines.append(f"- After repair: {'fixed — manifest now matches the actual archive contents' if report.get('content_types_fixed') else 'STILL PRESENT — needs manual investigation'}") else: lines.append("- No dangling part declarations found. Package structure was already consistent.") lines.append("") lines.append("## 2. Off-canvas shapes") changes = report.get("geometry_changes", []) flags = report.get("geometry_flags_for_manual_review", []) if changes: lines.append(f"- Repositioned/resized {len(changes)} shape(s) that extended past the slide boundary:") for c in changes: lines.append(f" - Slide {c['slide_index']}, \"{c['shape_name']}\" ({c['strategy']}): {c['before']} → {c['after']}") if c.get("note"): lines.append(f" {c['note']}") else: lines.append("- No shapes were found extending past the slide boundary.") if flags: lines.append(f"- {len(flags)} shape(s) inside groups were flagged for manual review (not auto-fixed, see references/defect-classes.md):") for f in flags: lines.append(f" - Slide {f['slide_index']}: {f['shape_name']}") lines.append("") lines.append("## 3. Fonts") font_status = report.get("font_status", {}) if font_status: for name, info in font_status.items(): lines.append(f"- **{name}**: {info['status']} — {info['detail']}") else: lines.append("- No custom fonts detected.") lines.append("") lines.append("## 4. Letter-spacing / LibreOffice text clipping") spc = report.get("letter_spacing_workaround", {}) if spc.get("attributes_neutralized_for_pdf_only"): lines.append(f"- {spc['attributes_neutralized_for_pdf_only']} tracked/letter-spaced text run(s) found. {spc['note']}") else: lines.append(f"- {spc.get('note', 'No letter-spacing attributes found.')}") lines.append("") lines.append("## Output") if report.get("pdf_conversion_ok"): lines.append(f"- PDF: `{report.get('pdf_path')}`") else: lines.append(f"- PDF conversion FAILED: {report.get('pdf_conversion_error', 'unknown error')}") if report.get("repaired_pptx_path"): lines.append(f"- Repaired PPTX: `{report['repaired_pptx_path']}`") validation = report.get("validation") if validation: lines.append("") lines.append("## PDF validation") lines.append(f"- Page count: {validation['page_count']} (expected {validation['expected_page_count']}) — {'OK' if validation['page_count_ok'] else 'MISMATCH'}") if validation.get("blank_pages"): lines.append(f"- Possibly-blank pages (verify these are intentional, e.g. section dividers): {validation['blank_pages']}") if validation.get("edge_overflow"): lines.append(f"- {len(validation['edge_overflow'])} text block(s) still touch/cross a page edge:") for o in validation["edge_overflow"]: lines.append(f" - Page {o['page']}: \"{o['text']}\"") else: lines.append("- No text blocks touch or cross a page edge.") return "\n".join(lines) + "\n" def main(): parser = argparse.ArgumentParser(description="Repair a Claude Design export and produce a clean PDF/PPTX.") parser.add_argument("input", help="Path to the export: .zip, .pptx, or .pdf") parser.add_argument("--out-dir", default="./design-export-repair-output", help="Where to write outputs") parser.add_argument("--no-network", action="store_true", help="Don't attempt to fetch missing fonts from the web") args = parser.parse_args() out_dir = Path(args.out_dir).resolve() out_dir.mkdir(parents=True, exist_ok=True) work_dir = out_dir / "_work" work_dir.mkdir(parents=True, exist_ok=True) resolved = unpack.resolve_input(args.input, work_dir) print(f"Detected input kind: {resolved.kind}\n{resolved.source_note}") if resolved.kind == "pptx": report = repair_pptx(resolved.primary_path, out_dir, allow_network_fetch=not args.no_network) elif resolved.kind == "html": report = repair_html(resolved, out_dir) elif resolved.kind == "pdf": report = pass_through_pdf(resolved, out_dir) else: print(f"ERROR: could not identify a repairable payload in {args.input}. {resolved.source_note}", file=sys.stderr) sys.exit(2) report["source_note"] = report.get("source_note", resolved.source_note) (out_dir / "repair_report.json").write_text(json.dumps(report, indent=2, default=str)) md = render_markdown_report(report) (out_dir / "repair_report.md").write_text(md) print("\n" + md) if report.get("pdf_conversion_ok") is False: sys.exit(1) if __name__ == "__main__": main() -
fonts.py 9.9 KB
""" Font availability repair (defect class 3 — see references/defect-classes.md). Claude Design decks routinely reference webfonts (Inter, Plus Jakarta Sans, JetBrains Mono are the ones observed in practice) purely by name, with no font data embedded in the .pptx. That's normal for a file meant to be opened in an app that already has those fonts. It becomes a problem the moment something *else* renders the file to a PDF/image — LibreOffice, a cloud converter, a CI box — because that renderer almost certainly does not have "Plus Jakarta Sans" installed and will silently substitute something else. The substitute has different metrics, so it doesn't just look wrong: it also changes text wrapping, which can turn a borderline shape (see geometry.py) from "fine" into "overflowing." This module makes sure the fonts a deck actually uses are installed where the PDF renderer will look for them, so the conversion step in convert.py produces something that matches the design instead of a best-effort substitute. """ import re import shutil import subprocess from pathlib import Path SKILL_ROOT = Path(__file__).resolve().parent.parent BUNDLED_FONTS_DIR = SKILL_ROOT / "assets" / "fonts" # Fonts that are safe to assume are present on basically any renderer # (they ship with LibreOffice / are standard Office/Core fonts). Anything # not in this set gets treated as "needs to be made available." SAFE_FONTS = { "calibri", "calibri light", "arial", "times new roman", "cambria", "cambria math", "segoe ui", "verdana", "georgia", "courier new", "helvetica", "liberation sans", "liberation serif", "liberation mono", "dejavu sans", "dejavu serif", "symbol", "wingdings", } # Maps a lowercased family name to the bundled TTF(s) that cover it. Keep # this in sync with assets/fonts/ — add a family here whenever you drop in # a new bundled font so ensure_fonts_available() picks it up automatically. BUNDLED_FAMILIES = { "inter": ["Inter.ttf"], "plus jakarta sans": ["PlusJakartaSans.ttf", "PlusJakartaSans-Italic.ttf"], "jetbrains mono": ["JetBrainsMono-Regular.ttf"], } # Placeholder/theme-reference strings that show up in typeface="" attributes # but aren't real font names (OOXML theme font slots) or are empty. NOT_A_FONT_NAME = {"", "+mj-lt", "+mn-lt", "+mj-ea", "+mn-ea", "+mj-cs", "+mn-cs"} def scan_typefaces(unpacked_pptx_dir: Path) -> set[str]: """Return the set of distinct Latin-script font family names actually used for rendering anywhere in an unpacked .pptx (theme major/minor font, and every a:latin typeface on rPr/defRPr/endParaRPr across masters, layouts, and slides). Deliberately scoped to <a:latin> only, not every typeface="..." in the file: every OOXML theme also carries a boilerplate list of ~10 per-script fallback fonts (<a:font script="Jpan" typeface="..."/>, script="Hang", "Thai", "Arab", etc., inside majorFont/minorFont) that are only used if the deck actually contains text in that script. For an English-language deck those are pure noise — matching them would make this skill "fix" font availability for a couple dozen CJK/Indic/ Southeast-Asian fonts nothing on the slide ever renders with. a:ea and a:cs (east-asian / complex-script) typeface overrides on individual runs are skipped for the same reason: they only matter for text in those scripts, and generators commonly set them to the same value as a:latin out of habit even on plain English runs. """ families: set[str] = set() latin_re = re.compile(r'<a:latin\b[^>]*?\btypeface="([^"]*)"') for xml_file in unpacked_pptx_dir.rglob("*.xml"): try: text = xml_file.read_text(encoding="utf-8", errors="ignore") except OSError: continue for m in latin_re.finditer(text): name = m.group(1).strip() if name and name not in NOT_A_FONT_NAME: families.add(name) return families def classify_fonts(families: set[str]) -> dict: """Split the fonts a deck uses into: already safe, covered by a bundled TTF, or unknown (will fall back silently unless fetched from the web).""" safe, bundled, unknown = [], [], [] for name in sorted(families): key = name.lower() if key in SAFE_FONTS: safe.append(name) elif key in BUNDLED_FAMILIES: bundled.append(name) else: unknown.append(name) return {"safe": safe, "bundled": bundled, "unknown": unknown} def _try_fetch_from_google_fonts(family: str, dest_dir: Path) -> list[str]: """Best-effort: fetch a family Claude didn't ship a copy of, using the open google/fonts OFL mirror. Network may not be available wherever this skill runs, so failure here is expected and non-fatal — the caller just ends up in the same place as if this function didn't exist: a logged warning instead of a silent substitution.""" import urllib.request import urllib.parse # Only worth attempting for plain ASCII family names — a font family # name containing non-Latin characters is never going to be a Google # Fonts slug, and building a URL from raw non-ASCII text throws deep # inside http.client rather than failing cleanly. if not family.isascii(): return [] slug = family.lower().replace(" ", "") compact = family.replace(" ", "") candidates = [ f"https://raw.githubusercontent.com/google/fonts/main/ofl/{slug}/{urllib.parse.quote(compact)}%5Bwght%5D.ttf", f"https://raw.githubusercontent.com/google/fonts/main/ofl/{slug}/{urllib.parse.quote(compact)}-Regular.ttf", ] saved = [] for url in candidates: try: dest = dest_dir / f"{compact}.ttf" urllib.request.urlretrieve(url, dest) if dest.stat().st_size > 1024: # got something real, not an error page saved.append(str(dest)) break dest.unlink(missing_ok=True) except Exception: # noqa: BLE001 - best-effort network fetch, any failure just falls through continue return saved def _default_font_install_dir() -> Path: """Pick a directory fontconfig already scans by default, so a plain `fc-cache -f` is enough to make installed fonts visible — no FONTCONFIG_PATH/FONTCONFIG_FILE trickery, which affects *config* lookup, not *font* lookup, and silently does nothing useful here. `/usr/local/share/fonts` is in the default <dir> list on effectively every Linux fontconfig config (confirmed via `fc-match`/fonts.conf at build time) and doesn't depend on which user/HOME the renderer runs as. Falls back to a per-user XDG font dir if that path isn't writable (e.g. running as a non-root user).""" import os candidate = Path("/usr/local/share/fonts/design-export-repair") try: candidate.mkdir(parents=True, exist_ok=True) probe = candidate / ".write_test" probe.touch() probe.unlink() return candidate except OSError: xdg = Path(os.environ.get("XDG_DATA_HOME", Path.home() / ".local" / "share")) fallback = xdg / "fonts" / "design-export-repair" fallback.mkdir(parents=True, exist_ok=True) return fallback def ensure_fonts_available(families: set[str], work_fonts_dir: Path | None = None, allow_network_fetch: bool = True) -> dict: """Install whatever fonts we can for the given family names into a directory fontconfig actually scans by default, then refresh the font cache so LibreOffice sees them. Returns a per-family status report to include in the repair report — this is meant to be visible to the user, not just logged, because "your PDF used a substitute font for X" is exactly the kind of silent breakage this skill exists to surface instead of hide. `work_fonts_dir` is accepted for backwards compatibility / explicit override but is no longer where the fonts need to end up for LibreOffice to find them — see _default_font_install_dir(). """ work_fonts_dir = _default_font_install_dir() classification = classify_fonts(families) status = {} for name in classification["safe"]: status[name] = {"status": "safe", "detail": "assumed present on any renderer"} for name in classification["bundled"]: key = name.lower() installed = [] for fname in BUNDLED_FAMILIES[key]: src = BUNDLED_FONTS_DIR / fname if src.is_file(): shutil.copy2(src, work_fonts_dir / fname) installed.append(fname) status[name] = { "status": "installed" if installed else "missing_bundled_file", "detail": f"copied {', '.join(installed)}" if installed else "expected bundled TTF not found on disk", } for name in classification["unknown"]: fetched = _try_fetch_from_google_fonts(name, work_fonts_dir) if allow_network_fetch else [] if fetched: status[name] = {"status": "fetched", "detail": f"downloaded {fetched[0]} from Google Fonts (OFL)"} else: status[name] = { "status": "fallback", "detail": ( "not bundled and could not be fetched — the PDF renderer will substitute " "a fallback font for this family, which may shift text position/wrapping" ), } _refresh_font_cache(work_fonts_dir) return status def _refresh_font_cache(fonts_dir: Path) -> None: try: # Refresh globally (-f forces it even if fontconfig thinks its cache # is current) rather than scoped to fonts_dir: fonts_dir is only # picked up at all once it's inside a directory fontconfig already # scans (see _default_font_install_dir), and a global refresh is # cheap and avoids any doubt about scoping. subprocess.run(["fc-cache", "-f"], capture_output=True, timeout=60) except (OSError, subprocess.TimeoutExpired): pass -
geometry.py 6.7 KB
""" Off-canvas shape geometry repair (defect class 2 — see references/defect-classes.md). PowerPoint's own on-screen renderer is forgiving about a shape whose box extends past the slide edge — it just draws the part that's off-canvas into empty space and nobody notices. Rasterizing renderers (what actually runs when you "export to PDF" through most non-PowerPoint pipelines) clip precisely at the slide boundary. So a shape that has *always* been 0.8in too wide only becomes visibly broken — text sliced off mid-word — at export time, which is exactly the "looked fine in the app, broken in the PDF" complaint this skill exists to fix. Strategy, cheapest-safest first: 1. Translate only. If the shape's own size is <= the slide's size in that dimension, sliding it back inside the slide fixes the overflow with zero visual change to the shape itself — no resizing, no distortion, nothing for autofit to recompute. 2. Scale down (fallback). Only reached when the shape is simply larger than the slide in some dimension (rare — usually a copy/paste from a different-sized template). Scales width and height by the same factor so aspect ratio holds, and if the shape is a table, scales every column width / row height by that same factor so the table stays internally consistent (a table's rendered width comes from its <a:tblGrid> column widths, not from the graphicFrame's own extent, so resizing one without the other leaves a table box that doesn't match its contents). Shapes nested inside a group are flagged for manual review rather than auto-fixed: python-pptx exposes child-shape offsets in the group's own child coordinate space, and correctly mapping that back to slide-absolute coordinates (and then writing a corrected value back through the group's chOff/chExt transform) is easy to get subtly wrong in a way that's worse than leaving a rare, already-cosmetic issue alone. If you hit one, the report tells you exactly which slide/shape to open and nudge by hand. """ from dataclasses import dataclass, field TOLERANCE_EMU = 3175 # ~0.0035in / ~1/3 pt — filters out rounding noise, not real overflow @dataclass class GeometryChange: slide_index: int shape_name: str shape_type: str strategy: str before: tuple after: tuple note: str = "" @dataclass class GeometryFlag: slide_index: int shape_name: str reason: str def _overflow(left, top, width, height, slide_w, slide_h): over_right = (left + width) - slide_w over_bottom = (top + height) - slide_h over_left = -left over_top = -top return over_right, over_bottom, over_left, over_top def _scale_table(shape, factor: float) -> None: """Scale a table's column widths and row heights by `factor` so the table's internal grid stays consistent with its resized frame.""" tbl = shape.table for col in tbl.columns: col.width = int(col.width * factor) for row in tbl.rows: row.height = int(row.height * factor) def fix_slide_geometry(slide, slide_index: int, slide_w: int, slide_h: int) -> tuple[list[GeometryChange], list[GeometryFlag]]: changes: list[GeometryChange] = [] flags: list[GeometryFlag] = [] for shape in slide.shapes: # Skip anything without a normal top-level position (placeholders that # inherit position from the layout report None here in python-pptx). if shape.left is None or shape.top is None or shape.width is None or shape.height is None: continue if shape.shape_type is not None and str(shape.shape_type) == "GROUP (6)": # See module docstring: intentionally not auto-fixed. for child in shape.shapes: flags.append(GeometryFlag( slide_index=slide_index, shape_name=f"{shape.name} > {getattr(child, 'name', '?')}", reason="shape is inside a group; skipped auto-fix, review position by hand", )) continue left, top, width, height = shape.left, shape.top, shape.width, shape.height over_right, over_bottom, over_left, over_top = _overflow(left, top, width, height, slide_w, slide_h) if max(over_right, over_bottom, over_left, over_top) <= TOLERANCE_EMU: continue # within tolerance, nothing to do before = (left, top, width, height) new_left, new_top, new_width, new_height = left, top, width, height strategy = "translate" note = "" # --- horizontal --- if width <= slide_w: if over_right > TOLERANCE_EMU: new_left = max(0, left - over_right) elif over_left > TOLERANCE_EMU: new_left = 0 else: strategy = "scale" # --- vertical --- if height <= slide_h: if over_bottom > TOLERANCE_EMU: new_top = max(0, top - over_bottom) elif over_top > TOLERANCE_EMU: new_top = 0 else: strategy = "scale" if strategy == "scale": factor = min(slide_w / width if width > slide_w else 1.0, slide_h / height if height > slide_h else 1.0) factor *= 0.98 # tiny safety margin so the scaled box clears the edge new_width = int(width * factor) new_height = int(height * factor) new_left = 0 new_top = 0 note = f"shape ({width}x{height} EMU) exceeded slide size ({slide_w}x{slide_h} EMU); scaled by {factor:.3f}" if shape.has_table: try: _scale_table(shape, factor) note += "; table columns/rows scaled to match" except Exception as exc: # pragma: no cover - defensive, table API can vary note += f"; WARNING could not scale table grid ({exc}) — table contents may not match frame" shape.left, shape.top, shape.width, shape.height = int(new_left), int(new_top), int(new_width), int(new_height) changes.append(GeometryChange( slide_index=slide_index, shape_name=shape.name or f"shape#{shape.shape_id}", shape_type=str(shape.shape_type), strategy=strategy, before=before, after=(shape.left, shape.top, shape.width, shape.height), note=note, )) return changes, flags def fix_presentation_geometry(prs) -> tuple[list[GeometryChange], list[GeometryFlag]]: all_changes: list[GeometryChange] = [] all_flags: list[GeometryFlag] = [] for i, slide in enumerate(prs.slides, start=1): changes, flags = fix_slide_geometry(slide, i, prs.slide_width, prs.slide_height) all_changes.extend(changes) all_flags.extend(flags) return all_changes, all_flags -
render_workaround.py 2.6 KB
""" LibreOffice-specific render workaround for letter-spacing clipping. Verified against the real sample export: any text run with an `<a:rPr spc="N">` (character tracking, in hundredths of a point — a deliberate, common styling choice for small tracked-out uppercase "kicker"/label/eyebrow text) gets its trailing 1-4 characters silently clipped by LibreOffice's headless PDF renderer, even when the containing shape's box is far wider than the text needs. Confirmed by A/B test: stripping `spc` from an otherwise-untouched copy of a real broken slide made every previously-clipped label render in full ("HOW WE WORI" -> "HOW WE WORK", "ONGOIN" -> "ONGOING", "MOST POPUL" -> "MOST POPULAR", etc.) This is independent of and *not* fixed by the shape-boundary repair in geometry.py — those boxes were already many times wider than their text, so this is a genuine LibreOffice text-layout/clip-rect bug tied to the spc attribute itself, not a sizing problem this skill can fix by resizing anything. There's no way to ask LibreOffice to render `spc` correctly here, so the fix is to render a spacing-neutralized COPY through LibreOffice, while leaving the actual returned .pptx untouched — the design's intended letter-spacing survives for anyone opening the deck in PowerPoint, Keynote, or Google Slides (none of which have this bug), and the PDF gets to be legible instead of visually broken. This is a real, visible trade-off (tracked-out labels render at normal spacing in the PDF only) and the repair report says so explicitly rather than changing it silently. """ import re import shutil import tempfile import zipfile from pathlib import Path _SPC_RE = re.compile(r'\s*spc="\d+"') def make_render_copy(pptx_path: Path) -> tuple[Path, int]: """Return (path_to_render_only_copy, number_of_spc_attributes_removed). The caller should feed the returned path to the PDF converter and discard it afterward — it is not a deliverable, only a rendering aid. """ pptx_path = Path(pptx_path) fd, tmp_name = tempfile.mkstemp(suffix=".render.pptx") render_path = Path(tmp_name) removed = 0 with zipfile.ZipFile(pptx_path) as src, zipfile.ZipFile(render_path, "w", zipfile.ZIP_DEFLATED) as out: for item in src.infolist(): data = src.read(item.filename) if item.filename.startswith("ppt/slides/slide") and item.filename.endswith(".xml"): text = data.decode("utf-8", errors="ignore") text, n = _SPC_RE.subn("", text) removed += n data = text.encode("utf-8") out.writestr(item, data) import os os.close(fd) return render_path, removed -
soffice.py 6.2 KB
""" Helper for running LibreOffice (soffice) headless in sandboxed environments where AF_UNIX sockets may be blocked, and where custom fonts need to be visible to the conversion. Adapted from the pattern used by Anthropic's bundled `pptx` skill (scripts/office/soffice.py) — same AF_UNIX shim, plus a `fonts_dir` argument so this skill's repaired decks render with the correct typefaces instead of silently falling back to Calibri/Arial. Call soffice through run_soffice(), not through subprocess directly: the shim and the per-run user profile both matter for headless reliability in a locked-down container. """ import contextlib import os import socket import subprocess import tempfile from collections.abc import Iterable from pathlib import Path def get_soffice_env() -> dict: env = os.environ.copy() env["SAL_USE_VCLPLUGIN"] = "svp" if _needs_shim(): shim = _ensure_shim() env["LD_PRELOAD"] = str(shim) return env def run_soffice(args: Iterable[str], **kwargs) -> subprocess.CompletedProcess: args = list(args) with contextlib.ExitStack() as stack: if not any(str(a).startswith("-env:UserInstallation") for a in args): profile = stack.enter_context( tempfile.TemporaryDirectory(prefix="lo_profile_", ignore_cleanup_errors=True) ) args = [f"-env:UserInstallation={Path(profile).as_uri()}"] + args return subprocess.run(["soffice"] + args, env=get_soffice_env(), **kwargs) def convert_to_pdf(input_path: str, out_dir: str, fonts_dir: str | None = None, timeout: int = 180) -> Path: """Convert a single office document to PDF via headless LibreOffice. `fonts_dir` is accepted for call-site compatibility but installing fonts is fonts.ensure_fonts_available()'s job — by the time this runs, the fonts a deck needs should already be sitting in a directory fontconfig scans by default (see fonts.py), and a plain `fc-cache -f` (already run there) is what makes soffice see them. There's nothing soffice-specific to configure here. Returns the path to the produced PDF. Raises RuntimeError on failure. """ input_path = Path(input_path).resolve() out_dir = Path(out_dir).resolve() out_dir.mkdir(parents=True, exist_ok=True) result = run_soffice( ["--headless", "--norestore", "--convert-to", "pdf", "--outdir", str(out_dir), str(input_path)], capture_output=True, text=True, timeout=timeout, ) produced = out_dir / (input_path.stem + ".pdf") if result.returncode != 0 or not produced.is_file(): raise RuntimeError( "soffice conversion failed " f"(exit {result.returncode}).\nstdout:\n{result.stdout}\nstderr:\n{result.stderr}" ) return produced _SHIM_SO = Path(tempfile.gettempdir()) / "der_socket_shim.so" def _needs_shim() -> bool: try: s = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM) s.close() return False except OSError: return True def _ensure_shim() -> Path: if _SHIM_SO.exists(): return _SHIM_SO src = Path(tempfile.gettempdir()) / "der_socket_shim.c" src.write_text(_SHIM_SOURCE) subprocess.run( ["gcc", "-shared", "-fPIC", "-o", str(_SHIM_SO), str(src), "-ldl"], check=True, capture_output=True, ) src.unlink() return _SHIM_SO _SHIM_SOURCE = r""" #define _GNU_SOURCE #include <dlfcn.h> #include <errno.h> #include <signal.h> #include <stdio.h> #include <stdlib.h> #include <sys/socket.h> #include <unistd.h> static int (*real_socket)(int, int, int); static int (*real_socketpair)(int, int, int, int[2]); static int (*real_listen)(int, int); static int (*real_accept)(int, struct sockaddr *, socklen_t *); static int (*real_close)(int); static int (*real_read)(int, void *, size_t); static int is_shimmed[1024]; static int peer_of[1024]; static int wake_r[1024]; static int wake_w[1024]; static int listener_fd = -1; __attribute__((constructor)) static void init(void) { real_socket = dlsym(RTLD_NEXT, "socket"); real_socketpair = dlsym(RTLD_NEXT, "socketpair"); real_listen = dlsym(RTLD_NEXT, "listen"); real_accept = dlsym(RTLD_NEXT, "accept"); real_close = dlsym(RTLD_NEXT, "close"); real_read = dlsym(RTLD_NEXT, "read"); for (int i = 0; i < 1024; i++) { peer_of[i] = -1; wake_r[i] = -1; wake_w[i] = -1; } } int socket(int domain, int type, int protocol) { if (domain == AF_UNIX) { int fd = real_socket(domain, type, protocol); if (fd >= 0) return fd; int sv[2]; if (real_socketpair(domain, type, protocol, sv) == 0) { if (sv[0] >= 0 && sv[0] < 1024) { is_shimmed[sv[0]] = 1; peer_of[sv[0]] = sv[1]; int wp[2]; if (pipe(wp) == 0) { wake_r[sv[0]] = wp[0]; wake_w[sv[0]] = wp[1]; } } return sv[0]; } errno = EPERM; return -1; } return real_socket(domain, type, protocol); } int listen(int sockfd, int backlog) { if (sockfd >= 0 && sockfd < 1024 && is_shimmed[sockfd]) { listener_fd = sockfd; return 0; } return real_listen(sockfd, backlog); } int accept(int sockfd, struct sockaddr *addr, socklen_t *addrlen) { if (sockfd >= 0 && sockfd < 1024 && is_shimmed[sockfd]) { if (wake_r[sockfd] >= 0) { char buf; real_read(wake_r[sockfd], &buf, 1); } errno = ECONNABORTED; return -1; } return real_accept(sockfd, addr, addrlen); } int close(int fd) { if (fd >= 0 && fd < 1024 && is_shimmed[fd]) { int was_listener = (fd == listener_fd); is_shimmed[fd] = 0; if (wake_w[fd] >= 0) { char c = 0; write(wake_w[fd], &c, 1); real_close(wake_w[fd]); wake_w[fd] = -1; } if (wake_r[fd] >= 0) { real_close(wake_r[fd]); wake_r[fd] = -1; } if (peer_of[fd] >= 0) { real_close(peer_of[fd]); peer_of[fd] = -1; } if (was_listener) _exit(0); } return real_close(fd); } """ if __name__ == "__main__": import sys result = run_soffice(sys.argv[1:]) sys.exit(result.returncode) -
unpack.py 8.6 KB
""" Figure out what a Claude Design export actually is before trying to fix it. Two genuinely different things reach this skill in practice: - A plain Export (PDF/PPTX/HTML) from a Design deck project. In practice this has turned out to be a bare .pptx (produced under the hood by PptxGenJS) as often as an actual .zip wrapper around one — this is the shape verified against a real broken export, and the fully-repaired path (content-types, geometry, fonts, spc workaround). - Whatever lands in the working directory from Design's "Send to Claude Code" handoff, which is a *different* feature from plain Export (it's aimed at importing a design into a codebase / continuing prototype work, per Claude's own docs — "import a design into your codebase... or let Claude build the whole thing"). Two variants, confirmed against Claude Code's own issue tracker (anthropics/ claude-code#51980, #69246): * "Send to Claude Code Web" opens a new claude.ai/code session with the design bundle already attached at the working directory — no file to locate at all; this skill (or whatever's already there) just needs pointing at the directory. * "Send to local coding agent" generates a copy/paste prompt that depends on a Claude Design MCP connector local Claude Code does not ship — per the linked issue this fails silently for most people. The dialog's "Download zip instead" option is the reliable path: it downloads an actual zip of the design files (a bundle of `*.dc.html` canvas files + a README), which this module also has to recognize. Either way, this module accepts a file OR a directory, plus a couple of other shapes the export might take, and always returns a small, explicit description of what it found rather than guessing silently. """ import zipfile from dataclasses import dataclass from pathlib import Path PPTX_EXTS = {".pptx", ".potx"} PDF_EXTS = {".pdf"} HTML_EXTS = {".html", ".htm"} DC_HTML_SUFFIX = ".dc.html" # Claude Design's own canvas file format, seen in its "Send to local coding agent" handoff prompt @dataclass class ResolvedInput: kind: str # "pptx" | "pdf" | "html" | "unknown" primary_path: Path extra_files: list # for html: the sibling assets (css/js/images) it needs source_note: str # human-readable explanation of what was found and where def _is_zip(path: Path) -> bool: try: return zipfile.is_zipfile(path) except OSError: return False def _looks_like_pptx(path: Path) -> bool: """A .pptx is itself a zip. zipfile.is_zipfile() is true for both a real .pptx and a zip bundle exported from Design, so distinguish them by whether the zip has the OOXML presentation part.""" try: with zipfile.ZipFile(path) as zf: names = zf.namelist() return "ppt/presentation.xml" in names or any(n.startswith("ppt/slides/") for n in names) except (zipfile.BadZipFile, OSError): return False def _scan_directory(directory: Path) -> ResolvedInput: """Search a directory (not a zip) for a repairable payload — this is what "Send to Claude Code Web" leaves you with: the design bundle already sitting in the working directory, no zip to unpack. Reuses the same priority order as the zip-extraction path: pptx first (the fully-repaired path), then Design's own .dc.html canvas format, then generic HTML, then a bare PDF.""" pptx_candidates = sorted(directory.rglob("*.pptx")) + sorted(directory.rglob("*.potx")) if pptx_candidates: chosen = max(pptx_candidates, key=lambda p: p.stat().st_size) note = f"found {len(pptx_candidates)} pptx file(s) in {directory}; using the largest: {chosen.relative_to(directory)}" return ResolvedInput("pptx", chosen, pptx_candidates, note) dc_html_candidates = sorted(directory.rglob(f"*{DC_HTML_SUFFIX}")) if dc_html_candidates: chosen = dc_html_candidates[0] siblings = [p for p in directory.rglob("*") if p.is_file() and p != chosen] note = ( f"found {len(dc_html_candidates)} Claude Design canvas file(s) (*.dc.html) in {directory}; " f"using {chosen.relative_to(directory)}. This is Design's native canvas format from its " "\"Send to Claude Code\" handoff, not an Export — the pptx repair path (content-types/geometry/" "fonts/letter-spacing) doesn't apply here; falling back to the best-effort HTML print-to-PDF path." ) return ResolvedInput("html", chosen, siblings, note) html_candidates = sorted(directory.rglob("index.html")) or sorted(directory.rglob("*.html")) if html_candidates: chosen = html_candidates[0] siblings = [p for p in directory.rglob("*") if p.is_file() and p != chosen] return ResolvedInput("html", chosen, siblings, f"found an HTML/CSS/JS bundle in {directory}; entry point: {chosen.relative_to(directory)}") pdf_candidates = sorted(directory.rglob("*.pdf")) if pdf_candidates: chosen = pdf_candidates[0] return ResolvedInput( "pdf", chosen, pdf_candidates, f"found a PDF with no editable source ({chosen.relative_to(directory)}) in {directory} — see note above about limited repair for flattened PDFs", ) found = [str(p.relative_to(directory)) for p in directory.rglob("*") if p.is_file()][:20] return ResolvedInput("unknown", directory, [], f"{directory} did not contain a recognizable pptx/dc.html/html/pdf payload; top files: {found}") def resolve_input(input_path: str, work_dir: Path) -> ResolvedInput: input_path = Path(input_path) if not input_path.exists(): raise FileNotFoundError(f"{input_path} does not exist") if input_path.is_dir(): # This is the shape "Send to Claude Code Web" leaves behind: no # zip, no single file to point at — just a working directory with # the design bundle already in it. return _scan_directory(input_path) suffix = input_path.suffix.lower() # Design's own canvas format (seen in its "Send to local coding agent" # handoff prompt: "Implement: <FILE>.dc.html") handed over as a loose # file rather than inside a directory/zip. if input_path.name.endswith(DC_HTML_SUFFIX): return ResolvedInput( "html", input_path, [], f"input is a Claude Design canvas file ({input_path.name}) from the \"Send to Claude Code\" handoff, " "not a plain Export — falling back to the best-effort HTML print-to-PDF path rather than the " "verified pptx repair path", ) # Case 1: bare .pptx handed directly (this is what Design's own export # button has produced in practice, not just a zip wrapper around one). if suffix in PPTX_EXTS or (suffix == "" and _looks_like_pptx(input_path)): if _looks_like_pptx(input_path): return ResolvedInput("pptx", input_path, [], f"input is already a .pptx: {input_path.name}") if suffix in PDF_EXTS: return ResolvedInput( "pdf", input_path, [], f"input is a bare PDF ({input_path.name}) with no editable source alongside it — " "structural repair (content-types / off-canvas shapes / font substitution) needs the " "original .pptx to fix; this skill can still validate the PDF, but cannot repair a " "flattened export the way it repairs a .pptx", ) # Case 2: a zip. Could be a zip that directly *is* a pptx (rename # confusion), a zip wrapping a pptx/pdf, or an HTML+assets bundle. if _is_zip(input_path): if _looks_like_pptx(input_path): return ResolvedInput("pptx", input_path, [], f"input is a .pptx saved with a non-.pptx extension: {input_path.name}") extract_dir = work_dir / "unzipped" extract_dir.mkdir(parents=True, exist_ok=True) with zipfile.ZipFile(input_path) as zf: zf.extractall(extract_dir) # Same search this skill would do for a "Send to Claude Code Web" # working directory — a downloaded zip (whether a plain Export or # Design's "Download zip instead" fallback) is just that same # bundle shape, pre-zipped. resolved = _scan_directory(extract_dir) if resolved.kind == "unknown": resolved.source_note = resolved.source_note.replace("did not contain", "(from the uploaded zip) did not contain") else: resolved.source_note = f"zip extracted; {resolved.source_note}" return resolved return ResolvedInput("unknown", input_path, [], f"unrecognized input type: {input_path.name} (suffix {suffix!r})") -
validate_pdf.py 3.2 KB
""" Post-conversion validation. The repair steps upstream (content_types.py, geometry.py, fonts.py) fix known defect classes *before* conversion; this module checks the actual PDF that came out the other end, so the repair report reflects reality rather than just "we ran the fixes and assumed it worked." Three checks, cheapest first: 1. Page count matches the slide count. A mismatch almost always means the converter silently dropped or merged slides. 2. No unexpectedly blank page. A single blank divider slide is normal; several in a row, or one where the source slide clearly had content, is a sign the render failed for that slide specifically. 3. No text bounding box touches/crosses the page edge. This is the direct check for defect class 2 (off-canvas shapes) — if geometry.py did its job, this should come back clean. If it doesn't, that's a real signal the shape-level fix missed something (e.g. a group-nested shape that was flagged rather than auto-fixed) and needs a human look. """ from dataclasses import dataclass, field try: import pymupdf as fitz # modern import name except ImportError: import fitz # older pymupdf releases expose the module as `fitz` EDGE_TOLERANCE_PT = 1.0 # points; ignore sub-pixel rounding at the page boundary BLANK_TEXT_LEN_THRESHOLD = 3 # a page with <= this many non-whitespace chars and no images is "blank" @dataclass class PdfValidationResult: page_count: int expected_page_count: int page_count_ok: bool blank_pages: list = field(default_factory=list) # 1-indexed page numbers edge_overflow: list = field(default_factory=list) # {"page": n, "bbox": (...), "text": "..."} @property def ok(self) -> bool: return self.page_count_ok and not self.edge_overflow def validate_pdf(pdf_path: str, expected_page_count: int) -> PdfValidationResult: doc = fitz.open(pdf_path) try: page_count = doc.page_count blank_pages = [] edge_overflow = [] for i, page in enumerate(doc, start=1): text = page.get_text("text").strip() images = page.get_images() if len(text) <= BLANK_TEXT_LEN_THRESHOLD and not images: blank_pages.append(i) pw, ph = page.rect.width, page.rect.height for block in page.get_text("dict").get("blocks", []): bbox = block.get("bbox") if not bbox: continue x0, y0, x1, y1 = bbox if x0 < -EDGE_TOLERANCE_PT or y0 < -EDGE_TOLERANCE_PT or x1 > pw + EDGE_TOLERANCE_PT or y1 > ph + EDGE_TOLERANCE_PT: snippet = "".join( span.get("text", "") for line in block.get("lines", []) for span in line.get("spans", []) )[:80] edge_overflow.append({"page": i, "bbox": bbox, "text": snippet}) return PdfValidationResult( page_count=page_count, expected_page_count=expected_page_count, page_count_ok=(page_count == expected_page_count), blank_pages=blank_pages, edge_overflow=edge_overflow, ) finally: doc.close() -
__init__.py 0 B
-
-
SKILL.md 10.2 KB
--- name: design-export-repair description: "Fixes broken decks/PDFs exported from Claude's \"Design\" feature (or similar AI deck generators) — cut-off or clipped text, wrong/substituted fonts, and corrupted .pptx package structure that shows up only once you export to PDF, not while looking at it in the app. Use this skill whenever the user uploads a zip, .pptx, or .pdf exported from Claude Design (or mentions \"Design\" export, an AI-generated deck, or a generated slide deck/PDF) AND reports it looking broken, wrong, cut off, garbled, or different from the original when opened, printed, or converted to PDF — even if they don't use the word \"repair\" or name the specific defect. Also trigger for general \"my exported PDF/deck is broken, help fix it\" requests when the file was clearly produced by an AI design/deck tool rather than hand-authored. Produces both a repaired, still-editable .pptx and a clean PDF, plus a report of exactly what was wrong and what changed." --- # Design export repair Fixes AI-generated deck/PDF exports (built for Claude Design's export path, but the checks are generic OOXML/PDF hygiene, not Claude-specific) that look fine in the app but come apart once exported — text missing its last few characters, fonts that don't match the original design, or a `.pptx` that some converters choke on even though PowerPoint opens it fine. Four defect classes were found and fixed against a real broken export; `references/defect-classes.md` has the full story for each, including how they were confirmed and what the fix trades off. Skim it before extending this skill to a case it doesn't already handle — in particular, defect #4 there is a good example of why "the shape geometry looks right" is not the same claim as "the rendered PDF looks right," and it's worth reading before assuming a new complaint is a geometry problem. ## End-to-end: from Claude Design to a clean file This is written for the person running Claude Code, not just for Claude — skim it once so you know what to actually do. There are two different ways a Design project reaches Claude Code, and they behave differently. **Path A — plain Export (this is the one that matches "my deck/PDF looks broken"):** In Claude Design, use Export and pick PDF, PPTX, or HTML. Save the file, then bring it to whichever Claude Code you're using — attach it in the chat if your setup allows that, or save it to disk and give Claude the path. Then tell Claude to use this skill: plain language like "this deck looks broken when I export it to PDF, can you fix it" is enough on its own (the skill's description is written to match that phrasing), or name it directly (`design-export-repair`) if you want to be explicit. **Path B — the "Send to Claude Code" button.** This is a related but different feature from plain Export — it's aimed at continuing design/ prototype work in code, not at fixing a deck export, but you may end up here anyway. It offers two options, and they behave differently: - *"Send to Claude Code Web"* opens a brand-new claude.ai/code session with the design bundle already sitting in that session's working directory — there's no file to save or attach, it's just already there. In that new session, tell Claude to look at what's in the working directory and use this skill on it. - *"Send to local coding agent"* gives you a prompt to paste into your local terminal Claude Code. Be aware: as of this writing, that default prompt depends on a "Claude Design connector" that local Claude Code doesn't ship, and reliably fails silently (see [anthropics/claude-code#69246](https://github.com/anthropics/claude-code/issues/69246) if you want the details). The same dialog has a **"Download zip instead"** option — that one actually works: it downloads a real zip of the design files. Save that zip, point local Claude Code at it (or unzip it into your project folder first), and tell Claude to use this skill on it. Whichever path you took, `scripts/fix_export.py` accepts a file, a zip, or a whole directory — `unpack.py` figures out what it's looking at rather than requiring one specific shape. One thing worth knowing about Path B specifically: its bundle format is Design's own `*.dc.html` canvas files, which is a different animal from the `.pptx` this skill's repair pipeline was actually built and verified against (see `references/defect-classes.md`) — a `.dc.html` bundle gets routed through the best-effort HTML path (`convert_html.py`) instead of the fully-tested one. If you're trying to fix a broken deck/PDF export specifically, Path A is the one to reach for. After either path, Claude runs `scripts/fix_export.py` and hands back a repaired `.pptx` (Path A) or PDF, plus a report explaining exactly what was wrong and what changed. Open the repaired `.pptx` if you want to keep editing in PowerPoint/Keynote/Google Slides, or use the `.pdf` directly if that's the deliverable you needed. One honest caveat: buttons and flows in Claude products change over time, and this description of Path B is based on Claude Code's own public issue tracker rather than hands-on testing of that specific handoff. If what you see doesn't match this, don't get stuck on it — just get the actual design files to Claude Code by whatever route works, and tell Claude to use this skill; `unpack.py` is written to figure out what it's been handed. ## Quick summary of what gets fixed | # | Defect | Symptom | Fix | |---|--------|---------|-----| | 1 | `[Content_Types].xml` declares parts that don't exist in the zip | Some converters fail, drop slides, or garble output; PowerPoint may silently "repair" and hide it | Load + re-save through `python-pptx` (rebuilds the manifest from what's actually there) | | 2 | A shape's box extends past the slide edge | Can get clipped by renderers that treat the slide as a strict viewport | Translate the shape back in bounds (or scale down + rescale table grids, if the shape is bigger than the slide itself) | | 3 | Custom webfonts (Inter, Plus Jakarta Sans, JetBrains Mono seen in practice) used but not embedded | Renderer without those fonts substitutes something else — wrong look, different text measurements | Install matching bundled/fetched font files where the renderer will find them, before conversion | | 4 | LibreOffice clips the tail of any text run with letter-spacing (`a:rPr spc`) | Tracked-out labels/kickers/badges lose their last 1-4 characters — **this is usually the actual cause of "text is cut off" even when #2 is also present** | Render a letter-spacing-neutralized copy through LibreOffice only; the returned `.pptx` keeps its real letter-spacing for editing | ## How to run it Everything is one script: ``` python3 scripts/fix_export.py <input> --out-dir <where to write results> ``` `<input>` can be: - a `.pptx` handed directly (this is what Claude Design's export has produced in practice — don't assume it must be wrapped in a zip) - a `.zip` — it gets extracted and searched for a `.pptx` first, then an HTML/CSS/JS bundle (`index.html` + assets), then a bare `.pdf` - a bare `.pdf` — repair is limited without the source; see below Read the script's own docstring for the full pipeline description before running it if anything about the input is unusual — it's kept in sync with what the code actually does. After it runs, look at `<out-dir>/repair_report.md`. It documents, in order: what was wrong with the package structure, every shape that got repositioned/resized (slide, before/after coordinates), the status of every font the deck uses (installed / fetched / falling back), whether the letter-spacing workaround applied, and a validation pass on the final PDF (page count vs. slide count, any pages that look unexpectedly blank, any text that still touches a page edge). Read this back to the user — it's written to be shown, not just logged, because the whole point of this skill is "nothing changes silently." Outputs land in `<out-dir>`: - `*.repaired.pptx` — still fully editable, structurally clean, no off-canvas shapes, original letter-spacing intact - `*.repaired.pdf` — the clean PDF, converted from a letter-spacing- neutralized render copy (see defect #4) so tracked-out labels don't clip - `repair_report.md` / `repair_report.json` ## When the input is a bare PDF with no source file Structural repair (package structure, shape geometry, font substitution) needs the editable source — there's no way to move a shape or fix a font reference in something that's already been flattened to fixed vector paths and rasterized text. `fix_export.py` still runs the same PDF validation pass (page count sanity, blank-page check, edge-overflow check) so you at least get an honest read on whether the PDF itself looks broken, but say so plainly to the user rather than implying a fix happened: ask if they can re-export from Design as a `.pptx`/zip instead, since that's where the actual repair leverage is. ## When the input is an HTML/CSS/JS bundle This path (`scripts/convert_html.py`) is best-effort, not backed by a real defect investigation the way the `.pptx` path is — there was no sample of this shape to test against. It applies the generic fixes that most commonly break an HTML deck's print-to-PDF (forces background/color printing, un-hides anything relying on `overflow: hidden` + scrolling, avoids mid-element page breaks) and picks a page size from whatever the markup hints at. If you hit a real broken HTML export, treat this as a starting point and read `references/defect-classes.md`'s closing section on how to diagnose a new case (render, compare, isolate, don't guess from markup alone) rather than assuming the existing generic fixes cover it. ## Extending this skill If a deck comes through with a defect not in the table above, the method that found defect #4 generalizes: don't debug from the XML/CSS alone, because "looks structurally fine" and "renders fine" are different claims and only the second one is what the user is reporting. Render the untouched original to PDF, render a copy with one specific, isolated change, and compare the two — that's what confirmed the letter-spacing bug after the geometry and font fixes alone didn't resolve the visible clipping. `scripts/validate_pdf.py` already does the page-count/blank- page/edge-overflow checks; add to it if you find another cheap, objective signal worth checking automatically on every run.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.