Claude Skill

editable-figure

Analyze source material, find relevant paper, README, or awarded-proposal references, and design concise figures as editable PowerPoint objects. Use for overview, mechanism, workflow, or hero figures when an editable PPTX is wanted, including simplifying dense drafts and combinin

LLM Mart · 0 points · 0 views 40 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download yzhao062-anywhere-agents-skills_editable-figure-f842be2.zip · 46 KB
Part of yzhao062/anywhere-agents — 8 skills

Install

skills CLI npx skills add https://github.com/yzhao062/anywhere-agents/tree/main/skills/editable-figure
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install yzhao062-anywhere-agents@llmmart
Git git clone https://github.com/yzhao062/anywhere-agents.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole yzhao062/anywhere-agents collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Editable Figure

Overview

Turn a document's central idea into a figure that a reader can understand quickly, then deliver a PowerPoint source that the author can actually edit. The common workflow is analyze, find references, design, build native objects, inspect in context. Paper, proposal, and README figures share this workflow but serve different reader decisions.

Choose the scope

Respect the requested deliverable. An assessment or prompt-only request does not require generating a deck. Once figure creation is authorized, use subsequent feedback to revise the artifact without repeatedly asking permission. Choose a reasonable composition and produce a reviewable draft when the source is sufficient.

Use this skill for the figure's editorial and design decisions. If an installed presentation skill applies, read it for the current authoring runtime and validation requirements. Without one, use an available PPTX library or native PowerPoint automation and the principles in native-powerpoint.md. This workflow works with Codex, Claude Code, or another capable agent; it does not depend on one model or desktop plugin.

Platform requirement. Before starting a figure build, confirm that the session can use desktop PowerPoint on Windows or macOS for the rendering and editing checks in step 5. A PPTX library such as python-pptx or PptxGenJS can create native objects headlessly, and that is a real capability. It does not complete these checks: nothing on a headless machine can confirm the result opens, renders, and edits as intended, and an unverified figure is the failure this skill exists to prevent. The bundled scripts/render_powerpoint.ps1 uses Windows COM automation; macOS requires an available native PowerPoint workflow instead.

Without that capability, explain the limitation before building. Assessment and prompt-only work can continue unaffected. Offer ci-mockup-figure when its output meets the user's needs. Preserve an explicit PPTX requirement unless the user agrees to another deliverable; do not silently substitute a flattened image and call it an editable figure.

Nearby workflows have distinct outputs:

  • figure-prompt-builder: prompts and reference-guided concepts, when that is the requested endpoint.
  • ci-mockup-figure: HTML mockups, screenshots, TikZ, or other code-native figure formats.
  • A plotting workflow: scientific data plots with reproducible data and axes.
  • A presentation workflow: complete decks. Here a slide is a canvas for a document figure, not a presentation with a cover.

Do not replace an explicitly requested SVG, Illustrator document, or screenshot with PPTX merely because this skill is available.

1. Analyze the source and the reader's decision

Read the relevant section and its neighboring prose, the document's purpose, existing figures, and primary artifacts behind any claims. Identify the intended placement and display width before allocating space. Read only enough of a large repository to establish the contribution, terminology, evidence, and constraints.

Write a short working brief, in scratch space or beside the figure source when it needs to persist:

Decision Record
Audience and placement Who sees the figure, where, and at what size?
Reader takeaway One sentence the reader should remember after a quick look.
Usefulness What can the reader understand, evaluate, build, or decide because of this work?
Necessary visual evidence The example, relationship, mechanism, or measured result that supports that takeaway.
Content allocation What belongs in the drawing, caption, neighboring prose, or a table?
Claim boundaries Facts versus illustrative examples, proposed work, predictions, and measured results.

Apply the relevant mode in document-contexts.md. Do not require every figure to summarize the whole project.

For proposal figures, also read proposal-figures.md for distinct figure roles, visible aim-name consistency, shared-graph semantics, and reading at manuscript width. It links the author's preferred examples. Keep lessons from compact integration figures separate from untested large-overview designs.

For a benchmark, distinguish the system producing the record from the method being evaluated. For a proposal, distinguish planned capabilities from completed results. For a README, show the actual user workflow and supported behavior.

2. Find and study references before designing

For a new figure or substantial redesign, run a focused reference search before choosing the composition. Read reference-search.md and use the route that matches the document:

  • Paper: related top-venue papers, typically NeurIPS, ICML, and ICLR for ML/AI. Inspect the actual figure, caption, and nearby prose.
  • GitHub README: current relevant trending projects and related active repositories. Inspect the rendered README and its actual visual assets.
  • Proposal: start with the author's preferred examples when they fit the figure's job. Search the local awarded/funded collection further when needed, using figure purpose and agency/program as selection criteria. Check indexes and known collection paths first.

Select a small number of references and state what transfers: information hierarchy, a concrete example, a mechanism, or a useful visual structure. Keep exact source locators in the working brief. Popularity, acceptance, and funding are discovery signals, not proof of figure quality.

Use the user's supplied references when appropriate. Reuse already inspected examples for small revisions rather than restarting research. If a source is inaccessible, report that boundary and continue with available material; do not invent search results or funding status. Once a direction is supported, proceed to design without adding a reference-approval checkpoint.

3. Design for the takeaway

Make usefulness visible through an understandable problem, a consequential distinction, a concrete mechanism, or a supported result. Claims such as "powerful" and lists of components do not substitute for that evidence.

Choose the visual structure from the message: an example with an intervention point, evidence becoming available over time, a bottleneck and proposed mechanism, input-to-output transformation, comparison, or another justified structure. Do not default every task to three columns, a lifecycle, or a grid of cards.

For a position, framework, or concept paper, make the overview a compressed argument rather than only a taxonomy or a polished slogan. Retain the minimal claim graph and a concrete domain object needed to show why the position follows, even while removing prose. Read position-and-concept-figures.md for semantic compression, staged refinement, and links back to the manuscript.

Use graphics to carry relationships. Text should identify objects and explain only what the graphic cannot. A useful starting point for a compact overview is a few focal elements, each with an object and a short question or label. Adjust to the source; this is not a fixed panel or word quota.

For overview figures, retain a concrete domain object that participates in the mechanism, such as a circuit feeding measurements, rather than relying on a domain name in the title. Check that a reader can identify both the subject and the connection between panels. Shared variables, a carried example, or explicit input/output edges can establish that connection; adjacency and panel letters alone cannot. A section pointer printed in the figure, such as Section 3.2, helps navigation but does not explain the mechanism. Recheck it in the compiled document after structural edits, because section numbers move.

Concise explanation can use visually rich content. Maps, scene illustrations, screenshots, scientific plots, and structured panels can make a proposal easier to understand when they carry its substance. Evaluate their role and reading hierarchy rather than treating flat minimalism, a particular font, or the absence of rounded panels as a quality test.

Make occupied space earn its place

Space efficiency is a design requirement for every figure, especially in papers and proposals where page area is scarce. Judge useful information per occupied area at the final document size. Inspect both the figure's internal composition and the space reserved around it in the document, including captions, gutters, and wrapped prose. A figure can be legible and editable yet still waste substantial page space.

Distinguish whitespace that supports hierarchy and separation from broad unused bands or regions with no reading function. Check each panel's actual content bounds; a long footer, wide background, or oversized text box can hide wasted space when inspecting only the overall bounding box. Reflow or shorten the element that sets an unnecessary width, then reduce the canvas or panel dimensions while preserving meaningful content and readable type. Do not add labels, decoration, or repeated claims just to fill a hole, and do not compress necessary gutters to maximize a pixel-fill ratio.

Return reclaimed space to the destination. For a wrapfigure, reduce the placed width along with the canvas when the goal is to give room back to prose. Cropping the source while retaining its old placed width enlarges the content and can increase its height instead of saving page area. Recheck physical font size, caption wrapping, adjacent paragraphs, section transitions, and total pagination. Read space-efficiency.md when reflowing a sparse figure or adapting one to a compact document slot.

Use the lab's default palette

For new paper, proposal, and GitHub README figures without a specified visual identity, use the shared CatchBench, Cat-DPO, and No Attacker Needed color family: white background, mint and coral as the main contrast, teal and warm gold when additional categories are needed, and dark text with quiet gray context. The latter two papers are Tiankai Yang's reference examples. Read default-palette.md for color roles, source evidence, and the canonical reusable tokens in references/default-palette.json.

An explicit user palette or an established document palette takes priority. Preserve accepted figures' colors when editing them. Carry the same color-to-meaning mapping across the schematic, results, caption keys, and other figures in a document. Reference searches may supply composition ideas without changing this default visual identity; do not repeat a paper search just to recover the stored colors.

Reduce reading load before reducing font size

For each proposed label, identify the information it adds or the misunderstanding it prevents. Omit text whose removal leaves the correct reading intact, including sentences that repeat a visible relationship, panel title, or caption. Do not add explanations, formal terms, or caveat blocks merely to make the figure appear rigorous. Keep verification details in the working notes and methodological detail in the surrounding prose unless the reader needs it at that point in the graphic.

Keep a number when it carries the claim, such as a verified performance gap or scale comparison. Do not add counts merely to make the work appear substantial. Do not transfer every deleted sentence into an oversized caption.

Simplification must preserve the inference. Check whether a removed qualifier makes a diagnostic look like general evidence, an illustrative story look like a paired experiment, a forecast look like a result, or a highlighted answer look like an input available to the evaluated method.

Prefer short, concrete action labels; keep a formal term underneath only when it helps connect the drawing to the prose. Judge understanding rather than word count alone. A record edit should be described as a record edit: wording that implies the agent actually behaved differently requires evidence of that behavior.

Require secondary strips and extra panels to explain a distinct relationship or supply useful evidence. A generic sequence such as edit, rebuild, record may add little even when accurate. If its only contribution fits in one caption sentence, remove it and reclaim its space. When that secondary mechanism is itself central, show a concrete before/after example instead. After deleting a footer, reduce canvas height while keeping its width, retained object coordinates, and font sizes fixed. If the width or object scale changes, recalculate printed text size.

The reader should recognize the main contrast at a glance and explain the figure after reading its short caption. Inspect at the intended manuscript or README width. A legible full-screen slide can still fail as a paper figure.

When breadth of resources, research outcomes, or community is itself the argument, retain the supported inventory and give it sufficient page area and reading levels. For a compact destination, select the relevant part rather than shrinking the entire inventory below readable size. A proposal ecosystem overview can expose its organization at a glance while leaving individual publications or examples for a later read. The reduction criteria target detail that does not support the figure's claim; see proposal-figures.md and the preferred examples.

For mechanism diagrams, read mechanism-figures.md: it covers minimal examples, source checks, essential qualifiers, secondary content, and a concrete revision case.

When revising a dense first figure, read iteration-example.md. Its layout and counts are an example, not defaults.

Combine explanation with evidence when useful

A dominant schematic can explain what the work makes possible, while a smaller result panel gives the reader a numerical reason to care. Repeating a selected result from Results can serve this introductory purpose. Judge its contribution to the reading path rather than rejecting all repetition or optimizing for the fewest words.

When the author wants a large (a) and narrow (b), reserve the gutter first, then divide the remaining width according to their reading roles. Judge separation between the actual labels and artwork, including axes and callouts. A divider does not compensate for crowded panels. After changing the canvas, recalculate text size at the final display width.

Read panels-and-results.md when combining an explanatory panel with data, separating regions within a wide schematic, or revising panel spacing. It covers panel hierarchy, region separation, cross-panel edges, source-derived numbers, descriptive versus inferential claims, and focused revisions. Use the plotting workflow for scientific correctness and the native PowerPoint guidance for editable chart behavior.

Choose how to draft

Native PowerPoint can be the first draft when the figure consists of text, nodes, arrows, and simple geometry. A preliminary raster generation is optional, unless the user explicitly requests it. If an image generator helps explore a richer concept, use the available image-generation workflow, then rebuild the required semantic content as native objects. Do not treat a screenshot on a slide as an editable figure.

Show one recommended direction with a concrete reason. Generate alternatives only when a real design tradeoff remains or the user requests them. Save revisions under meaningful new names so the author can compare them.

4. Build the editable figure

Read native-powerpoint.md before authoring or converting a diagram. Discover current tools and fonts rather than copying cache paths or runtime versions from a previous session.

Use text boxes for text, native shapes for semantic objects, and attached connectors for relationships that should follow moved nodes. Keep imported photos or screenshots as such when they are part of the requested figure. Name objects and group them by meaningful region so both whole-block and individual editing remain practical.

Use a figure-sized canvas rather than inheriting a slide aspect ratio without reason. Preserve an existing palette or typography when supplied. Work out the final physical text size after scaling into the document. Do not shrink labels to rescue a composition that needs editing.

Retain the editable source and a reproducible builder when one was used. After manual PowerPoint edits, identify the current source of truth; a stale builder must not silently overwrite an author's changes.

5. Inspect semantics, appearance, and editability

These are separate checks:

  1. Meaning: Trace claims to the source, distinguish illustrative from empirical content, and verify the figure plus caption conveys the intended usefulness without overstating scope.
  2. Rendering and space: Export the final PPTX and inspect the actual result for clipping, wrapping, arrow direction, visibility, spacing, and unused regions. Evaluate each panel and the complete occupied document area, not only a tightly cropped preview. Recheck at the intended document width, and in the actual document when insertion is part of the task.
  3. Editability: Verify that important text and objects are native and independently selectable. On a disposable copy, edit representative text and move a connected node to check behavior. For a native result chart, check its data source and edit/restore a series value; verify the embedded workbook when portable data editing is required. Counts of shapes or a PNG preview alone do not prove this.

Use the installed presentation workflow's validators when available. The optional Windows helper scripts/render_powerpoint.ps1 exports a single-slide figure through local PowerPoint and reports native object counts. It is a renderer and inventory check, not a substitute for visual or editing inspection.

Fix issues that affect the requested result and stop when the checks pass. Do not expand a one-figure revision into unrelated document edits or a large review process. If native PowerPoint becomes unavailable, report which rendering and editing checks remain incomplete and follow the platform requirement above. A fallback preview does not complete this workflow.

Treat review suggestions as claims to reconcile with primary sources and the user's requirements. An optional reviewer does not decide the authoring format or override later user feedback. Record which artifact version a review covers. A small spacing revision calls for geometry, rendering, and relevant data-integrity checks; it need not restart reference research or a full review cycle.

For a visual-only pass, freeze wording, font sizes, semantic groups, and connector endpoints, and compare them with the previous version so polish cannot hide a redesign.

When another model is requested, prepare a frozen image, caption, relevant section, and only the source excerpts needed to verify the mechanism. Include the actual publication width and column layout so advice addresses the real destination. For a first-read comprehension check, initially send only the image and placement context, without revealing the desired takeaway. Obtain that reading before sending the caption and source for the semantic check; putting the answer later in the same prompt still primes the reader. Other review tasks can use the combined packet. Use existing authorization for that material and destination; resolve any missing permission only when required, and do not expand a limited review into sending the full manuscript or repository. Separate the reviewer's direct checks from locally reported PowerPoint and document checks.

6. Deliver for the destination

Save the editable .pptx and the appropriate viewing or publication export together, using the project's existing figure/assets folder or the user's named destination. Do not leave the only usable deliverable in temporary storage.

  • Papers and proposals: normally a vector .pdf, plus a PNG preview when useful.
  • GitHub README: a .png or compatible .svg for display, plus the .pptx source. Provide meaningful alt text when embedding it.

Check that the exports correspond to the final editable source. Supply a concise caption or nearby sentence when needed to make the figure interpretable. Do not insert into or rewrite the document unless authorized by the task. Link the editable file, show a preview, and briefly identify any material limitation.

Keep caption or adjacent explanatory text in one canonical place; for LaTeX insertion, use one canonical float fragment. After a revision, synchronize the PPTX, viewing exports, builder dimensions and group definitions, caption, and current verification notes. Before replacing files after a long build, check whether those targets changed concurrently. Preserve unrelated edits and render or compile the current document when insertion is authorized; a private snapshot is not a replacement for newer document work. Keep historical reviews labeled by version rather than implying they cover later edits.

Files (anywhere-agents)
  • agents
    • openai.yaml 458 B
      interface:
        display_name: "Editable Figure"
        short_description: "Editable PPTX figures; desktop PowerPoint required"
        default_prompt: "Use $editable-figure to analyze this source, find suitable references for its document type, and design a concise figure that shows why the work is useful. Deliver an editable PPTX with an appropriate preview or publication export unless another output format or an assessment or prompt-only deliverable was requested."
      
  • references
    • default-palette.json 414 B
      {
        "name": "lab-default",
        "tokens": {
          "background": "#FFFFFF",
          "mint": "#BFDFD2",
          "coral": "#ED8D5A",
          "teal": "#4198AC",
          "gold": "#ECB66C",
          "gray": "#C9C9C9",
          "ink": "#202923",
          "mutedText": "#67756D",
          "rule": "#DCE6DF",
          "mintStroke": "#365E49",
          "coralStroke": "#B66037"
        },
        "pairOrder": ["mint", "coral"],
        "categoricalOrder": ["mint", "teal", "gold", "coral"]
      }
      
    • default-palette.md 4.9 KB
      # Default figure palette
      
      The user selected CatchBench, Cat-DPO and No Attacker Needed as the lab's default color references on 2026-09-05. Apply this visual family to new paper, proposal and GitHub README figures unless the user or existing document specifies another palette. This preference concerns visual identity; it does not prescribe those papers' layouts, icon styles or scientific claims.
      
      For proposal composition, the separately selected [proposal examples](proposal-style-exemplars.md) guide visual richness and hierarchy. They do not replace these color defaults or the established palette of a manuscript being edited.
      
      `default-palette.json`, beside this file, is the canonical source of exact color values. Use it directly in a builder or copy its tokens into the figure's working brief. It is self-contained, so ordinary figure work does not depend on downloading the reference papers again.
      
      ## Roles and use
      
      | Tokens | Default use |
      |---|---|
      | `background` | White canvas and open space. |
      | `mint` | Main comparison, original record or supporting structure. |
      | `coral` | Focal method, intervention, changed object or highlighted result. |
      | `teal`, `gold` | Additional categories or methods when mint and coral are insufficient. |
      | `gray` | Context, baseline or inactive objects, with a readable border when necessary. |
      | `ink`, `mutedText` | Primary text and secondary labels. |
      | `rule` | Quiet separators and non-data structure. |
      | `mintStroke`, `coralStroke` | Dark outlines, connectors or labels associated with the corresponding light fills. |
      
      Use a subset suited to the message. Most explanatory figures need mint, coral and neutrals. For a new two-series comparison, use `pairOrder`; `categoricalOrder` is a four-series preset, not a sequence to truncate for fewer series. For three series, retain the main mint/coral contrast and choose an additional color by its role. Establish a semantic mapping in the working brief and retain it across panels and related figures; existing identities take priority over either preset's order. Coral identifies the focal element, so it does not universally mean success, failure, danger or the proposed method. If a comparison is also shown in a results panel, its identities should retain their colors.
      
      Pastel fills work with dark text and visible outlines. Use the darker stroke tokens where a light fill would disappear as a thin line. For small scientific curves, retain distinguishable markers, line styles and direct labels; do not rely on a mint/coral hue difference alone. Check the actual document width and grayscale legibility. Keep a white figure canvas for README exports unless the repository has an explicit dark theme.
      
      ## Source evidence and existing variants
      
      These papers share a visual family, not an identical set of historical HEX values. The canonical tokens standardize future work; do not claim every token was used in all three papers or recolor their existing artifacts without a request.
      
      Locators below are relative to the `internal-writing` repository unless a public link is given:
      
      - **CatchBench** ([paper](https://arxiv.org/abs/2608.22808)): `papers/iclr-2027-auditablebench/figures-preamble.tex` declares mint, coral and gray used by the paper. The accepted editable Figure 1 and Gold mechanism builders under `figure-src/` supply the dark ink, secondary text, rule and darker outline tokens. Their existing mint/coral fills are `#C2DFD1` and `#ED986C`, slightly softer variants of the canonical pair. Those exact fills govern both revisions and new figures in that document unless the user requests a palette change. Do not mix both variants in a single new figure.
      - **Cat-DPO: Category-Adaptive Safety Alignment** (Tiankai Yang and coauthors): [paper](https://arxiv.org/abs/2604.17299); local figures `references/style-exemplars/catdpo/figures/per_category_harm_bars_top8_floor.pdf` and `balance_panels.pdf`. RGB values recovered from their vector drawing objects are exactly the canonical mint, teal, gold and coral tokens. The overview `method_overview.pdf` also uses light mint and gold regions. Its illustrations are composition references, not artwork to copy.
      - **No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents** (Tiankai Yang and coauthors): [arXiv v1 PDF](https://arxiv.org/pdf/2604.01350v1). Inspected Figure 1 on page 1 and Figure 2 on page 3 for pale green/warm fills, teal/gold accents, dark outlines and white space. The PDF also contains pale green `#D7E4BD` and pink `#F2DCDB` in Figure 1, and blue `#779BB8` and rose `#E37F80` in later plots. These document the reference's variants; they are not additional default categorical colors. New work should use the canonical tokens to keep the three-reference visual family consistent.
      
      Source inspection date: 2026-09-05. CatchBench and Cat-DPO were checked from local source and figure PDFs; No Attacker Needed was checked from its public arXiv PDF. Existing document colors or an explicit user request override this starting palette.
      
    • document-contexts.md 4.8 KB
      # Adapt the figure to its document
      
      The invariant is the reader's decision, not a shared page template.
      
      | Context | Reader should quickly understand | Strong visual content | Detail usually better elsewhere |
      |---|---|---|---|
      | Paper first figure | The problem, contribution, and why the work is useful or surprising | A concrete failure, a new evaluation question, a mechanism, or one verified result | Complete corpus lists, every metric, implementation parameters, repeated contribution bullets |
      | Paper method figure | How the central mechanism works and which distinction is technically important | Named inputs, operations, constraints, outputs, and necessary dependencies | Boilerplate infrastructure, incidental class names, generic arrows between empty boxes |
      | Proposal overview | Why the gap matters, how the proposed approach addresses it, and what successful work produces | Gap linked to mechanisms and outcomes, with essential dependencies or validation | Every work package, staffing detail, acronym, and schedule item |
      | Proposal aim or timeline | Feasibility, logical dependency, or a concrete success criterion | The relevant aim's method and evaluation, or a schedule when timing is the actual question | Unrelated aims, speculative results presented as completed, repeated overview text |
      | README hero or overview | Who the project helps, what it does, and what using it yields | An actual input/output example, user workflow, or screenshot of an existing feature | Installation flags, internal module lists, unsupported performance claims, tiny terminal text |
      
      ## Papers
      
      Read the introduction around the figure callout. Allocate information across figure, caption, and surrounding prose as a unit. The figure need not restate the title, abstract, benchmark inventory, and methods table.
      
      For a benchmark, ask what meaningful evaluation becomes possible. Coverage counts can support that point, but do not automatically make the figure persuasive. A method-family comparison needs supported results and compatible metrics. Separate data tracks must not be drawn as a matched experiment unless such pairing exists.
      
      A large overview plus a compact result can make both the contribution and its quantitative evidence visible. Use [panels-and-results.md](panels-and-results.md) for hierarchy, whitespace, and statistical scope. A selected result can earn introductory space even if Results explains it again.
      
      Keep scientific distinctions that affect interpretation: input versus target label, prediction versus ground truth, natural data versus injected diagnostics, observation versus intervention, and measured versus illustrative quantities. Remove decorative qualifiers, not essential scope.
      
      Judge legibility at the final column or text width. Compute the scaled type size rather than relying on the font size in the PPTX. Inspect the actual manuscript when integrating; placement, caption length, and neighboring text change the space cost.
      
      ## Proposals
      
      Read the relevant aims and, when provided, the solicitation's criteria. Make the intellectual connection between the problem, proposed mechanism, and expected outcome clear. A diagram of three aims becomes useful only when it explains why those aims belong together or how their outputs support the proposed result.
      
      Use restrained labels such as "proposed" or "target" where they prevent anticipated outcomes from reading as achieved results. Distinguish preliminary evidence from planned experiments. Do not invent partner commitments, data access, milestones, or measured gains to strengthen the figure.
      
      Aim numbers, validation milestones, and responsibilities belong in the figure when they resolve feasibility or integration. Keep them out of a motivation figure if they divert attention from the problem. Match the proposal's print context and page constraints; do not assume a full presentation slide can be pasted at readable size.
      
      ## GitHub READMEs
      
      Read the opening paragraph, quick start, and actual supported behavior. Select one workflow or recognizable output that helps a visitor decide whether to keep reading or try the project. Build a visual example from genuine capabilities; label synthetic examples when they could be mistaken for measured output.
      
      Choose a web display export and retain the PPTX as its editable source. Inspect at a typical README content width and on a narrow viewport. Labels must survive scaling. Use a background that remains legible in the intended light/dark context; verify transparency deliberately rather than assuming it helps.
      
      A real screenshot may be a native image inside the figure. Surrounding callouts and workflow arrows can be editable without claiming that the screenshot's UI is editable. Avoid imitating controls that do not exist. Put installation instructions and detailed options in accessible Markdown, and supply alt text for the image.
      
    • iteration-example.md 5.3 KB
      # Lesson from a benchmark first figure
      
      ## Reduce reading load while keeping the scientific distinction
      
      In a September 2026 CatchBench drafting session, the first editable figure combined three evidence states with corpus names, board counts, label-process information, method families, and scoring details. Its contents were mostly accurate, but the drawing behaved like a benchmark specification table. Readers had to work through the inventory before seeing the reason to use the benchmark.
      
      The author's target was a first figure that, alongside the surrounding paper text, quickly made the benchmark understandable and useful. Completeness inside the image was secondary. An early lean revision retained three record illustrations and short audit questions. It contained about 34 English words, excluding step numbers; reducing the canvas from 1600 by 820 to 1600 by 480 cut height at the same width by about 41 percent. These describe an intermediate draft, not the accepted final composition or an ideal word budget.
      
      A later requested Claude review helped combine the old figure's evidence structure with the lean draft's concrete questions. LIVE and POST traces aligned, while PRE remained a distinct configuration card because it used a separate corpus. A common heading identified the evaluated auditor. Nodes, dependency arcs, and masked future evidence carried the mechanism with minimal text.
      
      Simplification preserved scope. The diagram did not supply an outcome label to a detection entrant, treat all three states as one empirical run, or present Gold cause attribution as a general natural-data claim. The question "Why?" was removed from the general POST heading because attribution belonged to a specific diagnostic board.
      
      ## Restore a small result for a different reading role
      
      The reviewed schematic initially omitted the previous LIVE result inset because Results already discussed it. The author then explicitly requested side-by-side (a) and (b), with the explanatory panel much larger. That feedback changed the design: the schematic explained what the benchmark enables, and a compact result gave a numerical reason to care. Repetition in Results was not sufficient reason to exclude that useful introductory evidence.
      
      The result values came from the authoritative board through the existing plotting parser. SWE-Gym gains were `[0.113, 0.103, 0.131, 0.141]`; tau-bench gains were `[0.000, -0.001, 0.020, 0.046]`, at 25/50/75/100 percent prefixes. These were differences of printed mean AUCs for the named methods. The quarter-prefix registered SWE-Gym paired effect was `0.101` under a different averaging procedure. The inset therefore made a descriptive claim and did not label its `0.113` point as that registered effect. Inferential details remained in Results.
      
      Review recommendations were reconciled with sources and user intent. The requested editable PPTX format was retained despite a suggestion to use TikZ. Native export demonstrated searchable vector output. Earlier Claude reviews applied to the preceding schematic; the subsequent result-panel and spacing revisions were checked locally, without claiming a new Claude review of those versions.
      
      ## Separate panels with whitespace and check the paper
      
      The first combined draft felt too crowded. Panel widths already had the requested hierarchy; the missing element was enough space between them. The accepted revision increased the nominal gutter from 27 to 91 design pixels, removed the vertical divider, and retained widths of approximately 1180 and 389 pixels, a 3.03:1 ratio.
      
      The canvas widened from 1600 by 580 to 1664 by 580. At the manuscript's fixed 396 pt figure width, the smallest figure text became approximately 7.16 pt. This was inspected in the compiled paper, where Figure 1 remained on page 2 and the caption occupied five lines. These are case measurements, not minimum font sizes, mandatory ratios, or standard gutters.
      
      The lesson is to allocate the gutter independently from panel widths and inspect visible labels, not only panel rectangles. Expanding the canvas trades more separation for smaller printed text; recalculate and view the actual page.
      
      ## Deliver editable data as well as editable shapes
      
      The accepted master contained five groups, native diagram shapes and connectors, one native chart, one embedded data workbook, and no picture objects. The Figure 1 PDF contained no raster images. Representative text, node, and chart-series edits were tested in PowerPoint and restored; the final chart values and workbook passed validation. These checks did not claim a separate interactive Edit Data UI test.
      
      PowerPoint saves exposed two non-obvious issues: resizing the canvas recentered existing groups, and even a geometry-only save reserialized chart numbers. The final spacing adjustment preserved chart/workbook parts byte-for-byte while changing the slide and panel geometry. The master PPTX and matching PDF/PNG were saved together in `figure/`; the builder, parsed data, and validation notes were retained beside the figure sources.
      
      Transfer the editorial decision, not the exact layout: give readers a clear reason to care, use visual structure and selected evidence to support it, and retain the qualifiers needed for a correct interpretation. A proposal may pair a mechanism with preliminary evidence; a README may pair a workflow with an actual output. Once the content is accepted, keep spacing feedback a focused layout revision.
      
    • mechanism-figures.md 6 KB
      # Make a mechanism easier to understand
      
      Use this reference when a figure must explain how a method, intervention or evaluation works, especially when a draft has become a list of operations. It complements the first-figure case in [iteration-example.md](iteration-example.md).
      
      ## Build the smallest example that carries the distinction
      
      Identify the original object, the operation, what changes, and how the output or target is determined. A minimal before/after example can carry more meaning than a longer abstract pipeline. Keep the unchanged original visible when it establishes the comparison or control. Use parallel examples for mutually exclusive interventions; merging them into one graph can imply both happen together.
      
      Trace each drawn change to the implementation or other authoritative source for that exact setting. Check arrow direction, the object being edited, label construction, and any eligibility or selection rule. Evidence from a related version is insufficient when the illustrated version has a different mechanism. A small read-only check of the relevant functions can settle a concrete question without rerunning the full experiment.
      
      For proposals, use the same before/after structure to explain a proposed operation and label its status. For READMEs, use an actual supported input/output transformation. Neither requires pretending that a schematic is an observed experiment.
      
      ## Translate terminology without losing the claim
      
      Use a short verb and a concrete object for the main label. Keep the formal name as a secondary label when the reader needs to find it in the paper. Check the graphic, caption and adjacent prose together rather than demanding that every qualification appear inside the artwork.
      
      The CatchBench Gold revision illustrates the difference:
      
      | Draft wording | Revision and reason |
      |---|---|
      | `Use older information` | `Point to older step` describes the dependency-record edit without claiming the agent consumed different data during execution. |
      | `Changed step: 3` | `True label: step 3`, with a caption explaining that the edit determines the label independently of a detector. |
      | `Before the change` | `From the original` under `Steps to rank`, with the caption explaining that candidates are fixed from the original record. This clarifies selection rather than suggesting a timeline. |
      
      Preserve short qualifiers when they constrain the example. In this case `Example: remove a link` matters: original steps 2 and 3 are both eligible for dropped grounding, but only step 3 is eligible for the illustrated stale-state edit. The source functions confirmed that distinction. The text also limits the candidate pool to the matched control and states that controlling eligibility does not eliminate every construction artifact. Neither the figure nor its caption turns an injected diagnostic into validated natural-failure evidence.
      
      These are examples of distinctions to retain, not labels every future figure must use.
      
      ## Remove secondary content that does not explain enough
      
      The initial figure included `Named-value v2 -> Edit one value -> Rebuild the links -> Record step + field`. This distinguished the second substrate, but it did not show which value changed or why its dependencies changed. Recording the target also overlapped with the label box above. The user chose to remove the strip and retain one caption sentence explaining that v2 edits an argument value and rebuilds dependencies from the edited record.
      
      This decision focused the image on injection, label construction and candidate selection. The canvas changed from 1664 by 600 to 1664 by 490, reducing height by about 18 percent at the same manuscript width. The retained panels and fonts stayed fixed. Non-node labels fell from 60 to 48 words, and minimum printed text stayed approximately 7.62 pt at a 396 pt width. These measurements describe this case; they are not layout targets or a general minimum font size.
      
      If value editing were the main contribution, the better next figure would show a specific value before and after the edit and the resulting link change. Generic operation lists do not substitute for that explanation. Apply this criterion to optional proposal workstreams and README internals too, while retaining pipelines that actually explain a dependency or user workflow.
      
      ## Use review and verification at the right scope
      
      The requested Claude review received only the authorized image, caption, Section 4.3 and necessary code excerpts. It identified missing distinctions and a gap in the supplied file-level pool evidence. Adding the exact relevant excerpt closed the gap. Its warning about a two-column layout was conditional; the actual manuscript was single-column, so the existing full-textwidth float was appropriate.
      
      Shorter wording can resolve a review finding without copying the reviewer's longer label. Include scientific substance and the author's reading-load goal in the reconciliation. The closing review inspected the then-current image; the later user-requested strip removal was verified locally and recorded as a subsequent version. It did not require another broad review.
      
      Native PowerPoint inspection found a curved edge that collapsed onto the middle node despite looking correct in the library preview. The fix used attached segments and transparent routing anchors; see [native-powerpoint.md](native-powerpoint.md). Editing/restoring a label and moving/restoring a connected node tested behavior separately from package validity and vector PDF export.
      
      Finally, update the caption in its canonical place and all figure exports, then render or compile the document and inspect the actual page. If other figures changed during the work, preserve them and distinguish their resulting page reflow from a defect in this figure. Report checks against the current artifact; keep earlier reviews and measurements as dated evidence.
      
      Case evidence: `papers/iclr-2027-auditablebench/figure-spec/gold-mechanism-vet/` and `figure-src/gold-mechanism/` in `internal-writing`, 2026-09-05. The guidance above is usable without those project files.
      
    • native-powerpoint.md 14.3 KB
      # Native PowerPoint authoring and inspection
      
      ## Tool choice and source structure
      
      Use the current presentation skill and its documented API when one is available. In an environment without it, an available library such as PptxGenJS or native PowerPoint automation can produce editable drawings. Check local capabilities before selecting a backend. Do not hard-code a user's cache paths, install a replacement runtime without need, or claim another application was tested because a PPTX package opens as ZIP.
      
      A reproducible builder should separate the figure's labels, palette, coordinates, and primitives. Use a figure-sized canvas, semantic object names such as `LIVE_boundary` and `POST_event_3`, and groups that correspond to meaningful regions. Keep labels separately editable unless text belongs inside its node. An explicit source image may remain an image; a flattened diagram does not meet a native editability requirement.
      
      Use attached connectors when moving a node should move its edges. Static curves or freeforms are appropriate for decorative arcs or relationships that do not need attachment. If a curve is only a freeform, do not describe it as an automatically rerouting connector.
      
      Choose an available font and verify actual rendered text. Numeric dimensions may be pixels, points, or EMUs depending on the API. In OOXML, 1 inch is 914400 EMUs; in a 96 DPI authoring canvas, 1 pixel is 9525 EMUs. Check the library's conventions. A figure scaled to one third of its slide width also scales its text to one third of its point size.
      
      ## Lessons from an actual native build
      
      These were observed in September 2026 with the Artifact Tool deck-authoring runtime and native PowerPoint. Artifact Tool is one authoring runtime, not a required dependency of this skill. Recheck API-specific remedies against the installed version; they are diagnostics, not reasons to patch every deck.
      
      | Symptom | What to inspect | Narrow remedy |
      |---|---|---|
      | Connector exists but is invisible | Z-order relative to panel backgrounds | Bring that connector forward or send the relevant background back. Keep intended masks above hidden content. |
      | An overhead curved connector collapses onto an intermediate node in PowerPoint | The native export, not only the library preview | If that renderer cannot preserve the curve, use attached straight segments with transparent native routing anchors. Group the route and its nodes; verify attachments and clearance after moving them. |
      | Arrow points backward | The exported arrowhead and actual start/end attachment | Fix the head/tail option that controls the intended endpoint. API naming can differ; inspect the rendered direction. |
      | PowerPoint disables grouping | Group locks on the specific native shapes | Prefer the authoring API's unlock option. If the exporter wrote `a:spLocks noGrp="1"`, remove only that grouping restriction in a new draft copy, then revalidate. |
      | COM rejects restoring a position | A `Single` property passed to a setter expecting `Double` | Preserve and restore as an explicit `[double]`, for example `[double]$originalLeft = $shape.Left`. |
      | PowerPoint opens a path but export fails | Absolute export path and native Windows separators | Normalize export paths with .NET path utilities or `Join-Path`. |
      | COM reports `80070520` in a sandbox | Whether the current runtime has an interactive logon session | Use the supported execution/approval mechanism if authorized. Otherwise export with the available renderer and disclose the missing native check. Do not disable controls or repeatedly retry unchanged calls. |
      
      If XML repair is necessary, keep it confined to the known defect in a new file. Preserve relationships and other package parts. Run the active presentation workflow's package, layout, font, and import checks after the final repair. Validate the delivered bytes, not an earlier version.
      
      ## Native charts and their data
      
      When data editing is part of the deliverable, create a native chart with a functioning data source. Editable lines and labels do not by themselves provide editable chart data. Check the series values, category labels, chart relationships, and embedded workbook when portable PowerPoint data editing is expected. A native chart count or a successful package import does not prove that the workbook exists or agrees with the plot.
      
      In the observed Artifact Tool finalizer, `nativeChartTargetApplication: 'powerpoint'` accepted literal native charts but did not create a workbook even with `materializeLiteralChartWorkbooks: true`. For that version, `'portable'` together with `materializeLiteralChartWorkbooks: true` materialized and validated a workbook from complete literal data. Verify the installed API and the resulting package instead of relying on the option name. Preserve existing formulas and linked data sources; do not replace them with a literal snapshot merely to pass a check.
      
      Inspect both literal series (`c:numLit`) and referenced series caches (`c:numRef/c:numCache`) when validating OOXML. Follow the chart's relationship to its workbook and the series' referenced cells; do not assume a particular embedded filename or worksheet.
      
      ### Saving can rewrite data representations
      
      A native edit/restore test, and even a later geometry-only COM save, can reserialize `0.103` as `0.10299999999999999`. Numerical equality within tolerance and exact workbook/cache agreement are different checks. In the observed workflow, this changed representation failed the finalizer's precision check.
      
      If reconciliation is needed, first establish the authoritative source for that chart: the original decimal values for a literal-data build, or its actual referenced worksheet for a workbook-backed chart. Confirm point counts, series/category identities, and numerical agreement within a justified tolerance before restoring the source representation. Never round every chart number or weaken validation to conceal a real discrepancy. If source and cache disagree materially, resolve that conflict before exporting.
      
      For a layout-only edit, preserve the chart and embedded workbook parts byte-for-byte where practical. Otherwise recheck them after the save. Run the final package validator and export only after the last repair has succeeded.
      
      ### Compact chart labels need native inspection
      
      Inspect after changing chart fonts: PowerPoint can wrap a short series name or split a numeric label across lines. On affected data labels in the observed COM workflow, `Format.TextFrame2.WordWrap = 0` and `Format.TextFrame2.AutoSize = 1` resolved wrapping. Apply label styling, then check and position the updated text bounds. These are targeted remedies, not blanket settings for every text box.
      
      Set axis number formatting explicitly when its default obscures the intended precision. Inspect unwanted leader lines and marker contrast. Changing marker style reset marker colors in the observed build, so restore foreground/background colors when that occurs. A failed `PlotArea.InsideLeft` setter is a reason to inspect the supported chart-layout API, not to retry the same call repeatedly.
      
      ## Canvas and group geometry
      
      For a schematic that should read as level, derive node centers, label baselines, and connection ports from shared horizontal and vertical guides. Equal bounding boxes alone do not align the visible artwork. Check each arrow's final segment and its contact with the node silhouette. Use orthogonal branches where they clarify the relationship, and matched radii with horizontal or vertical end tangents for symmetric loops. Preserve deliberate diagonals when the structure calls for them.
      
      Inspect this geometry in the native export at manuscript width. In a September 2026 Stellar proposal revision, individually reasonable coordinates produced sloping venue links, uneven title baselines, and an asymmetric bot loop. Shared guides and equal-radius corners corrected the appearance while preserving the mechanism. Native connector tests had passed before this visual correction; editability and visual alignment were separate checks.
      
      Align paired comparison rows by deriving both rows' node centers and their connection ports from the same guides. When one downstream node serves both rows, its outgoing edge can leave from their midpoint if that edge represents their joint result. Place a fan-out junction clear of the source node's silhouette; a technically attached route can still overprint that node's border. Inspect the rendered fork and arrowhead contact alongside endpoint attachment flags. A small geometry check can assert the intended row equality, axis alignment, and attachment with a declared tolerance. Visual inspection remains necessary.
      
      Capture top-level object positions before changing the slide size. In the observed PowerPoint COM workflow, increasing `PageSetup.SlideWidth` recentered the existing groups by half the increase. Applying the planned panel shift afterward displaced the result twice and clipped its right edge. Compare before/after absolute positions rather than assuming a canvas change leaves objects fixed.
      
      For a narrow spacing adjustment, a targeted OOXML edit can preserve chart data and other unaffected parts. Change the slide size and the intended panel transforms in a new copy, then compare package parts and render. Group `a:xfrm` uses both parent `off`/`ext` and child `chOff`/`chExt` coordinates; changing only one extent can unintentionally scale its children. Maintain the intended mapping and account for panel titles that live outside the moved group. Do not apply a case-specific pixel shift to every group.
      
      After a wider canvas or changed group transform, inspect clipping and the effective font size at the document width. See [panels-and-results.md](panels-and-results.md) for the visual tradeoff between separation and scaling.
      
      When deleting a footer or secondary strip, remove its separator, arrows and group definition as well as its text. Reduce canvas height to the remaining artwork with a modest margin, keeping canvas width, retained object coordinates and font sizes fixed. This preserves printed text size when the figure is inserted at the same document width. If the width or object scale changes, recalculate printed text size. Update the builder, expected canvas in validators, and export aspect ratio together. Native PDF export may quantize page dimensions slightly; use a small declared geometry tolerance and inspect the rendered bounds rather than requiring exact floating-point equality with the slide size.
      
      ## Native rendering helper
      
      On Windows with PowerPoint installed, the bundled helper opens a single-slide PPTX read-only, exports PNG and PDF, and prints a JSON inventory. Supply absolute paths. It refuses to replace existing outputs and does not edit the input deck.
      
      ```powershell
      & '<skill-root>/scripts/render_powerpoint.ps1' `
        -InputPptx '<absolute-path>/figure.pptx' `
        -OutputBase '<absolute-path>/exports/figure' `
        -PngWidth 2400
      ```
      
      An optional `-ReportPath '<absolute-path>/build/native-render.json'` saves the inventory outside the deliverable folder. A normal run creates `<OutputBase>.png` and `<OutputBase>.pdf`. Use fresh output and report paths for a revision or failed partial export. Check each stage's exit status before using its files: a finalizer can write an artifact and then fail while writing its receipt. File existence alone is not success. The helper supports one figure slide deliberately; use a presentation renderer for a multi-slide deck.
      
      The inventory records native shapes, text-bearing objects, connectors, charts, groups, pictures, and `otherObjects` for remaining object types such as tables or placeholders. A nonzero `otherObjects` count is not itself an error. The `charts` count reports native chart objects, not embedded workbooks or successful data editing. Native counts cannot prove that every required label is editable, that a connector follows a particular node, or that the content is correct.
      
      ## Check actual editing
      
      On a disposable copy of the final source, select and change representative text, change a meaningful shape property, and move a connected node. Confirm its connectors still attach to the correct endpoints. Inspect whether semantic groups can be moved together and ungrouped for detailed changes. Restore or discard the disposable copy; keep the final source unchanged by the test.
      
      For scripted COM checks, name the exact objects being checked, verify the change, and restore it. Example for a known shape:
      
      ```powershell
      [double]$originalLeft = $shape.Left
      $shape.Left = $originalLeft + 2.0
      if ([math]::Abs([double]$shape.Left - ($originalLeft + 2.0)) -gt 0.1) {
          throw 'Node movement did not take effect.'
      }
      $shape.Left = $originalLeft
      ```
      
      This movement check alone does not establish connector behavior; inspect the relevant `ConnectorFormat` attachments and the moved rendering as well.
      
      For a native chart, also change and restore a known series value on the disposable copy and check the visible update. Test the workbook's Edit Data path when that behavior is required. Record the specific behavior tested; do not present a COM series test alone as proof that workbook editing was tested. Discard the test copy or validate any serialization changes before using it as a source.
      
      Close only the deck opened by the task. Avoid terminating a PowerPoint process that may own the user's work. The helper exits the application only when no PowerPoint process was present at startup and no presentation remains open.
      
      ## Inspect the final export
      
      Open the final PNG or rendered PDF. Look for text clipping, changed line breaks, missing masks, wrong arrows, weak contrast, and accidental overlap. Inspect at the intended print or web size as well as full resolution. For a document figure, font sizes and page area matter more than how large the slide looks in the editor.
      
      Retain the native PPTX even when delivering a PDF or PNG. A PDF may be vector while some included assets remain raster; describe that accurately. A future author who changes the PPTX must regenerate the viewing export. Keep generation code when useful, but identify manual PPTX edits so regenerating does not erase them.
      
      Native effects and vector exports are separate properties. In an observed PowerPoint export, editable panel shadows became raster images in the PDF while foreground text and geometry remained vector. When effects are used, inspect the actual PDF before describing it as wholly vector. Use a native effect in the PPTX when editability is required. If a destination requires an entirely vector figure, simplify or remove the effect and verify the export.
      
    • panels-and-results.md 6.6 KB
      # Explanatory panels with compact results
      
      ## Give each panel a reading role
      
      An introductory figure can answer two questions together: what does this work enable, and why should the reader care? Let the schematic carry the organizing idea and let a compact measured result provide scale or a consequential contrast. Use this composition when those roles support the same takeaway; a result panel is not mandatory.
      
      A result already discussed later can still earn space in the first figure. Keep the selected comparison interpretable at the smaller size. Move the full experimental inventory and inferential details to their existing table or Results section. Do not hide a qualifier needed to understand the visible claim.
      
      Apply the same decision to proposals and READMEs without copying the paper template. A proposal might pair its mechanism with clearly identified preliminary evidence. A README might pair a workflow with a verified output example. Neither needs a performance chart merely to look impressive.
      
      ## Allocate width and whitespace independently
      
      1. Establish the usable width after outer margins.
      2. Reserve a gutter that reads as separation between panels.
      3. Divide the remaining width according to each panel's information load. For a requested large schematic and narrow result, about 3:1 is a reasonable first draft.
      4. Inspect the actual visible bounds, including axis labels, markers, legends, annotations, and panel titles. Plot-area rectangles alone underestimate crowding.
      5. Compare the inter-panel gap with gaps inside the schematic. The two panels should read as distinct units without losing their relationship.
      
      Do not assign 75 percent and 25 percent of the full usable width and then discover that the gutter has no room. Avoid a divider when whitespace already separates the panels; when they look crowded, test more whitespace before adding a line.
      
      Expanding the slide can preserve the drawing's internal geometry, but it reduces every label's printed size when the figure is still inserted at the same width. Compute:
      
      `printed font pt = slide font pt * document figure width pt / slide width pt`
      
      Inspect the resulting manuscript page or README viewport. Widening the canvas is a tradeoff, not free space. Tighten low-value content if the final labels become hard to read. See [native-powerpoint.md](native-powerpoint.md) for canvas resizing and group-transform pitfalls.
      
      ## Separate panels without breaking the reading path
      
      When a wide schematic reads as one uninterrupted strip, give its conceptual regions visible boundaries while preserving the cross-panel relationships. Start with gutters and aligned headings. Pale surfaces, a thin border, or a restrained shadow can add separation when whitespace alone is insufficient. These are optional treatments, not a requirement for cards or three columns. Keep them subordinate to semantic colors, and inspect at manuscript width so the treatment does not become visual clutter or disappear entirely.
      
      For the PDF export implications of shadows, see [native-powerpoint.md](native-powerpoint.md#inspect-the-final-export).
      
      Route cross-panel edges through the gutters to explicit destinations. An edge feeding a region with several readouts should not accidentally point at one readout and imply that only it receives the result. Keep qualifications next to the object they qualify; a nearby label from another row can appear to complete the wrong sentence even when their text boxes do not overlap.
      
      ## Keep the numerical claim attached to its estimator
      
      Read plotted values from an authoritative table, board, or experiment artifact. Prefer a parser shared with the paper's existing plotting code. Retain the input locator, series labels, transformation, and precision beside the builder. Do not infer values from image pixels when source data exists.
      
      Before choosing the chart or caption, establish:
      
      | Question | Why it matters |
      |---|---|
      | What metric, units, and sign are plotted? | A positive gap needs a named reference and direction. |
      | Which methods, population, and observation budget are compared? | A baseline or prefix change can alter the meaning. |
      | How are runs, seeds, or predictions aggregated? | Differences of reported means may differ from a registered paired estimator. |
      | Are these point estimates or an inferential statement? | A significance label must correspond to the displayed estimand and applicable test. |
      | Which test family or diagnostic scope applies? | A result in another state, corpus, or family does not establish this claim. |
      
      For example, one CatchBench curve showed a quarter-prefix difference of printed mean AUCs of `0.113`, while the registered paired effect was `0.101` under a different averaging procedure. Both can be correct. Labelling the plotted `0.113` as the registered effect would be incorrect. The accepted inset showed descriptive gains and left inference in Results.
      
      Preserve small negative values, ties, and non-monotonic behavior. Do not clip an inconvenient point or smooth a measured series into an invented trend. Two curves alone do not establish a cross-corpus significance claim. An unregistered comparison is not automatically evidence of no effect; it simply does not support that registered-test claim.
      
      ## Preserve just enough semantic guidance
      
      A small legend or caption phrase can carry essential interpretation: nodes are steps, arcs are dependencies, hatching is unavailable future evidence. A candidate prediction must not look like a target label supplied to an evaluated method. Separate corpora must not look like a matched run across stages unless the source supports that pairing.
      
      Remove repeated inventories and decorative text before removing these distinctions. The aim is fast understanding with sufficient evidence, not minimum word count.
      
      ## Keep revisions proportional to the feedback
      
      If the author accepts the content and asks for more separation, change the spacing and inspect its consequences. Preserve the source values and chart/workbook parts where possible. Recheck the panel balance, label placement, effective font size, and final export. Do not reopen a settled concept or request approval for a reversible layout adjustment already authorized by the feedback.
      
      When an independent review is requested, reconcile recommendations against the source and the user's requirements. Record accepted and rejected suggestions with short reasons, and identify the reviewed version. A later result-panel restoration or spacing change is a new local revision, not evidence that the earlier reviewer inspected those final bytes.
      
      See [iteration-example.md](iteration-example.md) for the complete design progression and its case-specific measurements.
      
    • position-and-concept-figures.md 5.3 KB
      # Position and concept paper overview figures
      
      Use this reference when a paper's main contribution is a position, framework, taxonomy, or research agenda rather than a single implemented method. The first figure should make the paper's argument easier to reconstruct. It need not summarize every section.
      
      ## Compress the argument before the text
      
      Start from the smallest claim graph that explains why the position follows. A useful graph often contains:
      
      - a concrete carrier, such as an agent task, record, scientific object, or deployment setting;
      - the necessary concepts and relationships, including any dependency or boundary central to the claim; and
      - a consequence, diagnosis, decision, or unresolved case that gives those concepts a reason to matter.
      
      These are semantic roles, not required panels or a fixed left-to-right layout. A paper may need a different structure. The test is whether a reader can trace the argument through visible objects and relations rather than recover it from a list of labels.
      
      Avoid two opposite failures:
      
      | Failure | What remains visible | What is lost |
      |---|---|---|
      | Taxonomy display | The names of the concepts | Why they belong together, how they constrain one another, and what follows from them |
      | Poster-like minimalism | A polished contrast or slogan | The evidence or reasoning that makes the position informative |
      
      A dense operational workflow can fail for a third reason: it may preserve the logic but bury it in field names, repeated annotations, and procedural detail. Reduce that reading load by replacing prose with object silhouettes, attached edges, containment, a small exception path, or a diagnostic output. Do not delete the few relations that let the reader explain what leads to what.
      
      ## Run a meaning checkpoint after simplification
      
      Treat semantic density and text density as separate variables. A figure can contain few words and still carry a rich argument; it can also contain many correct words and explain very little.
      
      After a substantial simplification, inspect the image without relying on the caption and ask:
      
      1. Is the concrete setting or object identifiable?
      2. Can the reader trace the central dependencies, intervention, or information flow?
      3. What consequence or limitation does the figure show?
      4. Are illustrative objects, proposed mechanisms, and measured evidence distinguishable?
      
      If any answer is no, restore the missing object, relation, consequence, or claim boundary before polishing the layout. A cleaner canvas is not an improvement when it no longer explains the paper's position. Use the caption for qualifications and scope, but do not require it to reconstruct the entire claim graph.
      
      When another model is requested, use the image-only first-read protocol in the skill's inspection step instead of relying on this self-check.
      
      ## Refine in semantic stages
      
      Keep revisions ordered so visual polish does not reopen settled reasoning:
      
      1. **Argument structure:** establish the claim graph, concrete carrier, and claim boundaries, using the manuscript's names and order for its central concepts.
      2. **Content economy:** remove repeated prose and secondary detail while preserving the graph.
      3. **Visual polish:** align corresponding edges and baselines, regularize padding, reclaim unused bands, and assign colors by semantic role.
      4. **Navigation aids:** add compact numbering, markers, or section references only when they help a reader find the matching text. They can change label widths and heading hierarchy, so rerun the stage 3 alignment and text-bound checks afterward.
      
      Use the skill's visual-only freeze-and-compare rule to keep alignment, footprint, and palette changes separate from semantic redesign. Recheck rendering and text bounds after reflow.
      
      Assign colors by semantic role, starting from [default-palette.md](default-palette.md) unless the document has its own palette. Give each accent one meaning across the figure; when two different states share an accent, recolor one of them. Keep structural objects and headings neutral when coloring them would imply a category membership they do not have.
      
      ## Link the figure back to the paper carefully
      
      When a framework has a stable set of central concepts, use consistent names, ordering, or small markers to make the mapping visible. A marker can connect a concept to a table, checklist, contribution, or later analysis without repeating that material in the figure. Do not add a marker for every subsection.
      
      Section references are optional navigation aids with the limits described in the skill's design step. Numbers printed inside the artwork cannot follow a renumbered manuscript. When the destination is LaTeX, prefer carrying section references in the caption with `\ref`, and recheck any number printed in the figure after structural edits.
      
      ## Case note
      
      For reference selection, follow [reference-search.md](reference-search.md). In the Auditable Agents case, AgentDojo supplied a concrete agent and tool setting, and an authenticated-delegation position paper supplied a focal-document hierarchy; no artwork was copied. The paper's own enterprise example, five dimensions, palette, and layout belong to that case and are not defaults.
      
      Case evidence: `papers/aisummit-2026-auditable-agents/figure-src/audit-redraw-v1/` through `audit-redraw-v7/` in `internal-writing`, 2026-09-13. The guidance above is usable without those project files.
      
    • proposal-figures.md 11.9 KB
      # Proposal figures that belong in the manuscript
      
      Use this reference for proposal overviews, aim-integration figures, and revisions whose main problem is reading load or visual mismatch. It supplements the source and reference workflow in [document-contexts.md](document-contexts.md) and [reference-search.md](reference-search.md).
      
      ## Give each figure a distinct job
      
      Read the existing figures and their callouts before adding another overview. A motivation figure may already explain the problem and how it divides into aims. An additional integration figure earns space by showing a shared scientific object, dependencies, or how an intervention changes an example. Three aim titles linked by arrows usually repeat the outline without explaining the integration.
      
      A three-part input, framework, and evaluation composition can serve a different purpose. It works when those relationships carry the argument and a dominant technical center contains a concrete domain object, as in the selected OpenLibSignal framework example. Judge that content rather than rejecting a panel count.
      
      For a graph view, identify what each node and edge represents before choosing icons. Make integration visible through actual dependencies and shared records; a colored path cannot substitute for a missing scientific dependency. Retain enough structure to distinguish a graph from a three-step workflow. Do not imply a mandatory execution order merely because the aims are numbered.
      
      Choose the smallest example that still shows each relevant aim doing its distinct work on that structure. Simplifying an integration figure to one aim's mechanism, with the other aim labels attached, loses its purpose. Conversely, merging every aim's detailed worked case into one graph can bury the integration. Defer cases already explained nearby while retaining the versions, branches, and cross-aim links necessary for the inference.
      
      These compact-figure lessons do not establish a template for a large research or ecosystem overview. The latter may need domain scenes, resources, evidence, and multiple reading levels. The author-selected FIRE and OpenLibSignal figures remain references for that separate design problem.
      
      ## Keep aim names separate from operations
      
      Read the current subsection headings and definitions. An operation such as re-anchor, repair, or trace is not the aim's official name. Place action words by the relevant marks instead of presenting them as replacement aim titles. A fluent substitute can also change the mechanism: reducing reliance is different from redirecting a request or changing partners.
      
      When shortening is appropriate, preserve distinctive source wording and keep a label-to-source mapping in the working brief. If the author flags ambiguity or requests exact naming, restore the full visible titles and reflow around them. Compare the exported title text with the current headings, ignoring line breaks only for an exact match. Correct wording hidden in speaker notes does not fix the image. Title consistency and quick comprehension require separate checks.
      
      ## Design for a brief reviewer skim
      
      Use the author's reading budget. In the motivating case, the target was no more than 30 seconds. This is a design constraint, not an agency rule or evidence that comprehension was measured.
      
      At manuscript width, the first glance should reveal the main object and organizing relationship. After the short caption, the reader should be able to explain the central relationship or change, including why the aims belong together when integration is the figure's job. If that requires reading every operation or decoding many symbols, revise the hierarchy or remove details that do not support the claim before reducing type size. A compact before/after example often helps a mechanism figure. A larger ecosystem overview can support a quick structural skim followed by optional detail reading; preserve breadth when breadth itself is the argument.
      
      When a reviewer pass is already planned and comprehension is in doubt, use the image-only first-read check in the skill's inspection step, "Inspect semantics, appearance, and editability". A technically correct graph can still fail this reading task. Judge retain, revise, or remove by the relationship it adds beyond existing figures and nearby prose, not by the effort already spent drawing it.
      
      Native editability does not require a figure made entirely of prose inside boxes. Use recognizable scientific objects, useful subgraphs, or supported plots to carry information. Each borrowed element should answer a reader question; adding a small chart or icon solely to make the figure look less generated adds clutter.
      
      ## Borrow the document's visual vocabulary
      
      For a short sponsor proposal with room for one figure, try a dominant domain scene with a smaller evidence or evaluation area when the argument centers on shared objects. In a September 2026 three-page Stellar proposal, equal research-question columns felt like stacked modules. A shared exchange scene made the trading relationships visible, and the author accepted the subsequent alignment revision for this lightweight proposal. This is a composition option, not a universal one-figure layout.
      
      Small domain miniatures can help readers recognize what the system handles: a transaction receipt, a bid/ask book, paired asset symbols, or an instrument image. Integrate them with the corresponding objects and preserve the main reading path. Use consistent native glyphs for simple schematic objects; retain authentic crops as replaceable pictures when their actual detail matters. A schematic remains illustrative. Real plots and screenshots need traceable sources, and invented prices, transaction hashes, or trend lines must not serve as visual texture. Keep details that survive manuscript scaling and avoid adding captions or panels merely to accommodate ornament.
      
      Inspect actual awarded-proposal pages when they are supplied, then compare candidate choices with the current manuscript's figures. Match more than color: typography, stroke weight, information-block shapes, arrow treatment, and actor/tool icons affect whether a figure belongs in the document.
      
      If the author points to recent paper figures, inspect those actual assets as additional donors. Equal node footprints, outside labels, simple record glyphs, and clear connection ports can transfer to a compact proposal graph. Keep the proposal's established color meanings.
      
      The author clarified that the unwanted "AI feel" comes from adding text where the figure does not need it, often to signal rigor. Audit that text allocation before restyling shapes or fonts. In the earlier revision, reusing original icons improved document continuity but did not resolve this editorial problem. The preferred examples support rich visual content and varied shapes with purposeful labels. A figure can contain many useful objects without needing an explanation beside each one.
      
      Read [proposal-style-exemplars.md](proposal-style-exemplars.md) when choosing a proposal composition or interpreting this author's visual preferences. The useful distinction is whether visual content explains the domain and its relationships. A column of generic icons and prose can feel templated; a structured panel with an informative map or mechanism can earn its space. Preserve the current document's color meanings while adapting the preferred examples' hierarchy and visual content. A preference for their composition does not automatically request their exact palette.
      
      | Candidate element | Reuse when | Preserve |
      | --- | --- | --- |
      | Actor or tool icon | It identifies the same kind of source or participant | Its role as context or as a graph node; adding an actor icon must not silently redefine information-record nodes |
      | State block or graph fragment | It explains the same information object or dependency | Node and edge meaning, version history, and the actual scope of any repair |
      | Plot or result panel | Its evidence supports this figure's specific claim | Values, axes, uncertainty, conditions, and preliminary-versus-planned status |
      
      Prefer original editable components. If only a PDF or image is available, an authorized crop can be a replaceable picture while surrounding labels and connections remain native. Inspect the crop for neighboring arrows, borders, clipped details, and background halos. Record the source region and disclose the picture boundary instead of claiming every component is editable. Do not copy whole panels when a smaller element serves the purpose.
      
      ## Keep graphic semantics and print behavior intact
      
      Separate anomaly flags, withheld use, repair, and verification visibly. A route described as held should end at the hold marker; an arrowhead entering its destination still suggests delivery. Put feedback at the object it actually updates. If later outcomes calibrate future scores, arrows into historical records can imply an unsupported retroactive change.
      
      Keep ownership and dependency distinct. Each aim may detect local anomalies while only one performs graph-wide risk assessment or repair. A selected region can span record types without transferring their local checks to that aim. Show the dependencies needed to justify the selection. In the motivating case, a retrieved claim directly supported memory as well as a handoff; omitting the first edge hid the cross-aim coupling.
      
      Separate artifacts eligible for repair from outcomes used as attribution evidence. An already delivered outcome need not be a repair target merely because it is downstream. Likewise, revising one field of a record does not certify every field or dependency as correct. Avoid success marks that imply more than the source supports.
      
      Distinguish a stored dependency's direction from the direction in which it is inspected. A `trace back` label beside a forward arrow can contradict the visible reading. Separate the traversal annotation from the dependency, or otherwise clarify the two directions through the marks. A detached upstream chevron was one case-specific solution, not a required notation.
      
      Inspect lines against visible text and symbols, especially at converging routes. Place labels near the relationships they name without making adjacent action labels read as one phrase. Inspect selection boundaries against entire node silhouettes, not just centers. Enlarging a text box does not remove an ink overlap; when label nudges trade one collision for another, reroute the edge or reflow the local cluster.
      
      For a revision inside a tight page budget, check the current compiled manuscript. A single extra caption line can split a running example or move later figures. Compare page count, the caption's page and vertical position, and neighboring content. A wrap is useful when it returns space to prose while keeping the graph legible, not as a universal placement rule.
      
      ## Record partial improvement honestly
      
      The author considered the earlier reuse revision still less than ideal and subsequently challenged naming, graph integration, and appearance. Restoring familiar elements improved continuity but did not settle those objections. Record the specific improvement and any remaining uncertainty without promoting a reviewed draft to an approved template.
      
      Keep visual acceptance separate from semantic review and native-object checks. A reviewer pass or successful PowerPoint edit test does not establish beauty or fast comprehension. Label reviews by the artifact version they inspected; a later local style revision does not inherit a new external review. Deliver the actual preview with the source so the author can assess the result.
      
      Case basis: September 2026 Figure 2 iterations reused the manuscript's Figures 1, 3, and 5, then inspected recent internal-writing paper figures. Later revisions restored full aim titles, simplified the shared graph, and corrected repair and attribution semantics. No new large overview figure was designed or validated in this round. Figure numbers, node counts, colors, font sizes, wrap fractions, and particular annotations belong to the case rather than the general workflow.
      
    • proposal-style-exemplars.md 7.4 KB
      # Author-selected proposal examples
      
      The author explicitly identified these three figures as preferred examples on 2026-09-05. They guide proposal composition and the use of concrete visual content. They supersede any inference that the preceding Figure 2 revision's plain rectangles and serif type are the author's ideal style.
      
      Locators are relative to the `NSF-Proposal-Template-Yue` repository. Open the actual PDFs when that collection is available. Do not copy the private artwork into this skill or send additional proposal content to a reviewer by default. The descriptions below keep the design lessons usable without the local collection.
      
      ## NSF FIRE: a visually concrete research overview
      
      **Preferred asset:** `proposals/awarded/nsf-fire/figure/main_figure.pdf`, page 1. Its inclusion and overview caption are in `00-project-description.tex` beside that proposal's four-thrust introduction.
      
      The figure connects people and an agent at the top, four research thrusts in the middle, and the wildfire/transportation setting along the bottom. Each thrust contains visual substance: a spread surface and map, a road-network scene, information delivery to different recipients, or a reporting loop. Short labels identify the mechanism without replacing it with a paragraph. Rounded colored regions organize the material, and the broad scene establishes the application setting.
      
      Borrow the hierarchy from beneficiaries and coordination to mechanisms and domain setting. Use a concrete visual object within each research component, with connectors that explain a relationship. Do not copy the four-column layout into a proposal with a different integration story, or assume that a surface, map, or scene is measured evidence merely because it looks scientific. Tiny inset labels and every decorative element need not transfer. This inspection did not establish a matching editable master; the nearby `paper-visuals.pptx` has different slide content.
      
      ## OpenLibSignal: an ecosystem with visible substance
      
      **Preferred asset:** `proposals/awarded/pose-phase1-openlibsignal/figs/ecosystem-pic-all.pdf`, page 1. `000description.tex` includes it as an ecosystem overview; its caption identifies research outcomes at the top, integrated tools in the middle, and proposed aims at the bottom.
      
      The upper branching structure organizes research directions and publications. The center makes dataset, simulator, and algorithm components visibly interdependent, surrounded by concrete tasks, scenarios, and simulator imagery. The bottom ties aims to result insets, development/community material, and governance. This is a dense figure with several reading levels, not a compact mechanism diagram. The publication legend distinguishes the PIs' group from the wider community. When showing a breadth of outcomes, keep team contributions distinguishable from external work so the inventory does not imply unsupported team credit.
      
      Borrow the large-scale organization and the connection from existing resources to research and proposed work. Screenshots, graph structures, and selected evidence give the ecosystem specific content. The reader can recognize that organization before reading each leaf. Allocate sufficient page area or select a smaller relevant part for a compact destination; do not reproduce every label at unreadable size. The visible publication, organization, and result labels belong to the reference and do not establish claims for another proposal.
      
      The nearby `ecosystem-pic.pptx` is a component source candidate, not a verified master of the entire preferred PDF. Inspect it before editing: its single slide does not contain all the preferred PDF's aim and outcome content.
      
      ## OpenLibSignal plot.pdf: a compact framework with a domain anchor
      
      **Preferred asset:** `proposals/awarded/pose-phase1-openlibsignal/figs/plot.pdf`, page 1.
      
      Despite its filename, this is a conceptual framework diagram, not a statistical result plot. It connects community engagement, a computational framework, and performance criteria. The center is visually larger and uses a prominent map to make the application concrete. Colored headings, short lists, and strong directional arrows establish a fast reading path; rounded panels and shadows are part of the liked example.
      
      Borrow the hierarchy between a dominant technical center and concise input/evaluation sides when that matches the proposal's argument. The map carries domain context rather than an unsupported performance claim. In this reference, the map crosses the dashed box borders and crowds the ends of titles and subtitles. Preserve the map's emphasis while improving text clearance when adapting it; an intentional frame overlap does not require accepting obscured text. Do not turn listed evaluation criteria into measured outcomes.
      
      `figs/plot.pptx` exists, but its wording differs from the preferred PDF, including the listed metrics. Use the PDF as the visual reference; verify the intended text from the active proposal before adapting the PPTX. No active inclusion of this PDF was found in the inspected LaTeX sources, so the file is a directly selected style example without an inferred current figure number.
      
      ## Supplementary paper donors for compact graphs
      
      The author later directed the Figure 2 revision to recent editable-figure work in `internal-writing`. The following assets were inspected as donors; they were not selected as large proposal-overview templates. Paths in this subsection are relative to that repository:
      
      - `papers/aisummit-2026-graph-view-for-agents/figure/grade-first-figure-a-v2.png`: equal graph-node footprints, labels outside nodes, light strokes, and recognizable record/resource shapes.
      - `papers/iclr-2027-auditablebench/figure/CatchBench-figure1.png`, `CatchBench-gold-mechanism.png`, and `CatchBench-research-landscape-wrap-v2.png` in the same figure directory: Arial hierarchy, compact graph marks, and local highlighting.
      
      Inspect the current assets and, when useful, their editable sources before borrowing. These locators identify inspected versions, not an instruction to rescan all recent papers. Transfer useful hierarchy and graph vocabulary while preserving the destination proposal's palette. The resulting Figure 2 still needed separate corrections to aim names, dependencies, and reading load; the donor inspection alone did not establish acceptance.
      
      ## What the preference changes
      
      The common lesson is meaningful visual richness: identifiable domain objects, scientific or system structure, purposeful labels, and a hierarchy that guides the first glance. The author's later clarification locates the unwanted "AI feel" in unnecessary text added to signal rigor. Shape geometry is not the criterion. Keep visual substance while removing redundant explanation; a dense collection of useful evidence does not justify explaining every element again.
      
      Select the example by the figure's job: integrated research overview, ecosystem argument, or compact framework. Keep editable text, arrows, and semantic objects alongside replaceable image assets when a mixed composition is appropriate. Apply the manuscript's established palette, or [default-palette.md](default-palette.md) for a new identity, unless the author asks for a different color treatment.
      
      Visual preference is not verification of the references' scientific claims, partner commitments, or empirical results. These examples also do not establish a universal federal-proposal style or require every proposal figure to be as dense as the ecosystem overview.
      
    • reference-search.md 7.1 KB
      # Find references before choosing a composition
      
      For a new figure or substantial redesign, search for relevant examples after understanding the source and before committing to a layout. The goal is a defensible design choice, not an exhaustive literature review. Reuse a relevant reference already inspected in the current task; a small correction to an accepted layout does not need another search.
      
      ## Route by the requested document
      
      | Target | Search first | Inspect | Prefer |
      |---|---|---|---|
      | ML/AI paper | Related papers in NeurIPS, ICML, and ICLR; use other leading field venues when appropriate | The actual first figure or matching method figure, its caption, and nearby introduction | Similar scientific contribution and figure purpose, with a clear problem, example, mechanism, or result |
      | GitHub README | Current GitHub Trending in a relevant category, then related active repositories | The rendered README, actual hero/workflow asset, opening description, and quick start | A recognizable user benefit, honest input/output example, and readable composition at README width |
      | Proposal | The author's selected proposal examples, then the local awarded/funded collection and its indexes as needed | Overview or aim figure with surrounding narrative, funding context, and available editable source | A match to the figure's job and the author's preference; agency/program and audience provide additional context |
      
      ### Paper references
      
      Search with both topic and figure role, for example `agent evaluation benchmark ICLR overview figure` or `scientific machine learning framework ICML`. Prefer official proceedings and author-hosted originals. Confirm publication venue and year from the source when calling a paper accepted or published; an arXiv upload alone does not establish acceptance.
      
      Read the actual figure rather than inferring its content from the abstract. A first figure may be an example, architecture, or result plot, and those serve different purposes. Inspect its relation to the surrounding text. Record the page and figure number so another agent can find it again.
      
      ### README references
      
      Treat "trending" as a dated observation. Check the current listing and record the date and daily/weekly/monthly window when available. Stars or historical popularity alone do not prove a repository is currently trending. Activity and visual quality are separate selection criteria.
      
      Use trending projects as candidates, then inspect related repositories whose users and outputs resemble the target. A popular project with a decorative logo may offer little guidance for an explanatory figure. Prefer the actual rendered README and source asset over a search thumbnail. Record a commit or retrieval date where practical because README layouts change.
      
      ### Local awarded proposal references
      
      Start with the author's [selected proposal examples](proposal-style-exemplars.md) and choose by figure job: integrated research overview, ecosystem argument, or compact framework. Search further when those examples do not serve the current figure's role. Use the user-named collection, project README, or existing local index and targeted searches for terms such as `awarded`, `funded`, agency names, and the proposal topic. Prefer indexed final awarded versions over unrelated drafts. Do not assume every proposal in a figure bank was funded.
      
      Use the user's explicit identification, a trusted local award index, or an award record to establish funding status. Otherwise mark the status unverified. If the collection path is unknown after checking the available project indexes, ask for its location while continuing source analysis and any available reference inspection. Do not invent a standard local path or scan unrelated personal folders.
      
      When the author asks to retain reference preferences or lessons in a local skill, record the source locators and concise design descriptions. This does not automatically authorize copying private artwork, PDF pages, or verbatim passages into that skill, a public reference bank, or an external prompt. Use any existing authorization for the actual material and destination; keep broader proposal contents outside a limited figure review.
      
      Inspect the current proposal's figures alongside the awarded examples. The former establish the document's own typography, palette, object shapes, and icon vocabulary; the latter can supply useful composition ideas. Resolve figure numbers through the current manuscript's actual asset inclusions. A similarly named small export or an older PPTX may contain different artwork. When the author asks to reuse their own elements, look for the editable source first and retain exact asset or crop locations. See [proposal-figures.md](proposal-figures.md) for choosing what to reuse.
      
      An existing `figure-references/index.md` can help discover visual references, but its style tags do not establish award status. If no awarded example is accessible, state that limitation and continue with user-supplied or clearly labeled alternative references. Do not claim to have completed an awarded-proposal comparison.
      
      ## Select what transfers
      
      Inspect a small candidate set, often three to five examples, and select one to three useful references. Expand only if none solves the design question. These are effort guides, not quotas. The best structural reference and the best style reference may be different examples.
      
      Compare candidates on the following questions:
      
      - Does the figure serve the same reader decision and document role?
      - What makes its usefulness, novelty, or feasibility apparent?
      - How does it divide information between graphics, labels, caption, and surrounding prose?
      - Which relationships remain clear at the intended display size?
      - Can the useful structure be rebuilt as editable PowerPoint objects?
      
      Keep a compact reference record with the working figure brief:
      
      | Source and locator | Role in this task | What to adopt | What does not transfer |
      |---|---|---|---|
      | URL or local path, figure/page or README section, date/version as relevant | Structure, visual style, or concrete-example treatment | Specific hierarchy, evidence boundary, grouping, or input/output arrangement | Unsupported claims, domain-specific content, excessive density, or incompatible format |
      
      Recommend a composition in one or two sentences that connect the chosen references to the current reader takeaway. Proceed to a draft; reference selection is not a mandatory approval gate. When useful, show the inspected references or link their exact locations alongside the explanation.
      
      Use references to inform an original design. Reuse artwork from the author's own figures when requested or otherwise authorized, preserving its meaning and source. Do not assume permission to reproduce unrelated reference artwork or invent impressive-looking results. Acceptance, funding, and GitHub popularity help identify candidates but do not prove that their figures caused those outcomes.
      
      If web or local access is unavailable, describe which part of the search was not performed, use the material actually available, and continue within the authorized scope. A user instruction to skip research or use a supplied template takes precedence.
      
    • space-efficiency.md 3.7 KB
      # Space efficiency in document figures
      
      Use this reference when a figure leaves broad unused areas, one element forces an oversized canvas, or a paper/proposal layout needs a smaller footprint. The same principle applies to overview diagrams, mechanisms, plots, timelines, screenshots with callouts, and README visuals: preserve the information and reading hierarchy while avoiding space with no useful role. The preferred footprint depends on the destination, not the authoring application's slide default.
      
      ## Diagnose the source of the space cost
      
      Inspect at the intended document width. Mark the actual occupied bounds of each panel, not only text-box bounds or background shapes. Check for unused strips beside short labels, empty rows after a deletion, loose inter-panel gaps, and long footers that set the width for otherwise narrow content. A footer spanning the canvas can make an overall crop look tight while hiding a large empty region above it.
      
      Then inspect the document. Include the caption, float margins, wrapped line count, paragraph spacing, and nearby headings. Saving space inside the drawing does not necessarily save space on the page. A compact wrapfigure can still create an extra caption line, leave a heading at half width, or move a later float onto another page.
      
      ## Reflow before reducing type
      
      - Narrow or shorten the element that unnecessarily determines the canvas size. Keep all scientifically necessary labels and distinctions.
      - Rearrange branches or method groups to fit the intended slot. Preserve their relationships; reflow must not introduce a causal progression, an ordering, or a one-to-one assignment absent from the source.
      - Reduce the resulting canvas width or height. Keep retained content and physical type size stable where possible.
      - Preserve gutters that separate panels and clearances around labels, axes, and arrows. Useful breathing room serves comprehension; broad empty areas that do not do so should be reclaimed.
      
      Do not fill space with decorative objects, repeated conclusions, unsupported numbers, or extra detail. Do not stretch objects disproportionately or shrink all labels merely to fit an inherited aspect ratio. Uniform pixel coverage is not the objective.
      
      ## Return the space to the document
      
      If source width changes from W_old to W_new, keeping placed_width / source_width approximately constant preserves physical type size. For example, reducing an 820 px figure at 50 percent text width to 690 px at about 42 percent text width keeps the label scale almost unchanged. The freed width becomes prose width.
      
      If the placed width is left unchanged after cropping, the content becomes larger and the placed figure can become taller. That can be appropriate for readability, but it is not the same as reducing page cost.
      
      Recompile or render the current document. Inspect the full figure plus caption and the text wrapping beside it. Check the transition back to full-width text and subsequent section headings. Report any material pagination effect without promising that wrapfigure placement or cropping guarantees a lower page count.
      
      ## CatchBench example
      
      In the initial wrap version, the upper timeline's longest label ended near x=657 on an 820 px canvas. Wide method bands and a footer extended across the canvas, leaving a visible empty strip beside the timeline. The correction narrowed the canvas to 690 px and reflowed the method bands. Dates, sketches, works, and method names were retained. Reducing the LaTeX wrap width from 0.50 to 0.42 of the text block kept the original label scale and returned the removed strip to neighboring prose.
      
      The reusable lesson is to evaluate each panel and the final document together. These dimensions illustrate one repair; they are not default aspect ratios or mandatory whitespace thresholds for other figures.
      
  • scripts
    • render_powerpoint.ps1 5.2 KB · in bundle
  • SKILL.md 21.7 KB
    ---
    name: editable-figure
    description: Analyze source material, find relevant paper, README, or awarded-proposal references, and design concise figures as editable PowerPoint objects. Use for overview, mechanism, workflow, or hero figures when an editable PPTX is wanted, including simplifying dense drafts and combining a schematic with a compact result panel. Complements scientific plotting; does not replace screenshot capture, prompt-only work, or full slide-deck authoring. Completing a PPTX deliverable requires desktop PowerPoint on Windows or macOS for the native rendering and editing checks; the bundled export helper is Windows-only. Assessment and prompt-only requests need neither.
    ---
    
    # Editable Figure
    
    ## Overview
    
    Turn a document's central idea into a figure that a reader can understand quickly, then deliver a PowerPoint source that the author can actually edit. The common workflow is **analyze, find references, design, build native objects, inspect in context**. Paper, proposal, and README figures share this workflow but serve different reader decisions.
    
    ## Choose the scope
    
    Respect the requested deliverable. An assessment or prompt-only request does not require generating a deck. Once figure creation is authorized, use subsequent feedback to revise the artifact without repeatedly asking permission. Choose a reasonable composition and produce a reviewable draft when the source is sufficient.
    
    Use this skill for the figure's editorial and design decisions. If an installed presentation skill applies, read it for the current authoring runtime and validation requirements. Without one, use an available PPTX library or native PowerPoint automation and the principles in [native-powerpoint.md](references/native-powerpoint.md). This workflow works with Codex, Claude Code, or another capable agent; it does not depend on one model or desktop plugin.
    
    **Platform requirement.** Before starting a figure build, confirm that the session can use desktop PowerPoint on Windows or macOS for the rendering and editing checks in step 5. A PPTX library such as `python-pptx` or PptxGenJS can create native objects headlessly, and that is a real capability. It does not complete these checks: nothing on a headless machine can confirm the result opens, renders, and edits as intended, and an unverified figure is the failure this skill exists to prevent. The bundled `scripts/render_powerpoint.ps1` uses Windows COM automation; macOS requires an available native PowerPoint workflow instead.
    
    Without that capability, explain the limitation before building. Assessment and prompt-only work can continue unaffected. Offer `ci-mockup-figure` when its output meets the user's needs. Preserve an explicit PPTX requirement unless the user agrees to another deliverable; do not silently substitute a flattened image and call it an editable figure.
    
    Nearby workflows have distinct outputs:
    
    - `figure-prompt-builder`: prompts and reference-guided concepts, when that is the requested endpoint.
    - `ci-mockup-figure`: HTML mockups, screenshots, TikZ, or other code-native figure formats.
    - A plotting workflow: scientific data plots with reproducible data and axes.
    - A presentation workflow: complete decks. Here a slide is a canvas for a document figure, not a presentation with a cover.
    
    Do not replace an explicitly requested SVG, Illustrator document, or screenshot with PPTX merely because this skill is available.
    
    ## 1. Analyze the source and the reader's decision
    
    Read the relevant section and its neighboring prose, the document's purpose, existing figures, and primary artifacts behind any claims. Identify the intended placement and display width before allocating space. Read only enough of a large repository to establish the contribution, terminology, evidence, and constraints.
    
    Write a short working brief, in scratch space or beside the figure source when it needs to persist:
    
    | Decision | Record |
    |---|---|
    | Audience and placement | Who sees the figure, where, and at what size? |
    | Reader takeaway | One sentence the reader should remember after a quick look. |
    | Usefulness | What can the reader understand, evaluate, build, or decide because of this work? |
    | Necessary visual evidence | The example, relationship, mechanism, or measured result that supports that takeaway. |
    | Content allocation | What belongs in the drawing, caption, neighboring prose, or a table? |
    | Claim boundaries | Facts versus illustrative examples, proposed work, predictions, and measured results. |
    
    Apply the relevant mode in [document-contexts.md](references/document-contexts.md). Do not require every figure to summarize the whole project.
    
    For proposal figures, also read [proposal-figures.md](references/proposal-figures.md) for distinct figure roles, visible aim-name consistency, shared-graph semantics, and reading at manuscript width. It links the author's preferred examples. Keep lessons from compact integration figures separate from untested large-overview designs.
    
    For a benchmark, distinguish the system producing the record from the method being evaluated. For a proposal, distinguish planned capabilities from completed results. For a README, show the actual user workflow and supported behavior.
    
    ## 2. Find and study references before designing
    
    For a new figure or substantial redesign, run a focused reference search before choosing the composition. Read [reference-search.md](references/reference-search.md) and use the route that matches the document:
    
    - **Paper:** related top-venue papers, typically NeurIPS, ICML, and ICLR for ML/AI. Inspect the actual figure, caption, and nearby prose.
    - **GitHub README:** current relevant trending projects and related active repositories. Inspect the rendered README and its actual visual assets.
    - **Proposal:** start with the author's [preferred examples](references/proposal-style-exemplars.md) when they fit the figure's job. Search the local awarded/funded collection further when needed, using figure purpose and agency/program as selection criteria. Check indexes and known collection paths first.
    
    Select a small number of references and state what transfers: information hierarchy, a concrete example, a mechanism, or a useful visual structure. Keep exact source locators in the working brief. Popularity, acceptance, and funding are discovery signals, not proof of figure quality.
    
    Use the user's supplied references when appropriate. Reuse already inspected examples for small revisions rather than restarting research. If a source is inaccessible, report that boundary and continue with available material; do not invent search results or funding status. Once a direction is supported, proceed to design without adding a reference-approval checkpoint.
    
    ## 3. Design for the takeaway
    
    Make usefulness visible through an understandable problem, a consequential distinction, a concrete mechanism, or a supported result. Claims such as "powerful" and lists of components do not substitute for that evidence.
    
    Choose the visual structure from the message: an example with an intervention point, evidence becoming available over time, a bottleneck and proposed mechanism, input-to-output transformation, comparison, or another justified structure. Do not default every task to three columns, a lifecycle, or a grid of cards.
    
    For a position, framework, or concept paper, make the overview a compressed argument rather than only a taxonomy or a polished slogan. Retain the minimal claim graph and a concrete domain object needed to show why the position follows, even while removing prose. Read [position-and-concept-figures.md](references/position-and-concept-figures.md) for semantic compression, staged refinement, and links back to the manuscript.
    
    Use graphics to carry relationships. Text should identify objects and explain only what the graphic cannot. A useful starting point for a compact overview is a few focal elements, each with an object and a short question or label. Adjust to the source; this is not a fixed panel or word quota.
    
    For overview figures, retain a concrete domain object that participates in the mechanism, such as a circuit feeding measurements, rather than relying on a domain name in the title. Check that a reader can identify both the subject and the connection between panels. Shared variables, a carried example, or explicit input/output edges can establish that connection; adjacency and panel letters alone cannot. A section pointer printed in the figure, such as `Section 3.2`, helps navigation but does not explain the mechanism. Recheck it in the compiled document after structural edits, because section numbers move.
    
    Concise explanation can use visually rich content. Maps, scene illustrations, screenshots, scientific plots, and structured panels can make a proposal easier to understand when they carry its substance. Evaluate their role and reading hierarchy rather than treating flat minimalism, a particular font, or the absence of rounded panels as a quality test.
    
    ### Make occupied space earn its place
    
    Space efficiency is a design requirement for every figure, especially in papers and proposals where page area is scarce. Judge useful information per occupied area at the final document size. Inspect both the figure's internal composition and the space reserved around it in the document, including captions, gutters, and wrapped prose. A figure can be legible and editable yet still waste substantial page space.
    
    Distinguish whitespace that supports hierarchy and separation from broad unused bands or regions with no reading function. Check each panel's actual content bounds; a long footer, wide background, or oversized text box can hide wasted space when inspecting only the overall bounding box. Reflow or shorten the element that sets an unnecessary width, then reduce the canvas or panel dimensions while preserving meaningful content and readable type. Do not add labels, decoration, or repeated claims just to fill a hole, and do not compress necessary gutters to maximize a pixel-fill ratio.
    
    Return reclaimed space to the destination. For a wrapfigure, reduce the placed width along with the canvas when the goal is to give room back to prose. Cropping the source while retaining its old placed width enlarges the content and can increase its height instead of saving page area. Recheck physical font size, caption wrapping, adjacent paragraphs, section transitions, and total pagination. Read [space-efficiency.md](references/space-efficiency.md) when reflowing a sparse figure or adapting one to a compact document slot.
    
    ### Use the lab's default palette
    
    For new paper, proposal, and GitHub README figures without a specified visual identity, use the shared CatchBench, Cat-DPO, and No Attacker Needed color family: white background, mint and coral as the main contrast, teal and warm gold when additional categories are needed, and dark text with quiet gray context. The latter two papers are Tiankai Yang's reference examples. Read [default-palette.md](references/default-palette.md) for color roles, source evidence, and the canonical reusable tokens in `references/default-palette.json`.
    
    An explicit user palette or an established document palette takes priority. Preserve accepted figures' colors when editing them. Carry the same color-to-meaning mapping across the schematic, results, caption keys, and other figures in a document. Reference searches may supply composition ideas without changing this default visual identity; do not repeat a paper search just to recover the stored colors.
    
    ### Reduce reading load before reducing font size
    
    For each proposed label, identify the information it adds or the misunderstanding it prevents. Omit text whose removal leaves the correct reading intact, including sentences that repeat a visible relationship, panel title, or caption. Do not add explanations, formal terms, or caveat blocks merely to make the figure appear rigorous. Keep verification details in the working notes and methodological detail in the surrounding prose unless the reader needs it at that point in the graphic.
    
    Keep a number when it carries the claim, such as a verified performance gap or scale comparison. Do not add counts merely to make the work appear substantial. Do not transfer every deleted sentence into an oversized caption.
    
    Simplification must preserve the inference. Check whether a removed qualifier makes a diagnostic look like general evidence, an illustrative story look like a paired experiment, a forecast look like a result, or a highlighted answer look like an input available to the evaluated method.
    
    Prefer short, concrete action labels; keep a formal term underneath only when it helps connect the drawing to the prose. Judge understanding rather than word count alone. A record edit should be described as a record edit: wording that implies the agent actually behaved differently requires evidence of that behavior.
    
    Require secondary strips and extra panels to explain a distinct relationship or supply useful evidence. A generic sequence such as edit, rebuild, record may add little even when accurate. If its only contribution fits in one caption sentence, remove it and reclaim its space. When that secondary mechanism is itself central, show a concrete before/after example instead. After deleting a footer, reduce canvas height while keeping its width, retained object coordinates, and font sizes fixed. If the width or object scale changes, recalculate printed text size.
    
    The reader should recognize the main contrast at a glance and explain the figure after reading its short caption. Inspect at the intended manuscript or README width. A legible full-screen slide can still fail as a paper figure.
    
    When breadth of resources, research outcomes, or community is itself the argument, retain the supported inventory and give it sufficient page area and reading levels. For a compact destination, select the relevant part rather than shrinking the entire inventory below readable size. A proposal ecosystem overview can expose its organization at a glance while leaving individual publications or examples for a later read. The reduction criteria target detail that does not support the figure's claim; see [proposal-figures.md](references/proposal-figures.md) and the [preferred examples](references/proposal-style-exemplars.md).
    
    For mechanism diagrams, read [mechanism-figures.md](references/mechanism-figures.md): it covers minimal examples, source checks, essential qualifiers, secondary content, and a concrete revision case.
    
    When revising a dense first figure, read [iteration-example.md](references/iteration-example.md). Its layout and counts are an example, not defaults.
    
    ### Combine explanation with evidence when useful
    
    A dominant schematic can explain what the work makes possible, while a smaller result panel gives the reader a numerical reason to care. Repeating a selected result from Results can serve this introductory purpose. Judge its contribution to the reading path rather than rejecting all repetition or optimizing for the fewest words.
    
    When the author wants a large (a) and narrow (b), reserve the gutter first, then divide the remaining width according to their reading roles. Judge separation between the actual labels and artwork, including axes and callouts. A divider does not compensate for crowded panels. After changing the canvas, recalculate text size at the final display width.
    
    Read [panels-and-results.md](references/panels-and-results.md) when combining an explanatory panel with data, separating regions within a wide schematic, or revising panel spacing. It covers panel hierarchy, region separation, cross-panel edges, source-derived numbers, descriptive versus inferential claims, and focused revisions. Use the plotting workflow for scientific correctness and the native PowerPoint guidance for editable chart behavior.
    
    ### Choose how to draft
    
    Native PowerPoint can be the first draft when the figure consists of text, nodes, arrows, and simple geometry. A preliminary raster generation is optional, unless the user explicitly requests it. If an image generator helps explore a richer concept, use the available image-generation workflow, then rebuild the required semantic content as native objects. Do not treat a screenshot on a slide as an editable figure.
    
    Show one recommended direction with a concrete reason. Generate alternatives only when a real design tradeoff remains or the user requests them. Save revisions under meaningful new names so the author can compare them.
    
    ## 4. Build the editable figure
    
    Read [native-powerpoint.md](references/native-powerpoint.md) before authoring or converting a diagram. Discover current tools and fonts rather than copying cache paths or runtime versions from a previous session.
    
    Use text boxes for text, native shapes for semantic objects, and attached connectors for relationships that should follow moved nodes. Keep imported photos or screenshots as such when they are part of the requested figure. Name objects and group them by meaningful region so both whole-block and individual editing remain practical.
    
    Use a figure-sized canvas rather than inheriting a slide aspect ratio without reason. Preserve an existing palette or typography when supplied. Work out the final physical text size after scaling into the document. Do not shrink labels to rescue a composition that needs editing.
    
    Retain the editable source and a reproducible builder when one was used. After manual PowerPoint edits, identify the current source of truth; a stale builder must not silently overwrite an author's changes.
    
    ## 5. Inspect semantics, appearance, and editability
    
    These are separate checks:
    
    1. **Meaning:** Trace claims to the source, distinguish illustrative from empirical content, and verify the figure plus caption conveys the intended usefulness without overstating scope.
    2. **Rendering and space:** Export the final PPTX and inspect the actual result for clipping, wrapping, arrow direction, visibility, spacing, and unused regions. Evaluate each panel and the complete occupied document area, not only a tightly cropped preview. Recheck at the intended document width, and in the actual document when insertion is part of the task.
    3. **Editability:** Verify that important text and objects are native and independently selectable. On a disposable copy, edit representative text and move a connected node to check behavior. For a native result chart, check its data source and edit/restore a series value; verify the embedded workbook when portable data editing is required. Counts of shapes or a PNG preview alone do not prove this.
    
    Use the installed presentation workflow's validators when available. The optional Windows helper `scripts/render_powerpoint.ps1` exports a single-slide figure through local PowerPoint and reports native object counts. It is a renderer and inventory check, not a substitute for visual or editing inspection.
    
    Fix issues that affect the requested result and stop when the checks pass. Do not expand a one-figure revision into unrelated document edits or a large review process. If native PowerPoint becomes unavailable, report which rendering and editing checks remain incomplete and follow the platform requirement above. A fallback preview does not complete this workflow.
    
    Treat review suggestions as claims to reconcile with primary sources and the user's requirements. An optional reviewer does not decide the authoring format or override later user feedback. Record which artifact version a review covers. A small spacing revision calls for geometry, rendering, and relevant data-integrity checks; it need not restart reference research or a full review cycle.
    
    For a visual-only pass, freeze wording, font sizes, semantic groups, and connector endpoints, and compare them with the previous version so polish cannot hide a redesign.
    
    When another model is requested, prepare a frozen image, caption, relevant section, and only the source excerpts needed to verify the mechanism. Include the actual publication width and column layout so advice addresses the real destination. For a first-read comprehension check, initially send only the image and placement context, without revealing the desired takeaway. Obtain that reading before sending the caption and source for the semantic check; putting the answer later in the same prompt still primes the reader. Other review tasks can use the combined packet. Use existing authorization for that material and destination; resolve any missing permission only when required, and do not expand a limited review into sending the full manuscript or repository. Separate the reviewer's direct checks from locally reported PowerPoint and document checks.
    
    ## 6. Deliver for the destination
    
    Save the editable `.pptx` and the appropriate viewing or publication export together, using the project's existing figure/assets folder or the user's named destination. Do not leave the only usable deliverable in temporary storage.
    
    - Papers and proposals: normally a vector `.pdf`, plus a PNG preview when useful.
    - GitHub README: a `.png` or compatible `.svg` for display, plus the `.pptx` source. Provide meaningful alt text when embedding it.
    
    Check that the exports correspond to the final editable source. Supply a concise caption or nearby sentence when needed to make the figure interpretable. Do not insert into or rewrite the document unless authorized by the task. Link the editable file, show a preview, and briefly identify any material limitation.
    
    Keep caption or adjacent explanatory text in one canonical place; for LaTeX insertion, use one canonical float fragment. After a revision, synchronize the PPTX, viewing exports, builder dimensions and group definitions, caption, and current verification notes. Before replacing files after a long build, check whether those targets changed concurrently. Preserve unrelated edits and render or compile the current document when insertion is authorized; a private snapshot is not a replacement for newer document work. Keep historical reviews labeled by version rather than implying they cover later edits.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related