arxiv-doc-builder
Convert an arXiv paper to Markdown for reading or implementation reference. Use when asked to convert, fetch, or create documentation for an arXiv paper by its ID, or when a paper with a known arXiv ID needs to be read or referenced. Fetches the LaTeX source when available (plus
Install
npx skills add https://github.com/ultimatile/arxiv-skills/tree/main/skills/arxiv-doc-builder
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ultimatile-arxiv-skills@llmmart
git clone https://github.com/ultimatile/arxiv-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ultimatile/arxiv-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
arXiv Document Builder
Procedure
Run the converter:
# Using global command (recommended) convert-paper ARXIV_ID [--output-dir DIR] # Using script directly uv run --project "SKILL_DIR" "SKILL_DIR/arxiv_doc_builder/convert_paper.py" ARXIV_ID [--output-dir DIR]SKILL_DIR: replace with the absolute path of the directory this SKILL.md is in. It is a placeholder, not a shell variable. Keep the double quotes around it, so a path containing spaces stays one argument. The command then runs from any working directory.--output-dir: Directory where{SAFE_ID}/{SAFE_ID}.mdwill be created. Default: current working directory (not apapers/subdirectory).{SAFE_ID}, here and below, isARXIV_IDwith/replaced by_, as inhep-th/9711200→hep-th_9711200. A new-style ID contains no/, so it is used unchanged.- Use absolute paths to control output location precisely.
convert-paperdoes the metadata lookup, downloads, extraction, and directory creation itself; do not run curl, tar, or mkdir for them.On success,
convert-paperprints the Markdown file's path on itsOutput:line. The file opens with a YAML frontmatter block of provenance metadata;references/output-format.mddocuments its fields, including whatmetadata_statusrecords. The File Organization section ofreferences/output-format.mdshows the layout of the paper's directory,{SAFE_ID}/.
When Conversion Fails or Falls Back to PDF
If you edited a file under {SAFE_ID}/source/ and re-ran convert-paper, read references/source-edits.md first and follow it before any line below.
When you made no such edit, or once that file no longer tells you to edit again or re-run, read the file that the line matching the output points to. For an output not listed here, act on what the output itself says.
Whichever file you follow, change the source only as it directs. Do NOT attempt broad preprocessing (replacing documentclass, expanding \newcommand, removing environments, etc.) — pandoc handles revtex4/revtex4-2, custom commands, picture environments, and theorem environments correctly.
convert-paperexits with code 2 after printingError: Found N files with \documentclass→references/multiple-documentclass.md- It prints
Pandoc conversion failed:, and the pandoc message after it containsunexpected (orunexpected [→references/unknown-arity-macros.md - It prints
Pandoc conversion failed:, and the pandoc message after it contains neither →references/pandoc-failures.md - It prints
Pandoc did not finish withinorPandoc exceeded the <N> MB memory watchdog, or a pandoc run has not returned →references/pandoc-runaway.md - It prints
No LaTeX source, falling back to naive PDF conversion...→references/pdf-conversion.md
Files (arxiv-skills)
-
arxiv_doc_builder
-
arxiv_id.py 5.4 KB
"""Shared arXiv ID helpers. Canonical forms accepted: - New-style YYMM.NNNNN (5 digits) for YYMM >= 1501 (Jan 2015 onwards) - New-style YYMM.NNNN (4 digits) for 0704 <= YYMM <= 1412 - Legacy archive[.subject-class]/YYMMNNN for 9108 <= YYMM <= 0703 - Any of the above with an optional vN (N >= 1) version suffix Non-canonical new-style inputs (e.g. 2202.1173, which arXiv silently redirects to 2202.01173 — a *different* paper) are rejected. The older "zero-pad to canonical width" behaviour was removed because a 4-digit user input on a post-2015 paper is almost never an intentional request for the zero-padded 5-digit paper, and the silent remap masks typos. The validator also rejects semantically impossible inputs: all-zero sequence numbers, version zero, invalid months, and legacy-scheme YYMM values that fall outside arXiv's actual legacy window. """ import re from typing import Optional # New-style ID with split YY/MM so we can range-check the boundary month. _NEW_STYLE_RE = re.compile(r"^(\d{2})(\d{2})\.(\d{4,5})(v\d+)?$") # Legacy ID, e.g. hep-th/9901001, math.AG/0703001, cond-mat/0601234v2, # physics.optics/0501001, cond-mat.str-el/0601234. Subject classes come # in both short-uppercase (math.AG, nlin.CD) and lowercase-with-hyphens # (physics.optics, cond-mat.str-el, physics.comp-ph) forms. _LEGACY_RE = re.compile( r"^[a-z]+(?:-[a-z]+)?(?:\.[A-Za-z]+(?:-[A-Za-z]+)*)?" r"/(\d{2})(\d{2})(\d{3})(v\d+)?$" ) # Boundaries: # Aug 1991: first arXiv submission (hep-th/9108001) # Mar 2007: last legacy-scheme month # Apr 2007: first new-style month (0704.0001) # Jan 2015: first 5-digit sequence month _LEGACY_MIN_YYYYMM = 199108 _LEGACY_MAX_YYYYMM = 200703 _FIRST_NEW_YYMM = 704 _FIVE_DIGIT_YYMM = 1501 def safe_arxiv_id(arxiv_id: str) -> str: """Make an arXiv ID safe for use as a filesystem path component.""" return arxiv_id.replace("/", "_") def _check_version(arxiv_id: str, version_group: Optional[str]) -> None: """Reject v0. arXiv version numbering starts at v1.""" if not version_group: return if int(version_group[1:]) < 1: raise ValueError( f"arXiv ID {arxiv_id!r}: version numbers start at v1, " f"got {version_group!r}." ) def _check_sequence(arxiv_id: str, seq: str) -> None: """Reject all-zero sequence numbers. arXiv sequences start at 1.""" if int(seq) < 1: raise ValueError( f"arXiv ID {arxiv_id!r}: sequence numbers start at 1, got {seq!r}." ) def _legacy_yyyymm(yy: str, mm: str) -> Optional[int]: """Expand legacy YY/MM to a comparable YYYYMM integer, or None. Legacy 2-digit years map to 1991-1999 (91..99) or 2000-2007 (00..07). YY outside those ranges has no legacy interpretation and returns None, which the caller turns into a validation error. """ yy_int = int(yy) if 91 <= yy_int <= 99: year = 1900 + yy_int elif 0 <= yy_int <= 7: year = 2000 + yy_int else: return None return year * 100 + int(mm) def validate_arxiv_id(arxiv_id: str) -> str: """Return ``arxiv_id`` unchanged if canonical, else raise ``ValueError``. Callers are expected to invoke this at argparse boundaries; internal code paths may then trust that IDs are in canonical form (no further zero-padding required before building request URLs from them). """ legacy_m = _LEGACY_RE.match(arxiv_id) if legacy_m: yy, mm, seq, version = legacy_m.groups() month = int(mm) if not 1 <= month <= 12: raise ValueError(f"Invalid month {mm!r} in arXiv ID {arxiv_id!r}.") ym = _legacy_yyyymm(yy, mm) if ym is None or not (_LEGACY_MIN_YYYYMM <= ym <= _LEGACY_MAX_YYYYMM): raise ValueError( f"arXiv ID {arxiv_id!r}: legacy scheme covers Aug 1991 " "through Mar 2007. Outside that window, use the new-style " "YYMM.NNNN(N) form." ) _check_sequence(arxiv_id, seq) _check_version(arxiv_id, version) return arxiv_id m = _NEW_STYLE_RE.match(arxiv_id) if not m: raise ValueError( f"Unrecognized arXiv ID format: {arxiv_id!r}. " "Expected YYMM.NNNN / YYMM.NNNNN (optionally with vN), " "or legacy archive/YYMMNNN." ) yy, mm, seq, version = m.groups() month = int(mm) if not 1 <= month <= 12: raise ValueError(f"Invalid month {mm!r} in arXiv ID {arxiv_id!r}.") yymm = int(yy + mm) if yymm < _FIRST_NEW_YYMM: raise ValueError( f"arXiv ID {arxiv_id!r}: new-style IDs begin in April 2007 " "(YYMM=0704). For earlier papers use the legacy " "archive/YYMMNNN form." ) seq_len = len(seq) if yymm >= _FIVE_DIGIT_YYMM and seq_len != 5: raise ValueError( f"arXiv ID {arxiv_id!r}: papers from 2015-01 onwards use " "5-digit sequence numbers (YYMM.NNNNN). A 4-digit input on a " "post-2015 paper is refused to avoid silently resolving to a " "zero-padded neighbour (e.g. 2202.1173 → 2202.01173)." ) if yymm < _FIVE_DIGIT_YYMM and seq_len != 4: raise ValueError( f"arXiv ID {arxiv_id!r}: papers from 2007-04 through 2014-12 " "use 4-digit sequence numbers (YYMM.NNNN)." ) _check_sequence(arxiv_id, seq) _check_version(arxiv_id, version) return arxiv_id -
arxiv_metadata.py 33.8 KB
#!/usr/bin/env python3 """Fetch a paper's metadata record and render the document frontmatter. Single source of truth for the YAML frontmatter prepended to a converted paper by both conversion paths (``convert_latex.py`` and the PDF path through ``pdf_converter_lib.py``). The frontmatter is the provenance surface a downstream consumer reads, so it must carry the same key set regardless of which path produced it or whether the network was reachable. The record comes from arXiv's Atom API, or from the DOI arXiv registers at DataCite (``10.48550/arXiv.<id>``) when arXiv does not answer with one. ``ArxivMetadata.source`` says which; ``fetch_paper`` reads it, the document does not carry it. Design constraints: - **Runtime is dependency-free.** The LaTeX path runs as a plain ``python3`` subprocess with no third-party packages installed, so the frontmatter is hand-emitted rather than routed through PyYAML. A round-trip contract test (which *may* use PyYAML, a test-only dependency) pins the emitted text to valid, re-parseable YAML. - **The schema is total.** ``build_frontmatter`` always emits every key, even when the lookup failed. Unknown values render as YAML null (a bare ``key:``), which a parser reads as ``None``, except ``categories``, which renders as an empty list. - **Record-supplied fields are transcribed, not interpreted.** ``published``, ``categories``, ``doi``, ``journal`` and ``abstract`` are present exactly when the answering record supplied a value this module could parse, and empty otherwise; what an empty field says about the paper is the arxiv-lookup skill's question. See ``references/output-format.md``. - **The wait is bounded, not the request.** ``fetch_metadata`` stops waiting after a deadline and reports ``unavailable``; the request itself is not cancelled. """ from __future__ import annotations import argparse import dataclasses import html import html.entities import json import re import threading import time import urllib.error import urllib.parse import urllib.request import xml.etree.ElementTree as ET from dataclasses import dataclass, field from pathlib import Path from typing import Any, Callable, Optional, Union # Asked first: it is the system the downloads come from, so its revision matches. _ARXIV_API_URL = "https://export.arxiv.org/api/query" # arXiv asks API clients to identify themselves; urllib's default agent is # the first a public API throttles. _USER_AGENT = "arxiv-doc-builder (+https://github.com/ultimatile/arxiv-skills)" # The Atom feed's namespaces. arXiv's extension carries primary_category, doi # and journal_ref; everything else is plain Atom. _NS = { "atom": "http://www.w3.org/2005/Atom", "arxiv": "http://arxiv.org/schemas/atom", } # DataCite's endpoint for an arXiv DOI (10.48550/arXiv.<id>; case-insensitive). _DATACITE_URL = "https://api.datacite.org/dois/10.48550/arxiv." # Seconds ``fetch_metadata`` waits for both attempts together, by default. METADATA_DEADLINE_SECONDS = 5.0 # arXiv's share of the deadline; DataCite gets the rest. Over 31 measured # lookups arXiv answered in 0.31 s (median) and its stalls did not recur when # measured again, while DataCite took about 1.2 s, so an even split cuts off # no healthy answer and, at the default deadline, leaves the fallback room # even when arXiv spends its whole share. _ARXIV_DEADLINE_SHARE = 0.5 # The frontmatter's ``metadata_status`` values. METADATA_OK = "ok" METADATA_UNAVAILABLE = "unavailable" METADATA_NOT_REQUESTED = "not_requested" METADATA_STATUSES = (METADATA_OK, METADATA_UNAVAILABLE, METADATA_NOT_REQUESTED) # Which source answered. ``fetch_paper`` keeps a later recorded revision only # when DataCite did, since only DataCite can trail what arXiv serves. METADATA_SOURCE_ARXIV = "arxiv" METADATA_SOURCE_DATACITE = "datacite" METADATA_SOURCES = (METADATA_SOURCE_ARXIV, METADATA_SOURCE_DATACITE) # What a fetch can report; no fetch produces ``not_requested``. _FETCH_STATUSES = (METADATA_OK, METADATA_UNAVAILABLE) # A validated arXiv id ends in ``v<N>`` exactly when it names a revision. _VERSION_SUFFIX = re.compile(r"v(\d+)\Z") # A legacy id's subject class ("math.GT/0309136"). Both sources know the paper # only under the archive alone ("math/0309136"). _SUBJECT_CLASS = re.compile(r"^([a-z]+(?:-[a-z]+)?)\.[A-Za-z]+(?:-[A-Za-z]+)*/") # An entity reference with its terminating semicolon: named, decimal or hex. _ENTITY = re.compile(r"&(?:[A-Za-z][A-Za-z0-9]*|#[0-9]+|#[xX][0-9A-Fa-f]+);") # The revision at the front of a DataCite ``dateInformation`` label. _REVISION_LABEL = re.compile(r"v(\d+)") # Date types that list a revision. A withdrawal is still the latest revision. _REVISION_DATE_TYPES = ("Submitted", "Withdrawn") # A full date; DataCite also carries year-only and year-month values. _CALENDAR_DATE = re.compile(r"\d{4}-\d{2}-\d{2}") # DataCite writes an arXiv subject as "Human Name (code)". _SUBJECT_CODE = re.compile(r"\(([^()]+)\)\s*$") @dataclass class ArxivMetadata: """The fields of a metadata record that the frontmatter transcribes. ``None`` means only that no value was retained. The PDF path also builds one from a PDF's own title and author when no record was read. """ title: Optional[str] = None authors: list[str] = field(default_factory=list) version: Optional[str] = None published: Optional[str] = None primary_category: Optional[str] = None categories: list[str] = field(default_factory=list) doi: Optional[str] = None journal: Optional[str] = None abstract: Optional[str] = None # One of ``METADATA_SOURCES``, or ``None`` when no record backs this. # Last, so the other fields keep their positions. source: Optional[str] = None def __post_init__(self) -> None: if self.source is not None and self.source not in METADATA_SOURCES: raise ValueError(f"source {self.source!r} is not one of {METADATA_SOURCES}") @dataclass(frozen=True) class MetadataFetch: """The outcome of one metadata lookup: ``ok`` with a record, or ``unavailable`` with the cause the user is shown.""" status: str metadata: Optional[ArxivMetadata] = None error: Optional[str] = None @property def failure_cause(self) -> str: """Why no usable record was read. Defined only for a non-``ok`` outcome.""" if self.error is None: raise ValueError("an ok outcome has no failure cause") return self.error def __post_init__(self) -> None: if self.status not in _FETCH_STATUSES: raise ValueError( f"{self.status!r} is not a fetch outcome. " f"Expected one of {_FETCH_STATUSES}" ) if (self.status == METADATA_OK) != (self.metadata is not None): raise ValueError( f"status {self.status!r} disagrees with metadata=" f"{'present' if self.metadata is not None else 'absent'}" ) if self.status == METADATA_OK and self.error is not None: raise ValueError(f"an ok outcome carries no cause, got {self.error!r}") if self.status != METADATA_OK and not (self.error or "").strip(): # A blank cause would print a warning that says nothing. raise ValueError( f"status {self.status!r} needs a cause, got {self.error!r}" ) def _unavailable(cause: str) -> MetadataFetch: return MetadataFetch(METADATA_UNAVAILABLE, error=cause) def _cause(exc: BaseException) -> str: """An exception as the cause text a failed outcome carries.""" return f"{type(exc).__name__}: {exc}" def _request(url: str) -> urllib.request.Request: """``url`` as a request that names this client.""" return urllib.request.Request(url, headers={"User-Agent": _USER_AGENT}) def _is_yaml_printable(codepoint: int) -> bool: """Whether a code point may appear raw in YAML (its ``c-printable``).""" return ( codepoint in (0x09, 0x0A, 0x0D, 0x85) or 0x20 <= codepoint <= 0x7E or 0xA0 <= codepoint <= 0xD7FF or 0xE000 <= codepoint <= 0xFFFD or 0x10000 <= codepoint <= 0x10FFFF ) def _normalize(text: Optional[str]) -> Optional[str]: """Record text with non-YAML characters dropped and whitespace collapsed. Dropping is needed because the abstract is a literal block scalar, which cannot escape, and both sources can deliver control characters. Returns ``None`` for empty or whitespace-only input. """ if text is None: return None printable = "".join(ch for ch in text if _is_yaml_printable(ord(ch))) collapsed = re.sub(r"\s+", " ", printable).strip() return collapsed or None def _prose(text: Optional[str]) -> Optional[str]: """``_normalize`` for DataCite prose, which keeps entities like ``>``. Decodes only references ending in ``;``: numeric ones as HTML does (``€`` reads as ``€``), named ones when the name is known. ``html.unescape`` alone would also turn literal ``¬ation`` into ``¬ation``. """ if text is None: return None return _normalize(_ENTITY.sub(_decode_entity, text)) def _decode_entity(match: re.Match[str]) -> str: reference = match.group() if reference.startswith("&#") or reference[1:] in html.entities.html5: return html.unescape(reference) return reference def _archive_form(arxiv_id: str) -> str: """``arxiv_id`` without a legacy subject class, as both sources spell it. "math.GT/0309136v2" -> "math/0309136v2"; any other id is returned as is. """ return _SUBJECT_CLASS.sub(r"\1/", arxiv_id) def split_version(arxiv_id: str) -> tuple[str, Optional[int]]: """Split a validated id into its bare form and its revision number, if any. "2409.03108v2" -> ("2409.03108", 2) "hep-th/9711200" -> ("hep-th/9711200", None) """ match = _VERSION_SUFFIX.search(arxiv_id) if match is None: return arxiv_id, None return arxiv_id[: match.start()], int(match.group(1)) def _dicts(value: object) -> list[dict[str, Any]]: """The dict entries of a JSON array, or ``[]`` when ``value`` is not one.""" if not isinstance(value, list): return [] return [entry for entry in value if isinstance(entry, dict)] def _text(entry: dict[str, Any], key: str) -> Optional[str]: """``entry[key]`` when it is a string, else ``None``.""" value = entry.get(key) return value if isinstance(value, str) else None def _revisions( attributes: dict[str, Any], *, date_types: tuple[str, ...] ) -> list[tuple[int, Optional[str]]]: """The revisions DataCite lists under ``date_types``, with their raw dates. A label can carry a note after ``v<N>`` (``"v2; None"``). An entry without a date still lists its revision, with ``None`` as the date. """ revisions: list[tuple[int, Optional[str]]] = [] for entry in _dicts(attributes.get("dates")): if entry.get("dateType") not in date_types: continue match = _REVISION_LABEL.match(_text(entry, "dateInformation") or "") if match: revisions.append((int(match.group(1)), _text(entry, "date"))) return revisions def _parse_record(bare_id: str, attributes: dict[str, Any]) -> ArxivMetadata: """Map DataCite's ``data.attributes`` onto the frontmatter fields. DataCite holds one record per paper, so ``version`` is the latest revision it lists and ``published`` the first revision's date. """ title = next( ( text for text in ( _prose(_text(entry, "title")) for entry in _dicts(attributes.get("titles")) ) if text ), None, ) authors: list[str] = [] for creator in _dicts(attributes.get("creators")): given = (_text(creator, "givenName") or "").strip() family = (_text(creator, "familyName") or "").strip() # Not by nameType: some single-name people are tagged Organizational. name = f"{given} {family}" if given and family else _text(creator, "name") normalized = _prose(name) if normalized: authors.append(normalized) latest = max( (n for n, _ in _revisions(attributes, date_types=_REVISION_DATE_TYPES)), default=None, ) submitted = { n: date for n, date in _revisions(attributes, date_types=("Submitted",)) if date } version = f"{bare_id}v{latest}" if latest is not None else None first_date = _CALENDAR_DATE.match(submitted.get(1) or "") published = first_date.group() if first_date else None categories: list[str] = [] for subject in _dicts(attributes.get("subjects")): if subject.get("subjectScheme") != "arXiv": continue match = _SUBJECT_CODE.search(_text(subject, "subject") or "") code = _normalize(match.group(1)) if match else None if code: categories.append(code) dois = [ doi for doi in ( _normalize(_text(related, "relatedIdentifier")) for related in _dicts(attributes.get("relatedIdentifiers")) if related.get("relationType") == "IsVersionOf" and related.get("relatedIdentifierType") == "DOI" ) if doi ] abstract = next( ( text for text in ( _prose(_text(entry, "description")) for entry in _dicts(attributes.get("descriptions")) if entry.get("descriptionType") == "Abstract" ) if text ), None, ) return ArxivMetadata( title=title, authors=authors, version=version, published=published, primary_category=categories[0] if categories else None, categories=categories, doi=" ".join(dois) or None, abstract=abstract, source=METADATA_SOURCE_DATACITE, ) def _atom_text(entry: ET.Element, path: str) -> Optional[str]: """The text of the first matching child element, or ``None``.""" el = entry.find(path, _NS) return el.text if el is not None else None def parse_version_from_id(id_url: Optional[str]) -> Optional[str]: """The versioned arXiv id at the end of an Atom ``<id>`` URL, or ``None``. <id>http://arxiv.org/abs/2409.03108v2</id> -> "2409.03108v2" <id>http://arxiv.org/abs/hep-th/9901001v3</id> -> "hep-th/9901001v3" """ if not id_url: return None try: path = urllib.parse.urlparse(id_url).path except ValueError: # An unclosed IPv6 bracket raises; the tail split below still works. path = "" if path.startswith("/abs/"): return path[len("/abs/") :] or None return id_url.rsplit("/", 1)[-1] or None def _is_error_entry(entry: ET.Element) -> bool: """Whether the entry is arXiv's error report (its ``<id>`` is under ``/api/errors``), which otherwise reads like a record.""" try: path = urllib.parse.urlparse(_atom_text(entry, "atom:id") or "").path except ValueError: return False return path.startswith("/api/errors") def _parse_entry(entry: ET.Element) -> ArxivMetadata: """Map one Atom entry onto the frontmatter fields. ``_normalize``, not ``_prose``: the XML parser already decoded entities. """ primary = entry.find("arxiv:primary_category", _NS) published = _CALENDAR_DATE.match(_atom_text(entry, "atom:published") or "") return ArxivMetadata( title=_normalize(_atom_text(entry, "atom:title")), authors=[ name for name in ( _normalize(n.text) for n in entry.findall("atom:author/atom:name", _NS) ) if name ], version=parse_version_from_id(_atom_text(entry, "atom:id")), published=published.group() if published else None, primary_category=primary.get("term") if primary is not None else None, categories=[ term for term in (c.get("term") for c in entry.findall("atom:category", _NS)) if term ], doi=_normalize(_atom_text(entry, "arxiv:doi")), journal=_normalize(_atom_text(entry, "arxiv:journal_ref")), abstract=_normalize(_atom_text(entry, "atom:summary")), source=METADATA_SOURCE_ARXIV, ) def _identity_cause(version: Optional[str], query: str) -> Optional[str]: """Why an entry whose ``<id>`` reads ``version`` is not ``query``'s record. ``None`` when it is: the id names a revision, the bare ids match, and so does the revision when ``query`` names one. Checked before the entry's other fields are read, since another paper's feed parses just the same. ``query`` has no subject class, as arXiv spells entry ids. """ if version is None: return "arXiv's entry carries no id to identify the paper by" bare_entry, revision = split_version(version) if revision is None: return f"arXiv's entry id names no revision: {version}" bare_query, requested = split_version(query) if bare_entry != bare_query: return f"arXiv answered with a record for {bare_entry}, not {bare_query}" if requested is not None and revision != requested: return f"arXiv answered with {version}, not the requested {query}" return None def _http_cause(source: str, code: int) -> str: """What an HTTP failure from ``source`` says, worded alike for both sources.""" if code == 429: return f"{source} rate-limited the request (HTTP 429)" if 500 <= code <= 599: return f"{source} server error (HTTP {code})" return f"HTTP {code}" def _get( url: str, timeout: float, http_cause: Callable[[urllib.error.HTTPError], str] ) -> Union[bytes, MetadataFetch]: """The body ``url`` answers with, or ``unavailable`` saying why not. ``http_cause`` words an HTTP error, and may read its body. Anything other than an HTTP, URL or OS error (``http.client.IncompleteRead``, say) propagates to ``_bounded``. """ try: with urllib.request.urlopen(_request(url), timeout=timeout) as resp: return resp.read() except urllib.error.HTTPError as exc: # Before URLError, its base class. It holds an open socket, so close it. try: return _unavailable(http_cause(exc)) finally: exc.close() except urllib.error.URLError as exc: return _unavailable(f"URLError: {exc.reason}") except OSError as exc: return _unavailable(_cause(exc)) def _arxiv_error_cause(exc: urllib.error.HTTPError) -> str: """What an arXiv HTTP error says; a rejected id's 400 body explains why.""" try: entry = ET.parse(exc).find(".//atom:entry", _NS) except Exception: entry = None if entry is not None and _is_error_entry(entry): summary = _normalize(_atom_text(entry, "atom:summary")) if summary: return f"arXiv rejected the id: {summary}" return _http_cause("arXiv", exc.code) def _lookup_arxiv(arxiv_id: str, timeout: float) -> MetadataFetch: """One request to arXiv's Atom API, classified.""" query = _archive_form(arxiv_id) url = _ARXIV_API_URL + "?" + urllib.parse.urlencode({"id_list": query}) raw = _get(url, timeout, _arxiv_error_cause) if isinstance(raw, MetadataFetch): return raw try: feed = ET.fromstring(raw) except ET.ParseError as exc: return _unavailable(f"arXiv returned malformed XML: {exc}") entry = feed.find(".//atom:entry", _NS) if entry is None: return _unavailable("arXiv returned no record for this id") if _is_error_entry(entry): return _unavailable( _normalize(_atom_text(entry, "atom:summary")) or "arXiv reported an error for this id" ) mismatch = _identity_cause( parse_version_from_id(_atom_text(entry, "atom:id")), query ) if mismatch is not None: return _unavailable(mismatch) return MetadataFetch(METADATA_OK, metadata=_parse_entry(entry)) def _lookup_datacite(arxiv_id: str, timeout: float) -> MetadataFetch: """One request to DataCite, classified. Unanticipated exceptions propagate to ``_bounded``.""" bare_id, requested = split_version(arxiv_id) doi_id = _archive_form(bare_id) url = _DATACITE_URL + urllib.parse.quote(doi_id, safe="/") def http_cause(exc: urllib.error.HTTPError) -> str: if exc.code == 404: # ``doi_id``, not ``bare_id``: the DOI drops a subject class. return f"DataCite has no record for 10.48550/arXiv.{doi_id} (HTTP 404)" return _http_cause("DataCite", exc.code) raw = _get(url, timeout, http_cause) if isinstance(raw, MetadataFetch): return raw try: document = json.loads(raw.decode("utf-8")) except ValueError as exc: # Covers both JSONDecodeError and UnicodeDecodeError. return _unavailable(f"DataCite returned malformed JSON: {exc}") data = document.get("data") if isinstance(document, dict) else None attributes = data.get("attributes") if isinstance(data, dict) else None if not isinstance(attributes, dict): return _unavailable("DataCite response has no data.attributes") # Another paper's record parses just the same, so check its DOI first. expected = f"10.48550/arxiv.{doi_id}".lower() named = data.get("id") if isinstance(data, dict) else None if not isinstance(named, str) or named.lower() != expected: return _unavailable( f"DataCite answered with a record for {named!r}, not {expected}" ) # Under arXiv's spelling, so both sources' ``version`` compare equal. record = _parse_record(doi_id, attributes) if requested is not None: _, latest = split_version(record.version or "") if latest is not None and requested > latest: return _unavailable( f"{arxiv_id} is later than v{latest}, the latest revision " f"DataCite lists for {bare_id}" ) record = dataclasses.replace(record, version=f"{doi_id}v{requested}") return MetadataFetch(METADATA_OK, metadata=record) def _bounded( call: Callable[[], MetadataFetch], budget: float, what: str ) -> MetadataFetch: """``call``'s outcome, or ``unavailable`` once ``budget`` seconds pass. ``urlopen``'s timeout bounds each socket operation, not the request, so ``call`` runs in a daemon thread this stops waiting for; the request itself is not cancelled. Any failure of ``call`` or its thread comes back as ``unavailable``; an interrupt of the caller still propagates. """ outcome: list[MetadataFetch] = [] def run() -> None: try: outcome.append(call()) except Exception as exc: outcome.append(_unavailable(_cause(exc))) try: worker = threading.Thread( target=run, name=f"arxiv-metadata-{what}", daemon=True ) worker.start() except Exception as exc: return _unavailable(_cause(exc)) worker.join(budget) # Read before the outcome, so a result appended just after the join wins. finished = not worker.is_alive() if outcome: return outcome[0] if not finished: return _unavailable(f"{what} exceeded the {budget:g} s budget") # An exception escaped run(): a BaseException, or one its handler raised. return _unavailable(f"{what} ended without a result") def _lookup(arxiv_id: str, deadline: float) -> MetadataFetch: """arXiv's record, or DataCite's when arXiv does not supply one. arXiv gets ``_ARXIV_DEADLINE_SHARE`` of ``deadline`` and DataCite the rest. When both fail, the cause names what each said. """ started = time.monotonic() share = deadline * _ARXIV_DEADLINE_SHARE primary = _bounded( lambda: _lookup_arxiv(arxiv_id, share), share, "the arXiv lookup" ) if primary.status == METADATA_OK: return primary remaining = deadline - (time.monotonic() - started) if remaining <= 0: return _unavailable( f"arXiv: {primary.failure_cause}; no time left to ask DataCite" ) fallback = _bounded( lambda: _lookup_datacite(arxiv_id, remaining), remaining, "the DataCite lookup" ) if fallback.status == METADATA_OK: return fallback return _unavailable( f"arXiv: {primary.failure_cause}; DataCite: {fallback.failure_cause}" ) def fetch_metadata( arxiv_id: str, *, deadline: float = METADATA_DEADLINE_SECONDS ) -> MetadataFetch: """Look up the metadata record for ``arxiv_id``, waiting at most ``deadline`` s. A failed lookup is returned as ``unavailable`` with its cause, never raised (for an ``Exception``). A ``deadline`` that is not a positive ``int`` or ``float`` up to ``threading.TIMEOUT_MAX`` raises ``TypeError`` or ``ValueError`` before anything starts. ``arxiv_id`` must already be validated. """ # Decimal and Fraction would pass the range check, then break join(). if isinstance(deadline, bool) or not isinstance(deadline, (int, float)): raise TypeError( f"deadline must be an int or float, got {type(deadline).__name__}" ) # Also rejects NaN and infinity; above TIMEOUT_MAX join() overflows. if not 0 < deadline <= threading.TIMEOUT_MAX: raise ValueError( "deadline must be a positive number of seconds no greater than " f"threading.TIMEOUT_MAX ({threading.TIMEOUT_MAX:g}), got {deadline!r}" ) try: return _lookup(arxiv_id, deadline) except Exception as exc: return _unavailable(_cause(exc)) # The handoff carries every record field: lists, and text-or-null scalars. _HANDOFF_LISTS = tuple( f.name for f in dataclasses.fields(ArxivMetadata) if f.default_factory is list ) _HANDOFF_SCALARS = tuple( f.name for f in dataclasses.fields(ArxivMetadata) if f.name not in _HANDOFF_LISTS ) def write_metadata_handoff(path: Path, arxiv_id: str, fetch: MetadataFetch) -> None: """Write one lookup's outcome for a child process to read back. Refuses an ``ok`` outcome whose record names no source, which the reader would reject. """ if fetch.status == METADATA_OK and fetch.metadata.source is None: # type: ignore[union-attr] raise ValueError("an ok outcome names no source") payload = {"arxiv_id": arxiv_id, "fetch": dataclasses.asdict(fetch)} # ASCII-only, so a lone surrogate travels as \u rather than failing. path.write_text(json.dumps(payload), encoding="utf-8") def _handoff_metadata(value: object) -> Optional[ArxivMetadata]: if value is None: return None if not isinstance(value, dict): raise ValueError("metadata is neither null nor an object") expected = set(_HANDOFF_SCALARS) | set(_HANDOFF_LISTS) if set(value) != expected: raise ValueError( f"metadata keys {sorted(value)} differ from {sorted(expected)}" ) for key in _HANDOFF_SCALARS: if value[key] is not None and not isinstance(value[key], str): raise ValueError(f"metadata field {key!r} is not a string or null") for key in _HANDOFF_LISTS: items = value[key] if not isinstance(items, list) or not all(isinstance(i, str) for i in items): raise ValueError(f"metadata field {key!r} is not a list of strings") # The constructor rejects an unknown ``source``. return ArxivMetadata(**value) def _read_handoff(path: Path, arxiv_id: str) -> MetadataFetch: payload = json.loads(path.read_text(encoding="utf-8")) if not isinstance(payload, dict) or set(payload) != {"arxiv_id", "fetch"}: raise ValueError("top level is not an object with exactly arxiv_id and fetch") if payload["arxiv_id"] != arxiv_id: raise ValueError(f"it describes {payload['arxiv_id']!r}, not {arxiv_id!r}") fetch = payload["fetch"] if not isinstance(fetch, dict) or set(fetch) != {"status", "metadata", "error"}: raise ValueError( "fetch is not an object with exactly status, metadata and error" ) if not isinstance(fetch["status"], str): raise ValueError("status is not a string") if fetch["error"] is not None and not isinstance(fetch["error"], str): raise ValueError("error is not a string or null") # MetadataFetch.__post_init__ checks that status, metadata and error agree. outcome = MetadataFetch( fetch["status"], metadata=_handoff_metadata(fetch["metadata"]), error=fetch["error"], ) # Both lookups name a source, and ``fetch_paper`` branches on it. if outcome.status == METADATA_OK and outcome.metadata.source is None: # type: ignore[union-attr] raise ValueError("an ok outcome names no source") return outcome def read_metadata_handoff(path: Path, arxiv_id: str) -> MetadataFetch: """Read back what ``write_metadata_handoff`` wrote, raising no ``Exception``. A handoff that cannot be trusted yields ``unavailable`` with the reason. """ try: return _read_handoff(Path(path), arxiv_id) except Exception as exc: return _unavailable(f"metadata handoff unreadable: {_cause(exc)}") def add_metadata_handoff_option(parser: argparse.ArgumentParser) -> None: """Add the hidden ``--metadata-handoff`` option ``convert_paper`` passes. Not for users, so ``--help`` omits it. ``convert_paper`` passes an absolute path, so argparse never exits 2 on it, the code reserved for an ambiguous main ``.tex``. """ parser.add_argument("--metadata-handoff", type=Path, help=argparse.SUPPRESS) def resolve_metadata( arxiv_id: str, handoff: Optional[Path], lookup: Callable[[str], MetadataFetch], ) -> MetadataFetch: """The outcome handed over in ``handoff``, or a fresh ``lookup`` without one. ``lookup`` is passed in so a test that patches the caller's name catches it. """ if handoff is not None: return read_metadata_handoff(handoff, arxiv_id) return lookup(arxiv_id) def format_unavailable_warning(arxiv_id: str, *, cause: str) -> str: """The warning both conversion paths show when no usable record was read. ``cause`` is the ``MetadataFetch.error``. """ return ( f"WARNING: no usable metadata record for {arxiv_id}: {cause}\n" " The frontmatter fields a record supplies are left null. Those nulls " "mean the value is unknown, since no record was read.\n" " The document's frontmatter records " f"metadata_status: {METADATA_UNAVAILABLE}." ) def _yaml_dq(value: str) -> str: """Render any string as a YAML double-quoted scalar. Escapes backslash, double quote, ``\\n``, ``\\t``, ``\\r`` and every code point ``_is_yaml_printable`` rejects, so the result is valid YAML for any input, including fields that skip ``_normalize``. """ out: list[str] = [] for ch in value: codepoint = ord(ch) if ch == "\\": out.append("\\\\") elif ch == '"': out.append('\\"') elif ch == "\n": out.append("\\n") elif ch == "\t": out.append("\\t") elif ch == "\r": out.append("\\r") elif _is_yaml_printable(codepoint): out.append(ch) elif codepoint <= 0xFF: out.append(f"\\x{codepoint:02x}") else: out.append(f"\\u{codepoint:04x}") return '"' + "".join(out) + '"' def _scalar_line(key: str, value: Optional[str]) -> str: """One ``key: "value"`` line, or a bare ``key:`` (YAML null) when absent.""" if value is None: return f"{key}:" return f"{key}: {_yaml_dq(value)}" def _list_lines(key: str, items: list[str]) -> str: """A YAML block list, or ``key: []`` when empty.""" if not items: return f"{key}: []" body = "\n".join(f" - {_yaml_dq(item)}" for item in items) return f"{key}:\n{body}" def _block_lines(key: str, value: Optional[str]) -> str: """A literal block scalar (``key: |-``), or a bare ``key:`` when absent. ``value`` must already be ``_normalize``-d: a literal block cannot escape. """ if value is None: return f"{key}:" return f"{key}: |-\n {value}" def build_frontmatter( meta: Optional[ArxivMetadata], *, arxiv_id: Optional[str], source_type: str, conversion_date: str, metadata_status: str, fallback_title: Optional[str] = None, ) -> str: """Render the YAML frontmatter block for a converted paper. Every key is always present. ``arxiv_id`` is ``None`` only with ``metadata_status`` ``not_requested``; the status is passed in because the PDF path builds ``meta`` from the PDF itself when the lookup failed. """ if metadata_status not in METADATA_STATUSES: raise ValueError( f"unknown metadata_status {metadata_status!r}. " f"Expected one of {METADATA_STATUSES}" ) if (arxiv_id is None) != (metadata_status == METADATA_NOT_REQUESTED): raise ValueError( f"metadata_status {metadata_status!r} does not match " f"arxiv_id={'absent' if arxiv_id is None else 'present'}. " f"{METADATA_NOT_REQUESTED!r} is the only status a document with no " f"id can carry, and the only one a document with an id cannot" ) m = meta or ArxivMetadata() # Again here: the PDF path builds ``m`` from raw PDF strings. title = _normalize(m.title) or _normalize(fallback_title) author_names = [name for name in (_normalize(a) for a in m.authors) if name] authors = ", ".join(author_names) if author_names else None lines = [ "---", _scalar_line("title", title), _scalar_line("authors", authors), _scalar_line("arxiv_id", arxiv_id), _scalar_line("version", m.version), _scalar_line("published", m.published), _scalar_line("primary_category", m.primary_category), _list_lines("categories", m.categories), _scalar_line("doi", m.doi), _scalar_line("journal", m.journal), _scalar_line("source_type", source_type), _scalar_line("metadata_status", metadata_status), _scalar_line("conversion_date", conversion_date), _block_lines("abstract", _normalize(m.abstract)), "---", "", "", ] return "\n".join(lines) -
convert_latex.py 19.5 KB
#!/usr/bin/env python3 """ Convert LaTeX source to Markdown. Uses pandoc for conversion, with post-processing for better formatting. """ import argparse import os import re import shutil import subprocess import sys import tempfile import time from pathlib import Path from typing import Optional # Importable both as a package member and as a bare script. Narrow to # ModuleNotFoundError + name check so an ImportError raised *inside* # arxiv_id.py isn't silently masked by the script-mode fallback. try: from arxiv_doc_builder.arxiv_id import safe_arxiv_id, validate_arxiv_id from arxiv_doc_builder.arxiv_metadata import ( METADATA_UNAVAILABLE, add_metadata_handoff_option, build_frontmatter, fetch_metadata, format_unavailable_warning, resolve_metadata, ) except ModuleNotFoundError as _exc: if _exc.name != "arxiv_doc_builder": raise from arxiv_id import safe_arxiv_id, validate_arxiv_id from arxiv_metadata import ( METADATA_UNAVAILABLE, add_metadata_handoff_option, build_frontmatter, fetch_metadata, format_unavailable_warning, resolve_metadata, ) class AmbiguousMainTexError(Exception): """Raised when multiple \\documentclass files exist and no explicit --tex-file was given. Selecting the correct entry point cannot be done reliably by heuristics (file size, sort order, or trial conversion) because supplements and fragments can convert successfully just like the main paper. The caller must re-run with --tex-file pointing at the intended file. """ def __init__(self, candidates: list): self.candidates = candidates super().__init__(f"Found {len(candidates)} files with \\documentclass") def find_main_tex(source_dir: Path) -> Optional[Path]: """Find the main .tex file in source directory. Raises AmbiguousMainTexError if multiple files with \\documentclass are present and none of the conventional names (main.tex, paper.tex, ms.tex, article.tex) exists. The caller is expected to translate this into a fail-first error and require --tex-file on the re-run. """ # Conventional main file names take precedence over the ambiguity check: # if the source already uses a known layout, there is nothing ambiguous. candidates = ["main.tex", "paper.tex", "ms.tex", "article.tex"] for candidate in candidates: tex_file = source_dir / candidate if tex_file.exists(): return tex_file # Collect all .tex files with \documentclass doc_files = [] for tex_file in sorted(source_dir.glob("*.tex")): content = tex_file.read_text(encoding="utf-8", errors="ignore") if "\\documentclass" in content: doc_files.append(tex_file) if len(doc_files) == 1: return doc_files[0] if len(doc_files) > 1: raise AmbiguousMainTexError(doc_files) return None def _int_env(name: str, default: int) -> int: """Read an int env override, falling back to ``default`` when the var is unset, non-numeric, or non-positive. Parsing here (rather than inline at the constant) keeps a malformed override — e.g. ``ARXIV_PANDOC_TIMEOUT=3m`` — from raising at import time and crashing every command before it starts. """ raw = os.environ.get(name) if raw is None: return default try: value = int(raw) except ValueError: print( f"Ignoring non-integer {name}={raw!r}; using {default}.", file=sys.stderr, ) return default # Both bounds must be positive: a zero/negative timeout or RSS cap would trip # immediately and kill every conversion, so reject those like a parse error. if value <= 0: print( f"Ignoring non-positive {name}={raw!r}; using {default}.", file=sys.stderr, ) return default return value # A clean LaTeX->Markdown conversion finishes in seconds even for a 100-page, # 4000-line paper. When pandoc instead runs unbounded, it is trapped recursively # expanding a macro it cannot recognize as a no-op — typically a self-referential # redefinition pulled in from a *bundled style .sty* in the source dir, since # pandoc reads local .sty files matching a \usepackage. Two observed shapes: # a CPU spin at flat memory, and a slow leak (~10 MB/s, reaching tens of GB only # after many minutes). Both are caught by the WALL-CLOCK TIMEOUT, which is the # only reliable control here — GHC's `+RTS -M` heap cap and a macOS RLIMIT_AS # were both measured NOT to stop the runaway. The timeout also indirectly bounds # memory for the slow-leak shape (peak ~= timeout * leak-rate). The RSS WATCHDOG # is defense-in-depth for a hypothetical fast-allocating runaway the timeout # alone would not contain in time: it polls real resident memory (immune to the # RTS/rlimit quirks) and kills early. See references/pandoc-runaway.md in the # arxiv-doc-builder skill. Both bounds are env-overridable for odd edge cases. PANDOC_TIMEOUT_SECONDS = _int_env("ARXIV_PANDOC_TIMEOUT", 180) PANDOC_RSS_CAP_MB = _int_env("ARXIV_PANDOC_RSS_CAP_MB", 8192) _RSS_POLL_SECONDS = 1.0 # Grace to reap the child after SIGKILL before we stop waiting on it; bounding # this wait keeps a wedged child (e.g. uninterruptible sleep) from re-hanging # the very call these bounds exist to prevent. _KILL_REAP_GRACE_SECONDS = 10 # Shared remediation tail for both runaway-kill diagnostics (timeout + memory). # They describe the same root cause and fix, so the guidance is written once. _RUNAWAY_REMEDY = ( "Move the style-only .sty out of the source directory (or comment its " "\\usepackage line) and re-run. See references/pandoc-runaway.md in the " "arxiv-doc-builder skill." ) def _process_rss_mb(pid: int) -> Optional[int]: """Resident set size of a pid in MB, or None if it can't be read. Shells out to `ps -o rss=` (KB) rather than taking a psutil dependency; works on the macOS/Linux targets. RSS (not virtual size) is what causes swap thrash, and reading it externally sidesteps the GHC RTS reserving a huge virtual address space that defeats RLIMIT_AS-style caps. `ps` is resolved to an absolute path so a `.`-in-PATH setup can't run a planted binary from the (untrusted) source tree this tool runs subprocesses in. """ ps_bin = shutil.which("ps") # shutil.which can return a relative path when PATH holds a relative entry; # that would re-resolve against the untrusted cwd, so require an absolute one. if ps_bin is None or not os.path.isabs(ps_bin): return None try: out = subprocess.run( [ps_bin, "-o", "rss=", "-p", str(pid)], capture_output=True, text=True, timeout=5, ) except (subprocess.SubprocessError, OSError): return None out = out.stdout.strip() return int(out) // 1024 if out.isdigit() else None def convert_with_pandoc( tex_file: Path, output_md: Path, timeout: int = PANDOC_TIMEOUT_SECONDS, rss_cap_mb: int = PANDOC_RSS_CAP_MB, ) -> bool: """Convert LaTeX to Markdown using pandoc, bounded in wall-clock and memory. Returns False (with a diagnostic on stderr) instead of hanging when pandoc runs away. ``timeout`` is the primary, reliable bound; ``rss_cap_mb`` is a secondary watchdog against fast memory growth. """ print(f"Converting {tex_file.name} to Markdown with pandoc...") # Resolve pandoc to an absolute path. cwd below is the extracted (untrusted) # source tree, so a bare "pandoc" with `.` in PATH could exec a planted # binary; an absolute path skips PATH lookup in the child entirely. pandoc_bin = shutil.which("pandoc") # Require an absolute path: shutil.which can yield a relative one from a # relative PATH entry, which would re-resolve against the untrusted cwd. if pandoc_bin is None or not os.path.isabs(pandoc_bin): print("Error: pandoc not found on PATH as an absolute path.", file=sys.stderr) return False # Use absolute paths tex_file_abs = tex_file.resolve() output_md_abs = output_md.resolve() # Popen (not subprocess.run) so the loop below can poll RSS while pandoc # runs. pandoc writes the document to -o, so stdout stays empty; stderr goes # to a temp file rather than a PIPE so it can fill without us draining it — # a PIPE left unread during the wait loop could deadlock if pandoc emitted # enough to fill the pipe buffer. The temp file is drained once, after exit. with tempfile.TemporaryFile() as errf: proc = subprocess.Popen( [ pandoc_bin, str(tex_file_abs), "-f", "latex", "-t", "markdown", "--wrap=none", "--mathjax", "-o", str(output_md_abs), ], stdout=subprocess.DEVNULL, stderr=errf, cwd=str(tex_file.parent.resolve()), ) start = time.monotonic() killed_for = None # "timeout" | "memory" | None while True: try: proc.wait(timeout=_RSS_POLL_SECONDS) break # finished on its own except subprocess.TimeoutExpired: pass if time.monotonic() - start > timeout: killed_for = "timeout" else: rss = _process_rss_mb(proc.pid) if rss is not None and rss > rss_cap_mb: killed_for = "memory" if killed_for: proc.kill() try: proc.wait(timeout=_KILL_REAP_GRACE_SECONDS) except subprocess.TimeoutExpired: pass # reaping shouldn't outlast SIGKILL; never re-hang here break errf.seek(0) stderr = errf.read().decode(errors="replace") if killed_for == "timeout": # Not "slow" — a runaway. Name the most common root cause and the fix. print( f"Pandoc did not finish within {timeout}s and was killed. A normal " "conversion takes seconds; a runaway almost always means pandoc is " "recursively expanding a self-referential macro from a bundled " "style .sty in the source directory (pandoc reads local .sty files " f"matching a \\usepackage). {_RUNAWAY_REMEDY}", file=sys.stderr, ) return False if killed_for == "memory": print( f"Pandoc exceeded the {rss_cap_mb} MB memory watchdog and was " "killed (same runaway class as the timeout case — usually a " f"self-referential macro from a bundled style .sty). {_RUNAWAY_REMEDY}", file=sys.stderr, ) return False if proc.returncode != 0: print(f"Pandoc conversion failed: {stderr}", file=sys.stderr) return False print(f"✓ Converted to {output_md}") return True def extract_title_from_latex(tex_file: Path) -> Optional[str]: """Extract the paper title from LaTeX source. Reads the selected main ``.tex`` first so the fallback title belongs to the file actually being converted; only if that file carries no ``\\title`` (it may sit in an included preamble) does it scan sibling ``.tex`` files, rather than picking an arbitrary file from the directory. Returns ``None`` when no ``\\title`` is found, so an unknown title stays null in the frontmatter rather than a fabricated placeholder (matching the PDF path). Strips the markup that most often reaches the frontmatter verbatim, meaning commands, braces, and the ``~`` non-breaking space. Nothing else is resolved, and a title built from other constructs can still arrive mangled. """ candidates = [tex_file] candidates += sorted(p for p in tex_file.parent.glob("*.tex") if p != tex_file) for candidate in candidates: content = candidate.read_text(encoding="utf-8", errors="ignore") # Match \title{...} handling nested braces match = re.search(r"\\title\s*\{([^{}]*(?:\{[^{}]*\}[^{}]*)*)\}", content) if match: title = match.group(1) # Clean up LaTeX commands title = re.sub(r"\\[a-zA-Z]+\s*", "", title) # Remove commands # ``~`` is a non-breaking space and survives the other passes, # reaching the title glued between two words. The lookbehind spares # ``\~``, the tilde accent. title = re.sub(r"(?<!\\)~", " ", title) title = re.sub(r"[{}]", "", title) # Remove braces title = re.sub(r"\s+", " ", title).strip() # Normalize whitespace return title or None return None def post_process_markdown( md_file: Path, arxiv_id: str, tex_file: Path, *, metadata_handoff: Optional[Path] = None, ) -> None: """Post-process Markdown for better formatting. ``metadata_handoff`` is a lookup already made for this paper, written by ``convert_paper``. Without one, the record is looked up here. """ from datetime import datetime, timezone content = md_file.read_text(encoding="utf-8") # A single metadata record supplies the whole provenance frontmatter; the # LaTeX \title of the converted file is only a fallback for the title, and # only when the record did not provide one (lookup failed / no record). # Extracting it lazily avoids a needless file read on the common path. # conversion_date is UTC-aware so the provenance stamp is unambiguous # across environments. fetched = resolve_metadata(arxiv_id, metadata_handoff, fetch_metadata) meta = fetched.metadata if fetched.status == METADATA_UNAVAILABLE: print( format_unavailable_warning(arxiv_id, cause=fetched.failure_cause), file=sys.stderr, ) fallback_title = None if meta is None or not meta.title: fallback_title = extract_title_from_latex(tex_file) header = build_frontmatter( meta, arxiv_id=arxiv_id, source_type="latex", conversion_date=datetime.now(timezone.utc).isoformat(), metadata_status=fetched.status, fallback_title=fallback_title, ) # Fix figure paths (convert to relative paths) content = re.sub( r"!\[([^\]]*)\]\(([^)]+)\)", lambda m: f").name})", content, ) # Clean up excessive blank lines content = re.sub(r"\n{3,}", "\n\n", content) # Write back final_content = header + content md_file.write_text(final_content, encoding="utf-8") print("✓ Post-processed Markdown") def copy_figures(source_dir: Path, output_dir: Path): """Copy figure files to output directory. Only top-level figures are copied. Recursing is tempting but unsafe: the markdown post-processor rewrites all image references to ``figures/<basename>``, so two nested assets with the same basename (e.g. ``main/fig1.png`` and ``supp/fig1.png``) would silently collide and substitute the wrong asset for at least one reference. A proper fix requires collision-aware, path-preserving copying plus a smarter rewriter, which is out of scope here. """ figures_dir = output_dir / "figures" figures_dir.mkdir(exist_ok=True) # Common image extensions image_exts = [".png", ".jpg", ".jpeg", ".pdf", ".eps"] copied = 0 for ext in image_exts: for img_file in source_dir.glob(f"*{ext}"): dest = figures_dir / img_file.name dest.write_bytes(img_file.read_bytes()) copied += 1 if copied > 0: print(f"✓ Copied {copied} figure(s) to {figures_dir}") def main(): parser = argparse.ArgumentParser(description="Convert LaTeX to Markdown") parser.add_argument("arxiv_id", help="arXiv ID") parser.add_argument( "--source-dir", type=Path, help="LaTeX source directory (default: SAFE_ID/source, " "SAFE_ID being arxiv_id with '/' replaced by '_')", ) parser.add_argument( "--output", type=Path, help="Output Markdown file (default: SAFE_ID/SAFE_ID.md)", ) parser.add_argument( "--tex-file", type=Path, help="Specify the main .tex file directly (overrides auto-detection)", ) add_metadata_handoff_option(parser) args = parser.parse_args() try: validate_arxiv_id(args.arxiv_id) except ValueError as e: # Exit 1 (generic failure). Exit 2 is reserved below for the # "ambiguous main .tex" signal, which callers distinguish from # generic failures. print(f"Error: {e}", file=sys.stderr) sys.exit(1) # Run safe_arxiv_id through the default paths so legacy IDs like # "hep-th/9901001" don't smuggle a slash into the directory / filename # (which would produce hep-th/9901001/... and mismatch the # fetch-side cache at hep-th_9901001/). safe_id = safe_arxiv_id(args.arxiv_id) if args.source_dir: source_dir = args.source_dir else: source_dir = Path(safe_id) / "source" if args.output: output_md = args.output else: output_md = Path(safe_id) / f"{safe_id}.md" # Check source directory exists if not source_dir.exists(): print(f"Error: Source directory not found: {source_dir}") sys.exit(1) # Find main .tex file if args.tex_file: tex_file = args.tex_file if not tex_file.exists(): print(f"Error: Specified .tex file not found: {tex_file}") sys.exit(1) else: try: tex_file = find_main_tex(source_dir) except AmbiguousMainTexError as e: # Fail-first: do not guess. Require an explicit --tex-file on re-run. # Use absolute paths so the suggestion is portable across cwd and # --output-dir differences between the original run and the re-run. abs_candidates = [c.resolve() for c in e.candidates] print( f"Error: Found {len(abs_candidates)} files with \\documentclass " f"in {source_dir.resolve()}:", file=sys.stderr, ) for c in abs_candidates: print(f" - {c}", file=sys.stderr) print( "\nMain .tex selection is ambiguous. Re-run with --tex-file " "pointing at the correct file, e.g.:", file=sys.stderr, ) print( f" convert-paper <ARXIV_ID> --tex-file {abs_candidates[0]}", file=sys.stderr, ) print( "\nIf you originally passed --output-dir, include the same value " "in the re-run.", file=sys.stderr, ) sys.exit(2) if not tex_file: print(f"Error: No main .tex file found in {source_dir}") sys.exit(1) print(f"Found main file: {tex_file.name}") # Check pandoc is available result = subprocess.run(["which", "pandoc"], capture_output=True) if result.returncode != 0: print("Error: pandoc not found. Install with: brew install pandoc") sys.exit(1) # Convert output_md.parent.mkdir(parents=True, exist_ok=True) if not convert_with_pandoc(tex_file, output_md): sys.exit(1) # Post-process post_process_markdown( output_md, args.arxiv_id, tex_file, metadata_handoff=args.metadata_handoff, ) # Copy figures copy_figures(source_dir, output_md.parent) print() print("=" * 50) print(f"✓ Conversion complete: {output_md}") if __name__ == "__main__": main() -
convert_paper.py 7.1 KB
#!/usr/bin/env python3 """ Main orchestrator for converting arXiv papers to Markdown. Handles fetching and conversion automatically. """ import argparse import subprocess import sys import tempfile from pathlib import Path # Importable both as a package member (entry point) and as a bare script. # Narrow to ModuleNotFoundError + name check so an ImportError raised # *inside* arxiv_id.py isn't silently masked by the script-mode fallback. try: from arxiv_doc_builder.arxiv_id import safe_arxiv_id, validate_arxiv_id from arxiv_doc_builder.arxiv_metadata import fetch_metadata, write_metadata_handoff from arxiv_doc_builder._version import read_version except ModuleNotFoundError as _exc: if _exc.name != "arxiv_doc_builder": raise from arxiv_id import safe_arxiv_id, validate_arxiv_id from arxiv_metadata import fetch_metadata, write_metadata_handoff from _version import read_version def run_script(script_name: str, args: list, use_uv: bool = False) -> int: """Run a Python script with arguments and return its exit code. Returning the raw returncode (instead of a bool) lets callers distinguish semantically meaningful non-zero codes — e.g., exit code 2 from convert_latex.py signals "ambiguous main .tex" and must be propagated instead of collapsed into a generic failure. """ script_path = Path(__file__).parent / script_name if use_uv: cmd = ["uv", "run", "--no-project", str(script_path)] + args else: cmd = [sys.executable, str(script_path)] + args result = subprocess.run(cmd) return result.returncode def main(): parser = argparse.ArgumentParser( description="Convert arXiv paper to Markdown documentation" ) # action="version" prints and exits before any other parsing, so # `convert-paper --version` works without a positional arxiv_id. # read_version() runs at parser-build time, but it is a cheap metadata # / TOML read. argparse %-formats the whole version string, so each "%" # in the resolved version is doubled to come out literally. parser.add_argument( "-V", "--version", action="version", version=f"%(prog)s {read_version().replace('%', '%%')}", ) parser.add_argument("arxiv_id", help="arXiv ID (e.g., 2409.03108)") parser.add_argument( "--output-dir", type=Path, default=Path("."), help="Output directory (default: current directory)", ) parser.add_argument( "--tex-file", type=Path, help="Specify the main .tex file directly " "(required when multiple \\documentclass files are present)", ) args = parser.parse_args() try: validate_arxiv_id(args.arxiv_id) except ValueError as e: # Exit 1 (generic failure). Exit 2 is reserved for the # "ambiguous main .tex" signal propagated from convert_latex.py # below, which wrappers may retry with --tex-file. print(f"Error: {e}", file=sys.stderr) sys.exit(1) normalized_arxiv_id = safe_arxiv_id(args.arxiv_id) paper_dir = args.output_dir / normalized_arxiv_id source_dir = paper_dir / "source" print("=" * 60) print("arXiv Paper to Markdown Converter") print(f"Paper ID: {args.arxiv_id}") print("=" * 60) print() # One lookup serves every step, handed over in a file in a fresh per-run # directory, so no step reads an earlier run's result. print("Looking up the paper's metadata record...") fetched = fetch_metadata(args.arxiv_id) print() with tempfile.TemporaryDirectory(prefix="convert-paper-") as handoff_dir: # Absolute, so the child's argparse cannot read it as an option. handoff = Path(handoff_dir).resolve() / "metadata.json" write_metadata_handoff(handoff, args.arxiv_id, fetched) handoff_args = ["--metadata-handoff", str(handoff)] # Step 1: Fetch materials (reuses files already downloaded; fetch_paper # says when it downloads them again) print("Step 1: Fetching paper materials...") print("-" * 60) rc = run_script( "fetch_paper.py", [args.arxiv_id, "--output-dir", str(args.output_dir)] + handoff_args, ) if rc != 0: print("\n✗ Fetching failed") sys.exit(1) print() # Step 2: Convert to Markdown print("Step 2: Converting to Markdown...") print("-" * 60) # Decide LaTeX vs PDF path. An explicit --tex-file always forces the # LaTeX path: the auto-detection here only globs the top level of # source/, but some arXiv layouts put the real entrypoint in a # subdirectory, and --tex-file is advertised as a direct override. has_top_level_tex = source_dir.exists() and list(source_dir.glob("*.tex")) if args.tex_file or has_top_level_tex: if args.tex_file: print( f"Using explicit --tex-file {args.tex_file}, running LaTeX conversion..." ) else: print("LaTeX source detected, using LaTeX conversion...") latex_args = [ args.arxiv_id, "--source-dir", str(source_dir), "--output", str(paper_dir / f"{normalized_arxiv_id}.md"), ] if args.tex_file: latex_args += ["--tex-file", str(args.tex_file)] rc = run_script("convert_latex.py", latex_args + handoff_args) if rc != 0: # Exit code 2 signals ambiguous main .tex — the child already # printed a detailed stderr message with candidate paths and # a re-run suggestion, so we just propagate the code. if rc == 2: sys.exit(2) print("\n✗ LaTeX conversion failed") sys.exit(1) else: print("No LaTeX source, falling back to naive PDF conversion...") print("⚠ This uses single-column text extraction only.") print(" Output quality varies — inspect the result and consider") print(" using convert_pdf_double_column.py or convert_pdf_extract.py") print(" if the output is garbled.") pdf_file = paper_dir / "pdf" / f"{normalized_arxiv_id}.pdf" if not pdf_file.exists(): print(f"✗ PDF file not found: {pdf_file}") sys.exit(1) rc = run_script( "convert_pdf_simple.py", [ str(pdf_file), "-o", str(paper_dir / f"{normalized_arxiv_id}.md"), "--arxiv-id", args.arxiv_id, ] + handoff_args, use_uv=True, ) if rc != 0: print("\n✗ PDF conversion failed") sys.exit(1) print() print("=" * 60) print("✓ Conversion complete!") print(f"Output: {paper_dir / f'{normalized_arxiv_id}.md'}") print("=" * 60) if __name__ == "__main__": main() -
convert_pdf_double_column.py 1.6 KB
#!/usr/bin/env -S uv run # /// script # dependencies = ["pdfplumber", "pypdf"] # /// """ Double-column PDF to Markdown converter - converts all pages as double-column. This converter processes all pages with double-column layout (common in academic papers). For mixed layouts, use convert_pdf_extract.py with --double-column-pages option. """ import argparse import sys from pathlib import Path # Import shared library from pdf_converter_lib import convert_pdf_to_markdown import pdfplumber def main(): parser = argparse.ArgumentParser( description="Convert PDF to Markdown (all pages, double-column)", epilog="Example: %(prog)s paper.pdf -o output.md", ) parser.add_argument("pdf_path", type=Path, help="Path to PDF file") parser.add_argument( "-o", "--output", type=Path, help="Output Markdown file path (default: same name as PDF with .md extension)", ) args = parser.parse_args() if not args.pdf_path.exists(): print(f"Error: PDF file not found: {args.pdf_path}") sys.exit(1) output_path = args.output or args.pdf_path.with_suffix(".md") # Get all page numbers with pdfplumber.open(args.pdf_path) as pdf: total_pages = len(pdf.pages) all_pages = set(range(1, total_pages + 1)) # Convert all pages as double-column convert_pdf_to_markdown( pdf_path=args.pdf_path, output_path=output_path, pages_to_extract=None, # All pages double_column_pages=all_pages, # All as double-column ) if __name__ == "__main__": main() -
convert_pdf_extract.py 2.4 KB
#!/usr/bin/env -S uv run # /// script # dependencies = ["pdfplumber", "pypdf"] # /// """ Page-wise PDF extractor - extracts specific pages with optional double-column processing. This converter allows extracting specific pages from a PDF and optionally processing some of them as double-column layout. """ import argparse import sys from pathlib import Path # Import shared library from pdf_converter_lib import convert_pdf_to_markdown, parse_page_ranges def main(): parser = argparse.ArgumentParser( description="Extract specific pages from PDF to Markdown", epilog="Example: %(prog)s paper.pdf --pages 1-5,10 --double-column-pages 3-5", ) parser.add_argument("pdf_path", type=Path, help="Path to PDF file") parser.add_argument( "-o", "--output", type=Path, help="Output Markdown file path (default: same name as PDF with .md extension)", ) parser.add_argument( "--pages", type=str, required=True, help='Page numbers to extract (e.g., "1-5,7,9-12")', ) parser.add_argument( "--double-column-pages", type=str, help='Page numbers to process as double-column (e.g., "3-5"). Must be subset of --pages.', ) args = parser.parse_args() if not args.pdf_path.exists(): print(f"Error: PDF file not found: {args.pdf_path}") sys.exit(1) output_path = args.output or args.pdf_path.with_suffix(".md") # Parse page ranges pages_to_extract = parse_page_ranges(args.pages) double_column_pages = ( parse_page_ranges(args.double_column_pages) if args.double_column_pages else set() ) # Validate: double_column_pages must be subset of pages_to_extract if double_column_pages: invalid_pages = double_column_pages - pages_to_extract if invalid_pages: print( f"Error: --double-column-pages contains pages not in --pages: {sorted(invalid_pages)}" ) print(f" Pages to extract: {sorted(pages_to_extract)}") print(f" Double-column pages: {sorted(double_column_pages)}") print(f" Invalid pages: {sorted(invalid_pages)}") sys.exit(1) # Convert specified pages convert_pdf_to_markdown( pdf_path=args.pdf_path, output_path=output_path, pages_to_extract=pages_to_extract, double_column_pages=double_column_pages, ) if __name__ == "__main__": main() -
convert_pdf_simple.py 2.2 KB
#!/usr/bin/env -S uv run # /// script # dependencies = ["pdfplumber", "pypdf"] # /// """ Simple PDF to Markdown converter - converts all pages as single-column. This is the basic converter that processes all pages with single-column layout. For double-column papers, use convert_pdf_double_column.py or convert_pdf_extract.py. """ import argparse import sys from pathlib import Path # Import shared library from arxiv_metadata import add_metadata_handoff_option from pdf_converter_lib import convert_pdf_to_markdown def main(): parser = argparse.ArgumentParser( description="Convert PDF to Markdown (all pages, single-column)", epilog="Example: %(prog)s paper.pdf -o output.md", ) parser.add_argument("pdf_path", type=Path, help="Path to PDF file") parser.add_argument( "-o", "--output", type=Path, help="Output Markdown file path (default: same name as PDF with .md extension)", ) parser.add_argument( "--arxiv-id", help="arXiv ID for authoritative frontmatter metadata (optional)" ) add_metadata_handoff_option(parser) args = parser.parse_args() # A handoff describes one id's lookup, and --arxiv-id is optional here, so # this pairing is reachable from the command line. Refused by name at the # entry the user invoked: the converter refuses it too, but as a ValueError # naming neither the option nor this script. Exit 1, since exit 2 belongs # to the ambiguous-main-tex channel. if args.metadata_handoff is not None and not args.arxiv_id: print("Error: --metadata-handoff needs --arxiv-id", file=sys.stderr) sys.exit(1) if not args.pdf_path.exists(): print(f"Error: PDF file not found: {args.pdf_path}", file=sys.stderr) sys.exit(1) output_path = args.output or args.pdf_path.with_suffix(".md") # Convert all pages as single-column convert_pdf_to_markdown( pdf_path=args.pdf_path, output_path=output_path, pages_to_extract=None, # All pages double_column_pages=None, # Single-column arxiv_id=args.arxiv_id, metadata_handoff=args.metadata_handoff, ) if __name__ == "__main__": main() -
convert_pdf_split_columns.py 2.6 KB
#!/usr/bin/env -S uv run # /// script # dependencies = ["pdf2image", "pypdf", "pillow"] # /// """ Convert PDF to images with column splitting for better detail. For 2-column academic papers, this splits each page into left/right columns to get higher resolution details of small text and formulas. Thin argv shim: the conversion logic lives in pdf_image_lib.convert_pdf_split_columns. """ import argparse import sys from pathlib import Path # Importable both as a package member (e.g. `python -m # arxiv_doc_builder.convert_pdf_split_columns` after installing the `pdf` extra) # and as a bare sibling script under `uv run --no-project` (sys.path[0] is this # script's directory, where the package is not installed). Narrow to a missing # top-level package so an ImportError raised *inside* pdf_image_lib isn't masked. try: from arxiv_doc_builder.pdf_image_lib import convert_pdf_split_columns except ModuleNotFoundError as _exc: if _exc.name != "arxiv_doc_builder": raise from pdf_image_lib import convert_pdf_split_columns def main(): parser = argparse.ArgumentParser( description="Convert PDF to images with column splitting for better detail" ) parser.add_argument("pdf_path", type=Path, help="Path to PDF file") parser.add_argument( "-o", "--output-dir", type=Path, help="Output directory for images (default: PDFNAME/images_split)", ) parser.add_argument( "--dpi", type=int, default=300, help="Image resolution in DPI (default: 300, higher = better detail)", ) parser.add_argument( "--columns", type=int, default=2, help="Number of columns to split each page into (default: 2)", ) args = parser.parse_args() if not args.pdf_path.exists(): print(f"Error: PDF file not found: {args.pdf_path}") sys.exit(1) # Default output directory if args.output_dir: output_dir = args.output_dir else: paper_name = args.pdf_path.stem output_dir = Path(paper_name) / "images_split" # A PDF with no extension, run from its own directory, has its own # name as the stem, and no directory can be made under a file. stem_path = Path(paper_name) if not stem_path.is_dir() and (stem_path.exists() or stem_path.is_symlink()): print( f"Error: the default output directory {output_dir} is under " f"{paper_name}, which is not a directory. Pass -o DIR." ) sys.exit(1) convert_pdf_split_columns(args.pdf_path, output_dir, args.dpi, args.columns) if __name__ == "__main__": main() -
convert_pdf_with_vision.py 2.8 KB
#!/usr/bin/env -S uv run # /// script # dependencies = ["pdf2image", "pypdf", "pillow"] # /// """ Convert PDF to Markdown by converting pages to images and reading them. This script converts PDFs to images, which can then be read by Claude using the Read tool to extract text and mathematical formulas accurately. Thin argv shim: the conversion logic lives in pdf_image_lib.convert_pdf_to_images. """ import argparse import sys from pathlib import Path # Importable both as a package member (e.g. `python -m # arxiv_doc_builder.convert_pdf_with_vision` after installing the `pdf` extra) # and as a bare sibling script under `uv run --no-project` (sys.path[0] is this # script's directory, where the package is not installed). Narrow to a missing # top-level package so an ImportError raised *inside* pdf_image_lib isn't masked. try: from arxiv_doc_builder.pdf_image_lib import convert_pdf_to_images except ModuleNotFoundError as _exc: if _exc.name != "arxiv_doc_builder": raise from pdf_image_lib import convert_pdf_to_images def main(): parser = argparse.ArgumentParser( description="Convert PDF to images for Claude vision-based processing" ) parser.add_argument("pdf_path", type=Path, help="Path to PDF file") parser.add_argument( "-o", "--output-dir", type=Path, help="Output directory for images (default: PDFNAME/images)", ) parser.add_argument( "--dpi", type=int, default=300, help="Image resolution in DPI (default: 300)" ) parser.add_argument( "--no-split", action="store_true", help="Disable column splitting (default: split into 2 columns)", ) parser.add_argument( "--columns", type=int, default=2, help="Number of columns to split (default: 2)" ) args = parser.parse_args() if not args.pdf_path.exists(): print(f"Error: PDF file not found: {args.pdf_path}") sys.exit(1) # Default output directory if args.output_dir: output_dir = args.output_dir else: paper_name = args.pdf_path.stem output_dir = Path(paper_name) / "images" # A PDF with no extension, run from its own directory, has its own # name as the stem, and no directory can be made under a file. stem_path = Path(paper_name) if not stem_path.is_dir() and (stem_path.exists() or stem_path.is_symlink()): print( f"Error: the default output directory {output_dir} is under " f"{paper_name}, which is not a directory. Pass -o DIR." ) sys.exit(1) convert_pdf_to_images( args.pdf_path, output_dir, args.dpi, split_columns=not args.no_split, num_columns=args.columns, ) if __name__ == "__main__": main() -
fetch_paper.py 15.8 KB
#!/usr/bin/env python3 """ Fetch arXiv paper materials (source and/or PDF). Tries to fetch LaTeX source first, falls back to PDF if unavailable. """ import argparse import json import shutil import subprocess import sys from pathlib import Path from typing import Optional # Importable both as a package member (pytest / entry point) and as a # bare script launched via subprocess from convert_paper.py. Narrow to # ModuleNotFoundError + name check so that an ImportError raised *inside* # arxiv_id.py (transitive missing dep, partial load) is not masked by the # fallback — only a genuinely absent top-level package falls through. try: from arxiv_doc_builder.arxiv_id import safe_arxiv_id, validate_arxiv_id from arxiv_doc_builder.arxiv_metadata import ( METADATA_SOURCE_DATACITE, MetadataFetch, add_metadata_handoff_option, fetch_metadata, resolve_metadata, split_version, ) except ModuleNotFoundError as _exc: if _exc.name != "arxiv_doc_builder": raise # Script invocation: script dir is on sys.path[0], so arxiv_id.py is # importable as a top-level module. from arxiv_id import safe_arxiv_id, validate_arxiv_id from arxiv_metadata import ( METADATA_SOURCE_DATACITE, MetadataFetch, add_metadata_handoff_option, fetch_metadata, resolve_metadata, split_version, ) _METADATA_FILE = ".arxiv-fetch.json" def _probe_metadata(arxiv_id: str) -> MetadataFetch: """The lookup the drift check reads, when no handoff supplies one. The whole outcome, since only it tells a failed lookup from a record without a version. ``arxiv_id`` must already be validated. """ return fetch_metadata(arxiv_id) def _latest_version(probe: MetadataFetch) -> Optional[str]: """The version the probe reports, or ``None`` (failed or versionless).""" return probe.metadata.version if probe.metadata else None def _has_cached_source(paper_dir: Path) -> bool: """Whether a cached source tree is on disk, which ``fetch_source`` reuses. A cached PDF does not count: with only a PDF, the source is requested on every run, and a bogus recorded revision kept there would fail every time. So for a paper arXiv serves as a PDF alone, the recorded revision follows whichever source answered, moving back and forth while DataCite trails. """ source = paper_dir / "source" return source.is_dir() and any(source.rglob("*.tex")) def _target_version( paper_dir: Path, latest: Optional[str], cached: Optional[str], *, pinned: bool, source: Optional[str], ) -> Optional[str]: """The revision this run should have on disk and record. Normally ``latest``. ``cached`` wins instead when all of these hold: the id named no revision, DataCite answered (only it can trail arXiv), a source tree is cached, and ``cached`` is a later revision of the same paper. """ if pinned or latest is None: return latest if source != METADATA_SOURCE_DATACITE: # arXiv is authoritative about its own revisions. return latest if not _has_cached_source(paper_dir): return latest if cached is None: return latest cached_bare, cached_revision = split_version(cached) looked_up_bare, looked_up_revision = split_version(latest) if ( cached_revision is not None and looked_up_revision is not None and cached_bare == looked_up_bare and cached_revision > looked_up_revision ): return cached return latest def _read_cached_version(paper_dir: Path) -> Optional[str]: """Read the previously recorded arXiv version, or None.""" meta = paper_dir / _METADATA_FILE if not meta.exists(): return None try: data = json.loads(meta.read_text(encoding="utf-8")) except Exception: return None version = data.get("version") if isinstance(data, dict) else None # A hand edit can put anything here. return version if isinstance(version, str) else None def _record_version(paper_dir: Path, latest: Optional[str], *, fetched: bool) -> bool: """Record ``latest`` if there is one and material was fetched; say whether.""" if latest is None or not fetched: return False _write_cached_version(paper_dir, latest) return True def _format_sidecar_skip_warning(arxiv_id: str, probe: MetadataFetch) -> str: """The warning for a run that fetched material but recorded no version. Call only when ``_record_version`` returned ``False`` after a fetch. """ if _latest_version(probe) is not None: raise ValueError( "the probe reported a version, so the sidecar was not skipped for " "the reason this warning states" ) if probe.error is not None: situation = f"no usable metadata record for {arxiv_id}: {probe.error}" else: situation = f"the metadata record for {arxiv_id} carried no version" return ( f"WARNING: {situation}\n" f" Version drift was not checked, and {_METADATA_FILE} was not updated." ) def _write_cached_version(paper_dir: Path, version: str) -> None: """Record the fetched arXiv version.""" meta = paper_dir / _METADATA_FILE meta.write_text( json.dumps({"version": version}, ensure_ascii=False), encoding="utf-8", ) def _detect_file_type(path: Path) -> str: """Detect downloaded file type using the file command. arXiv source downloads come in several formats: - gzip-compressed tar archive (most common for multi-file submissions) - gzip-compressed single .tex file (common for older papers) - plain text .tex file (rare) The file command on the outer gzip layer cannot distinguish between a tar archive and a single file inside, so for gzip files we decompress and check the inner content. Returns one of: "tar", "gzip_single", "latex", "unknown" """ result = subprocess.run( ["file", "--brief", str(path)], capture_output=True, text=True ) desc = result.stdout.strip().lower() if "tar archive" in desc: return "tar" if "gzip" in desc: # Decompress and check the inner content type via pipe inner = subprocess.run( f'gunzip -c "{path}" | file --brief -', shell=True, capture_output=True, text=True, ) inner_desc = inner.stdout.strip().lower() if "tar archive" in inner_desc: return "tar" # Single file (LaTeX, text, etc.) return "gzip_single" if "latex" in desc or "tex" in desc or "ascii text" in desc: return "latex" return "unknown" def _extract_gzip_single(downloaded: Path, source_dir: Path) -> bool: """Extract a single gzip-compressed file (not a tar archive). The file command output often contains the original filename, e.g.: "gzip compressed data, was \"main.tex\", ..." We use that to name the output file, falling back to main.tex. """ # Try to recover the original filename from gzip metadata result = subprocess.run( ["file", "--brief", str(downloaded)], capture_output=True, text=True ) desc = result.stdout.strip() original_name = "main.tex" if 'was "' in desc: # Extract name between quotes: was "foo.tex" start = desc.index('was "') + 5 end = desc.index('"', start) original_name = desc[start:end] source_dir.mkdir(exist_ok=True) out_path = source_dir / original_name decompress = subprocess.run(["gunzip", "-c", str(downloaded)], capture_output=True) if decompress.returncode != 0: print(f"Failed to decompress: {decompress.stderr.decode()}") return False out_path.write_bytes(decompress.stdout) print(f"✓ Source extracted to {out_path} (single gzip file)") return True def _needs_refresh(cached: Optional[str], latest: Optional[str]) -> bool: """Decide whether cached artifacts should be re-fetched. Returns True when ``latest``, the revision this run targets (see ``_target_version``), differs from ``cached``, the one recorded locally. Returns False (trust cache) when there is no target revision or the two match. """ if latest is None: return False if cached is None: # No metadata — either a pre-metadata cache or first run. # Re-fetch to establish a version record. return True return cached != latest def fetch_source( arxiv_id: str, output_dir: Path, file_id: str, *, refresh: bool = False, ) -> bool: """ Fetch LaTeX source from arXiv. Handles three arXiv source formats: - tar.gz archive (most common) - single gzip-compressed .tex file (common for older papers) - plain text .tex file (rare) Idempotent: if the source directory already contains at least one ``.tex`` file, the network fetch is skipped — unless ``refresh`` is True (version drift detected). Returns: True if source is available (freshly fetched or already present), False if the fetch failed or the source is not available on arXiv. """ # export.arxiv.org is the host arXiv designates for programmatic access # (https://info.arxiv.org/help/bulk_data.html); arxiv.org/robots.txt # additionally disallows /src. source_url = f"https://export.arxiv.org/src/{arxiv_id}" downloaded = output_dir / f"{file_id}-src.tar.gz" source_dir = output_dir / "source" if _has_cached_source(output_dir) and not refresh: print(f"✓ Source already present at {source_dir}, skipping fetch") return True if refresh: # Clear stale source tree so renamed/deleted files don't persist shutil.rmtree(source_dir, ignore_errors=True) print(f"Fetching source from {source_url}...") result = subprocess.run( ["curl", "-f", "-L", "-o", str(downloaded), source_url], capture_output=True ) if result.returncode != 0: print("Source not available (paper may be PDF-only)") downloaded.unlink(missing_ok=True) return False # Detect file type and extract accordingly file_type = _detect_file_type(downloaded) print(f" Detected source format: {file_type}") extract_ok = False if file_type == "tar": source_dir.mkdir(exist_ok=True) result = subprocess.run( ["tar", "-xzf", str(downloaded), "-C", str(source_dir)], capture_output=True ) if result.returncode != 0: print(f"Failed to extract source: {result.stderr.decode()}") else: print(f"✓ Source extracted to {source_dir}") extract_ok = True elif file_type == "gzip_single": extract_ok = _extract_gzip_single(downloaded, source_dir) elif file_type == "latex": # Plain uncompressed .tex file source_dir.mkdir(exist_ok=True) dest = source_dir / "main.tex" dest.write_bytes(downloaded.read_bytes()) print(f"✓ Source saved to {dest} (uncompressed)") extract_ok = True else: print(f"Unknown source format: {file_type}") if not extract_ok: # Remove partial extraction so it cannot masquerade as a cache hit shutil.rmtree(source_dir, ignore_errors=True) downloaded.unlink(missing_ok=True) return False downloaded.unlink() # Clean up downloaded file return True def fetch_pdf( arxiv_id: str, output_dir: Path, file_id: str, *, refresh: bool = False, ) -> bool: """ Fetch PDF from arXiv. Idempotent: if the PDF file already exists and is non-empty, the network fetch is skipped — unless ``refresh`` is True (version drift detected). Returns: True if PDF is available (freshly fetched or already present), False if the fetch failed. """ pdf_url = f"https://arxiv.org/pdf/{arxiv_id}.pdf" pdf_dir = output_dir / "pdf" pdf_file = pdf_dir / f"{file_id}.pdf" if pdf_file.exists() and pdf_file.stat().st_size > 0 and not refresh: print(f"✓ PDF already present at {pdf_file}, skipping fetch") return True print(f"Fetching PDF from {pdf_url}...") pdf_dir.mkdir(exist_ok=True) result = subprocess.run( ["curl", "-f", "-L", "-o", str(pdf_file), pdf_url], capture_output=True ) if result.returncode != 0: print(f"Failed to fetch PDF: {result.stderr.decode()}") pdf_file.unlink(missing_ok=True) return False if pdf_file.stat().st_size == 0: print("Failed to fetch PDF: empty response") pdf_file.unlink(missing_ok=True) return False print(f"✓ PDF saved to {pdf_file}") return True def main(): parser = argparse.ArgumentParser(description="Fetch arXiv paper materials") parser.add_argument("arxiv_id", help="arXiv ID (e.g., 2409.03108)") parser.add_argument( "--output-dir", type=Path, default=Path("."), help="Output directory (default: current directory)", ) add_metadata_handoff_option(parser) args = parser.parse_args() try: validate_arxiv_id(args.arxiv_id) except ValueError as e: # Exit 1 (generic failure) rather than 2: exit 2 is reserved for # "ambiguous main .tex" per convert_paper.py's child-propagation # contract, and wrappers that retry on 2 with --tex-file would # otherwise misinterpret an ID typo as a source-selection problem. print(f"Error: {e}", file=sys.stderr) sys.exit(1) normalized_arxiv_id = safe_arxiv_id(args.arxiv_id) # Create paper directory paper_dir = args.output_dir / normalized_arxiv_id paper_dir.mkdir(parents=True, exist_ok=True) print(f"Fetching materials for arXiv:{args.arxiv_id}") print(f"Output directory: {paper_dir}") print() # Check for version drift before fetching probe = resolve_metadata(args.arxiv_id, args.metadata_handoff, _probe_metadata) cached = _read_cached_version(paper_dir) looked_up = _latest_version(probe) latest = _target_version( paper_dir, looked_up, cached, pinned=split_version(args.arxiv_id)[1] is not None, source=probe.metadata.source if probe.metadata else None, ) if looked_up is not None and latest != looked_up: print( f"Note: the metadata record names {looked_up}, but the cached {latest} " "is kept, since DataCite's record can trail arXiv. A conversion that " f"reads this lookup, as convert-paper's does, names {looked_up} in " "its frontmatter.", file=sys.stderr, ) refresh = _needs_refresh(cached, latest) if refresh: if cached is None: print( f"No version metadata found, re-fetching to establish record (latest={latest})" ) else: print(f"⚠ Version drift detected: cached={cached}, latest={latest}") print() # Download the revision this run records, so the sidecar matches the disk. # With no version, the id is downloaded as given and nothing is recorded. download_id = latest or args.arxiv_id has_source = fetch_source( download_id, paper_dir, normalized_arxiv_id, refresh=refresh, ) has_pdf = fetch_pdf( download_id, paper_dir, normalized_arxiv_id, refresh=refresh, ) # Record version after successful fetch fetched_any = has_source or has_pdf if not _record_version(paper_dir, latest, fetched=fetched_any) and fetched_any: print(_format_sidecar_skip_warning(args.arxiv_id, probe), file=sys.stderr) # Summary print() print("=" * 50) if has_source: print("✓ LaTeX source available") if has_pdf: print("✓ PDF available") if not fetched_any: print("✗ Failed to fetch any materials") sys.exit(1) print(f"\nMaterials saved to: {paper_dir}") if __name__ == "__main__": main() -
pdf_converter_lib.py 12.6 KB
#!/usr/bin/env python3 """ Shared library for PDF to Markdown conversion. This module provides common functions used by the PDF converter scripts. """ import sys from datetime import datetime, timezone from pathlib import Path from typing import Optional, Set import pdfplumber from pypdf import PdfReader # Importable both as a package member (pytest) and as a bare sibling module # under `uv run --no-project` (the PDF scripts add their own dir to sys.path). # Narrow to a missing top-level package so an error inside the module is not # masked. arxiv_metadata is stdlib-only, so it imports under the uv env too. try: from arxiv_doc_builder.arxiv_metadata import ( METADATA_NOT_REQUESTED, METADATA_UNAVAILABLE, ArxivMetadata, build_frontmatter, fetch_metadata, format_unavailable_warning, resolve_metadata, ) except ModuleNotFoundError as _exc: if _exc.name != "arxiv_doc_builder": raise from arxiv_metadata import ( METADATA_NOT_REQUESTED, METADATA_UNAVAILABLE, ArxivMetadata, build_frontmatter, fetch_metadata, format_unavailable_warning, resolve_metadata, ) def parse_page_ranges(range_str: Optional[str]) -> Set[int]: """ Parse page range string like '1-3,5,7-9' into a set of page numbers. Args: range_str: String containing page ranges (e.g., '1-3,5,7-9') Returns: Set of page numbers (1-indexed) """ if not range_str: return set() pages = set() for part in range_str.split(","): part = part.strip() if "-" in part: start, end = part.split("-") pages.update(range(int(start), int(end) + 1)) else: pages.add(int(part)) return pages def clean_text(text: str) -> str: """Clean extracted text by removing excessive whitespace.""" if not text: return "" # Replace multiple spaces with single space text = " ".join(text.split()) # Fix common PDF extraction issues text = text.replace("fi", "fi") text = text.replace("fl", "fl") text = text.replace("ff", "ff") return text def extract_metadata(pdf_path: Path) -> dict: """Extract PDF metadata. Missing title/author are returned as ``None`` rather than synthesized, so the frontmatter renders them as null ("unknown stays unknown") instead of fabricating a filename-derived title. Callers supply their own display fallback (e.g. the file stem) for human-facing output. """ reader = PdfReader(pdf_path) meta = reader.metadata return { "title": meta.title if meta and meta.title else None, "author": meta.author if meta and meta.author else None, "subject": meta.subject if meta and meta.subject else "", "creator": meta.creator if meta and meta.creator else "", } def is_likely_header(text: str) -> bool: """Check if text is likely a page header.""" if not text: return False # Common header patterns header_indicators = [ "PHYSICAL REVIEW", "VOLUME", "NUMBER", ] return any(indicator in text.upper() for indicator in header_indicators) def is_likely_footer(text: str) -> bool: """Check if text is likely a page footer.""" if not text: return False # Just page numbers or very short text if text.strip().isdigit(): return True # Copyright or journal info footer_indicators = [ "The American Physical Society", "©", "Copyright", ] return any(indicator in text for indicator in footer_indicators) def extract_column_text(page_or_crop, page_num: int, column_label: str = "") -> str: """Extract and clean text from a page or cropped region.""" text = page_or_crop.extract_text() if not text: return "" lines = text.split("\n") cleaned_lines = [] for i, line in enumerate(lines): # Skip headers on first few lines if i < 3 and is_likely_header(line): continue # Skip footers on last few lines if i >= len(lines) - 3 and is_likely_footer(line): continue cleaned = clean_text(line) if cleaned: cleaned_lines.append(cleaned) # Join lines with proper spacing content = "\n\n".join(cleaned_lines) return content def extract_page_content(page, page_num: int, is_double_column: bool = False) -> str: """Extract content from a single page, with optional double-column support.""" if is_double_column: # Split page into left and right columns width = page.width height = page.height mid_x = width / 2 # Extract left column left_crop = page.crop((0, 0, mid_x, height)) left_text = extract_column_text(left_crop, page_num, "Left") # Extract right column right_crop = page.crop((mid_x, 0, width, height)) right_text = extract_column_text(right_crop, page_num, "Right") # Combine columns if not left_text and not right_text: return f"\n<!-- Page {page_num}: No text extracted -->\n" content = "" if left_text: content += left_text if right_text: if content: content += "\n\n" content += right_text return content else: # Single column extraction (original behavior) text = page.extract_text() if not text: return f"\n<!-- Page {page_num}: No text extracted -->\n" lines = text.split("\n") cleaned_lines = [] for i, line in enumerate(lines): # Skip headers on first few lines if i < 3 and is_likely_header(line): continue # Skip footers on last few lines if i >= len(lines) - 3 and is_likely_footer(line): continue cleaned = clean_text(line) if cleaned: cleaned_lines.append(cleaned) # Join lines with proper spacing content = "\n\n".join(cleaned_lines) # Extract tables if any tables = page.extract_tables() if tables: table_content = "\n\n" for j, table in enumerate(tables): table_content += f"**Table {j + 1}:**\n\n" # Convert to markdown table if table and len(table) > 0: # Header header = table[0] table_content += ( "| " + " | ".join(str(cell) if cell else "" for cell in header) + " |\n" ) table_content += "|" + "|".join(["---"] * len(header)) + "|\n" # Rows for row in table[1:]: table_content += ( "| " + " | ".join(str(cell) if cell else "" for cell in row) + " |\n" ) table_content += "\n" content += table_content return content def convert_pdf_to_markdown( pdf_path: Path, output_path: Path, pages_to_extract: Optional[Set[int]] = None, double_column_pages: Optional[Set[int]] = None, arxiv_id: Optional[str] = None, *, metadata_handoff: Optional[Path] = None, ) -> None: """ Convert PDF to Markdown using pdfplumber. Args: pdf_path: Path to PDF file output_path: Path to output Markdown file pages_to_extract: Set of page numbers to extract (1-indexed). If None, extract all pages. double_column_pages: Set of page numbers to process as double-column (1-indexed) arxiv_id: arXiv ID for authoritative metadata. When given, the paper's metadata record drives the frontmatter; when omitted (manual PDF scripts) or the lookup fails, the PDF's embedded title/author are used and the record-only fields render as null. ``metadata_status`` records which of those happened. metadata_handoff: A lookup already made for ``arxiv_id``, written by ``convert_paper``. Without one, the record is looked up here. Requires ``arxiv_id``. """ if double_column_pages is None: double_column_pages = set() # One notion of "no id" for the whole function. The status branch below # tests truthiness while the frontmatter builder tests for None, so an # empty string would take the no-id branch here and then be rejected there # for disagreeing with the status it produced, aborting a conversion over # an id that is merely absent. arxiv_id = arxiv_id or None if metadata_handoff is not None and arxiv_id is None: raise ValueError("metadata_handoff describes an arXiv id, so it needs one") print(f"Converting PDF: {pdf_path}") print(f"Output: {output_path}") if pages_to_extract: print(f"Extracting pages: {sorted(pages_to_extract)}") if double_column_pages: print(f"Double-column pages: {sorted(double_column_pages)}") print() # Extract metadata metadata = extract_metadata(pdf_path) print(f"Title: {metadata['title'] or pdf_path.stem}") print(f"Author: {metadata['author'] or 'Unknown'}") print() # Open PDF with pdfplumber with pdfplumber.open(pdf_path) as pdf: total_pages = len(pdf.pages) print(f"Total pages in PDF: {total_pages}") print() # Prepare markdown content markdown_parts = [] # Unified YAML frontmatter (same schema as the LaTeX path). When an # arXiv id is available its record is authoritative; otherwise fall # back to the PDF's embedded title/author, leaving record-only fields # null. if arxiv_id: fetched = resolve_metadata(arxiv_id, metadata_handoff, fetch_metadata) metadata_status = fetched.status meta = fetched.metadata if metadata_status == METADATA_UNAVAILABLE: print( format_unavailable_warning(arxiv_id, cause=fetched.failure_cause), file=sys.stderr, ) else: metadata_status = METADATA_NOT_REQUESTED meta = None if meta is None: pdf_author = metadata["author"] meta = ArxivMetadata( title=metadata["title"], authors=[pdf_author] if pdf_author else [], ) header = build_frontmatter( meta, arxiv_id=arxiv_id, source_type="pdf", # UTC-aware so the provenance stamp is unambiguous across environments. conversion_date=datetime.now(timezone.utc).isoformat(), metadata_status=metadata_status, fallback_title=metadata["title"], ) markdown_parts.append(header) # Process each page for i, page in enumerate(pdf.pages, 1): # Skip if not in extraction set if pages_to_extract and i not in pages_to_extract: continue is_double_column = i in double_column_pages column_info = " (double-column)" if is_double_column else "" print(f"Processing page {i}/{total_pages}{column_info}...") try: page_content = extract_page_content( page, i, is_double_column=is_double_column ) # Add page marker and content markdown_parts.append(f"\n\n<!-- Page {i}{column_info} -->\n\n") markdown_parts.append(page_content) print(f"✓ Page {i} processed ({len(page_content)} chars)") except Exception as e: print(f"✗ Error processing page {i}: {e}") markdown_parts.append(f"\n\n<!-- Page {i}: Error - {e} -->\n\n") print() # Combine all content full_markdown = "".join(markdown_parts) # Add notes section notes = """ --- ## Notes - This document was converted from PDF using pdfplumber - Mathematical formulas are preserved from the PDF text layer - Some formatting may require manual adjustment - Complex equations may need to be formatted as LaTeX - Please review and verify the content accuracy """ full_markdown += notes # Save to file output_path.parent.mkdir(parents=True, exist_ok=True) output_path.write_text(full_markdown, encoding="utf-8") print("=" * 60) print("✓ Conversion complete!") print(f"Output saved to: {output_path}") print(f"Total size: {len(full_markdown)} characters") print("=" * 60) -
pdf_image_lib.py 6.3 KB
#!/usr/bin/env python3 """ Shared library for PDF-to-image conversion. Importable helpers lifted from convert_pdf_with_vision.py and convert_pdf_split_columns.py so the image tier can be unit-tested and typechecked. The scripts themselves stay thin argv shims that run under `uv run --no-project` with their own PEP 723 inline deps (pdf2image / pypdf / pillow). This module imports only those third-party packages plus stdlib, so — unlike pdf_converter_lib.py — it needs no package-or-bare dual import. The two conversion entry points keep distinct contracts on purpose: convert_pdf_to_images is metadata-aware (prints title/author/total, returns the metadata dict alongside the paths) for the vision workflow, while convert_pdf_split_columns is leaner (prints only the page count, returns the paths). Do not unify them. """ from pathlib import Path from pdf2image import convert_from_path from pypdf import PdfReader def extract_metadata(pdf_path: Path) -> dict: """Extract PDF metadata for display. Unlike pdf_converter_lib.extract_metadata (which returns None for missing title/author so the frontmatter renders null), this display-oriented variant falls back to the file stem / "Unknown" because its consumers print the values directly for the human running the vision workflow. """ reader = PdfReader(pdf_path) meta = reader.metadata return { "title": meta.title if meta and meta.title else pdf_path.stem, "author": meta.author if meta and meta.author else "Unknown", "total_pages": len(reader.pages), } def split_image_columns(image, num_columns=2): """Split image into vertical columns. The last column absorbs the integer-division remainder so the columns exactly tile the page width [0, width) with no gap or overlap. """ width, height = image.size column_width = width // num_columns columns = [] for i in range(num_columns): left = i * column_width right = (i + 1) * column_width if i < num_columns - 1 else width column = image.crop((left, 0, right, height)) columns.append(column) return columns def convert_pdf_to_images( pdf_path: Path, output_dir: Path, dpi: int = 300, split_columns: bool = True, num_columns: int = 2, ): """Convert PDF to images, optionally splitting each page into columns. Returns ``(image_paths, metadata)``; prints title/author/total pages plus per-page progress for the vision workflow. """ print(f"Converting PDF to images: {pdf_path}") print(f"Output directory: {output_dir}") print(f"DPI: {dpi}") print(f"Column splitting: {'Yes' if split_columns else 'No'}") if split_columns: print(f"Columns per page: {num_columns}") print() # Extract metadata metadata = extract_metadata(pdf_path) print(f"Title: {metadata['title']}") print(f"Author: {metadata['author']}") print(f"Total pages: {metadata['total_pages']}") print() # Create output directory output_dir.mkdir(parents=True, exist_ok=True) # Convert PDF to images print("Converting pages to images...") images = convert_from_path(pdf_path, dpi=dpi) image_paths = [] for i, image in enumerate(images, 1): # Save full page full_page_path = output_dir / f"page_{i:03d}_full.png" image.save(full_page_path, "PNG") image_paths.append(full_page_path) print(f"✓ Page {i}/{len(images)} saved: {full_page_path.name}") # Split into columns if enabled if split_columns: columns = split_image_columns(image, num_columns) for col_idx, column in enumerate(columns, 1): col_path = output_dir / f"page_{i:03d}_col{col_idx}.png" column.save(col_path, "PNG") image_paths.append(col_path) print(f" └─ Column {col_idx}: {col_path.name}") print() print() print("=" * 60) print("✓ Conversion complete!") print(f"Total images: {len(image_paths)}") if split_columns: print(f" - Full pages: {len(images)}") print(f" - Column images: {len(images) * num_columns}") print(f"Images saved in: {output_dir}") print("=" * 60) print() print("Next steps:") print("1. Use Claude's Read tool to view each image") print("2. Extract text and mathematical formulas in LaTeX format") print("3. Combine into a Markdown document") print() if split_columns: print( "Tip: Process columns separately for better detail on small text/formulas" ) print() return image_paths, metadata def convert_pdf_split_columns( pdf_path: Path, output_dir: Path, dpi: int = 300, num_columns: int = 2 ): """Convert PDF to images with column splitting. Returns ``image_paths``; prints only the page count and per-page progress. """ print(f"Converting PDF with column splitting: {pdf_path}") print(f"Output directory: {output_dir}") print(f"DPI: {dpi}") print(f"Columns per page: {num_columns}") print() # Extract metadata reader = PdfReader(pdf_path) total_pages = len(reader.pages) print(f"Total pages: {total_pages}") print() # Create output directory output_dir.mkdir(parents=True, exist_ok=True) # Convert PDF to images print("Converting pages to images...") images = convert_from_path(pdf_path, dpi=dpi) image_paths = [] for i, image in enumerate(images, 1): # Save full page full_page_path = output_dir / f"page_{i:03d}_full.png" image.save(full_page_path, "PNG") image_paths.append(full_page_path) print(f"✓ Page {i}/{len(images)} full: {full_page_path.name}") # Split into columns columns = split_image_columns(image, num_columns) for col_idx, column in enumerate(columns, 1): col_path = output_dir / f"page_{i:03d}_col{col_idx}.png" column.save(col_path, "PNG") image_paths.append(col_path) print(f" └─ Column {col_idx}: {col_path.name}") print() print("=" * 60) print("✓ Conversion complete!") print(f"Total images: {len(image_paths)}") print(f" - Full pages: {len(images)}") print(f" - Column images: {len(images) * num_columns}") print(f"Images saved in: {output_dir}") print("=" * 60) return image_paths -
_version.py 3.2 KB
"""Resolve the package version for the CLI ``--version`` output. The SSOT is the installed distribution metadata: once built, the version is baked into ``arxiv_doc_builder-<ver>.dist-info/METADATA`` and read at runtime via ``importlib.metadata.version``. That is the only source needed for an installed CLI. A fallback to parsing ``pyproject.toml`` directly *is* warranted here, unlike the usual uv-tool-install workflow. This skill's scripts sit in the checkout, and ``convert_paper.py`` can be run from there by an interpreter the package was never installed into — ``uv run --no-project``, or a bare ``python`` — where no ``.dist-info`` exists and ``importlib.metadata.version`` raises ``PackageNotFoundError``. The fallback is what makes ``--version`` report the real number in that mode instead of crashing. Note the distribution name passed to ``metadata.version`` is the hyphenated ``arxiv-doc-builder`` (``pyproject``'s ``[project] name``), not the underscored import name ``arxiv_doc_builder`` — they intentionally differ. Every failure degrades to ``"unknown"``: ``read_version`` lets no ``Exception`` out and returns only a ``str``. That covers an unexpected metadata error, any exception from locating, reading or parsing the fallback pyproject or looking up its ``[project] version``, and a value that is not a string (a TOML number, or the ``None`` that ``importlib.metadata.version`` returned on Python 3.11, 3.13 and 3.14 for a distribution with no ``Version`` field). The top-level ``import tomllib`` is left unguarded because ``requires-python`` is ``>=3.11``. """ import tomllib from pathlib import Path # Distribution name from pyproject's [project] name. A literal, not derived # from __package__, so a rename surfaces as a lookup miss instead of a wrong # answer (see the dist-name/import-name pitfall in the design notes). _DIST_NAME = "arxiv-doc-builder" _UNKNOWN = "unknown" def read_version() -> str: """Return the package version, or ``"unknown"`` if unresolvable.""" try: from importlib import metadata try: return _str_or_unknown(metadata.version(_DIST_NAME)) except metadata.PackageNotFoundError: # No dist-info — the common source-tree case. Fall through. return _version_from_pyproject() except Exception: # Anything else (corrupt metadata, a failed import) is not the "not # installed" signal, so skip the pyproject fallback. return _UNKNOWN def _version_from_pyproject() -> str: """Read ``[project] version`` from the sibling ``pyproject.toml``. Walks up from this module to the package root's parent, where the project's pyproject lives. Any exception from locating, reading or parsing it, or from looking up the key, and a value that is not a string, collapse to ``"unknown"``. """ try: pyproject = Path(__file__).resolve().parent.parent / "pyproject.toml" with pyproject.open("rb") as f: version = tomllib.load(f)["project"]["version"] except Exception: return _UNKNOWN return _str_or_unknown(version) def _str_or_unknown(value: object) -> str: return value if isinstance(value, str) else _UNKNOWN -
__init__.py 0 B
-
-
references
-
multiple-documentclass.md 1.3 KB
# Troubleshooting: Multiple \documentclass Files Some arXiv papers (e.g., PRL with supplemental material) contain multiple `.tex` files, each with its own `\documentclass`. Automatic selection is unreliable in this case — the canonical example is `1911.04882`, which ships both the main PRL paper and an independent PRL supplement, and either can convert successfully. Since pandoc succeeding is not evidence that the selected file is the correct entry point, `convert-paper` refuses to guess: it fails explicitly with **exit code 2** and lists all candidates. Example failure output: ``` Error: Found 2 files with \documentclass in /path/to/1911.04882/source: - /path/to/1911.04882/source/main_paper.tex - /path/to/1911.04882/source/supplemental_material.tex Main .tex selection is ambiguous. Re-run with --tex-file pointing at the correct file, e.g.: convert-paper <ARXIV_ID> --tex-file /path/to/1911.04882/source/main_paper.tex If you originally passed --output-dir, include the same value in the re-run. ``` To resolve, re-run `convert-paper` with `--tex-file` pointing at the correct main file: ```bash convert-paper 1911.04882 --tex-file /path/to/1911.04882/source/main_paper.tex ``` If the original run used `--output-dir`, pass the same value again so that `convert-paper` reconstructs the correct paper directory. -
output-format.md 9.8 KB
# Output Format Specification This document has two kinds of content: - **Frontmatter (code-generated).** The YAML frontmatter at the top of every converted paper is emitted by `arxiv_doc_builder/arxiv_metadata.py` (`build_frontmatter`), which is the single source of truth for its schema. The block below documents that schema; the code, not this prose, defines it. - **Body formatting (agent-facing).** Everything after the frontmatter — math, figures, tables, code, citations — is guidance for an agent cleaning up or authoring the Markdown by hand. There is no code enforcing it, so it lives here. ## Frontmatter `build_frontmatter` writes one YAML block keyed identically on both conversion paths (LaTeX and PDF). The schema is **total**: every key is always present. A value that is not known renders as YAML null (a bare `key:`), which a parser reads as `None` and not as a missing key. The one exception is `categories`, which renders as an empty list (`categories: []`) instead. The converter transcribes a metadata record. It looks the record up once per run, asking arXiv's own API first and the registration arXiv files at DataCite when arXiv does not answer with a usable one. The wait is bounded, and no outcome of the lookup stops the conversion. For the fields a record supplies — `published`, `categories`, `doi`, `journal`, `abstract` — a value is present exactly when the answering record supplied one the converter could parse and retain, and otherwise the field is empty, as the paragraph above renders it. **The document says nothing further about what a null implies about the paper.** Whether a paper is published, for instance, is a bibliographic question this converter does not answer; the arxiv-lookup skill is where that belongs. The rule does not reach three groups. `title` and `authors` survive on local sources when no record backs the document (see the field notes below). `arxiv_id`, `source_type`, `metadata_status` and `conversion_date` are generated by the converter. `version` and `primary_category` are derived rather than carried. `metadata_status` records how the lookup ended: - `ok`. A record was read. - `unavailable`. A record was sought and none the converter could use was read. The conversion names the cause on stderr as it runs: what each source it asked said, or why a lookup handed over from `convert-paper` could not be read. - `not_requested`. The conversion ran with no arXiv id, so no record was sought. Among the manual PDF conversion scripts, `convert_pdf_simple.py` is the one that takes an `--arxiv-id`, and documents from the rest always carry this token. ```yaml --- title: "Paper Title" authors: "Author A, Author B, Author C" arxiv_id: "2409.03108" version: "2409.03108v2" published: "2024-09-04" primary_category: "cs.AI" categories: - "cs.AI" - "cs.CL" doi: "10.1145/1234567.1234568" # or bare `doi:` (null) journal: "Phys. Rev. D 76, 013009 (2007)" # or bare `journal:` (null) source_type: "latex" # or "pdf" metadata_status: "ok" # or "unavailable" / "not_requested" conversion_date: "2025-12-08T10:00:00+00:00" abstract: |- Single-paragraph abstract, whitespace-normalized. --- ``` Field notes: - `version` is the full versioned arXiv id (e.g. `2409.03108v2`, legacy `hep-th/9901001v3`). For an id given without a revision it names the latest revision the answering record lists; for an id given with one it names that revision. A withdrawn revision is still the paper's latest: `version` names it, the abstract reads as the withdrawal notice, and `published` stays the first revision's date. Which revision the other fields describe depends on the source that answered — arXiv is asked for the revision requested and returns its entry, while DataCite holds one record per paper and its fields follow that paper's latest revision. `version` can also name a revision other than the one converted: when an id without a revision is answered by DataCite, whose record trails a later revision whose source is already cached, the fetch step keeps that revision and prints a note on stderr naming both, while `version` names the record's. - `published` is the paper's date (`YYYY-MM-DD`); `conversion_date` is when the conversion ran (UTC-aware ISO 8601). They are deliberately distinct. - `doi` holds the published DOIs the answering record carries, spelled as that record stores them, except that unprintable characters are dropped and whitespace is collapsed and trimmed. It can name several, separated by spaces. Resolving a DOI the answering record does not carry (e.g. via OpenAlex) is the arxiv-lookup skill's job, not this converter's. - `journal` is the journal reference the answering record carries, as the author entered it on arXiv. Only arXiv's record has such a field. - `categories` lists the arXiv category codes the answering record gives, in its order. Which code of an alias pair (`math-ph` and `math.MP`, say) appears is likewise whatever that record lists. - `primary_category` is the category the answering record marks as primary. Not every record marks one, and where none is marked the value is either the first code in `categories` or null — which of the two depends on the source, since one marks a primary category explicitly and the other does not. A null therefore does not imply that `categories` is empty. - Two fields survive on local sources when no record backs the document. `title` comes from the LaTeX `\title` or the PDF's embedded title on either path, and `authors` from the PDF's embedded author on the PDF path, staying null on the LaTeX path. Every other record-supplied field is empty, rendered as the top of this section describes. A populated `title` or `authors` is therefore no evidence that a record was read, and `metadata_status` is what answers that. ## Body Structure After the frontmatter, the converted body typically looks like: ```markdown # Paper Title ## Abstract Abstract text (when the source carries it inline). ## Table of Contents - [1. Introduction](#1-introduction) - [2. Related Work](#2-related-work) - ... --- ## 1. Introduction Content... ### 1.1 Subsection Content... --- ## References [1] Author et al. Title. Conference/Journal, Year. ``` The abstract is also captured in the frontmatter (`abstract:`), so a consumer can read it from a fixed location even when the body extraction drops it — a common case in the PDF-only fallback. ## Mathematics Formatting ### Inline Math Use single `$` delimiters: ```markdown The learning rate $\alpha$ controls convergence. ``` ### Display Math Use double `$$` delimiters: ```markdown $$ \mathcal{L}(\theta) = \sum_{i=1}^n \ell(y_i, f_\theta(x_i)) $$ ``` ### Numbered Equations ```markdown $$ E = mc^2 \tag{1} $$ ``` ## Figure Handling ### With Available Images ```markdown  *Figure 1: Overview of the proposed architecture* ``` ### Images Not Extracted ```markdown **Figure 1:** Overview of the proposed architecture *[Image not extracted - see PDF page X]* ``` ## Table Formatting Standard Markdown tables: ```markdown | Method | Accuracy | F1 Score | |--------|----------|----------| | BERT | 92.3 | 89.1 | | GPT-2 | 91.8 | 88.5 | *Table 1: Performance comparison on benchmark dataset* ``` ## Code Blocks For algorithms or code: ````markdown ```python def train_model(data, epochs): for epoch in range(epochs): loss = compute_loss(data) update_params(loss) ``` ```` ## Citations ### In-text Citations Prefer readable format: ```markdown This approach was introduced by Smith et al. [1]. ``` Or keep LaTeX format if context is needed: ```markdown The method \cite{smith2023} shows promising results. ``` ### References Section ```markdown ## References [1] Smith, J., Doe, A. (2023). Title of Paper. *Conference Name*, pp. 123-456. [2] Jones, B. (2022). Another Paper. *Journal Name*, 15(3), 789-801. ``` ## File Organization `convert-paper` keeps a paper's files in one directory. In the tree below, `{output-dir}` is the `--output-dir` value and `{SAFE_ID}` is the name of the paper's directory, both as SKILL.md's Procedure step 1 describes them. ``` {output-dir}/ └── {SAFE_ID}/ ├── {SAFE_ID}.md # Main document (frontmatter + paper content) ├── .arxiv-fetch.json # Fetch-side version record (drift detection) ├── figures/ │ └── ... # A LaTeX conversion copies the .png, .jpg, .jpeg, .pdf and .eps files at the top level of source/ here ├── images/ # Page images; convert-paper does not write them, see below │ └── ... ├── source/ # The paper's source as unpacked from arXiv; a PDF-only paper has none │ └── ... └── pdf/ └── {SAFE_ID}.pdf # Original PDF, when its download succeeded ``` The LaTeX conversion creates `figures/` even when it finds no file to copy there, and the PDF fallback does not create it. No conversion deletes a file from `figures/`, so a file an earlier run copied stays until a later run overwrites it. `convert-paper` does not create `images/`. It holds the page images `convert_pdf_with_vision.py` writes when that script's output directory is this one, and `references/pdf-conversion.md` says when it is. The provenance metadata lives in the document's YAML frontmatter (see above). `.arxiv-fetch.json` is an internal sidecar used only for version-drift detection (`{"version": "2409.03108v2"}`); it is not the metadata surface a consumer reads. The sidecar records a version only when the fetch obtained material and the metadata record supplied one. A run that obtained material without a version writes nothing to it, leaving any earlier value in place, and says so on stderr. A run that obtained no material at all exits non-zero. -
pandoc-failures.md 1.1 KB
# Troubleshooting: pandoc Conversion Failures When pandoc fails on a LaTeX source, the error may point to `\end{document}` with `unexpected \end`. This means pandoc's parser broke down due to a syntax issue elsewhere — `\end{document}` itself is not the cause. Keep to SKILL.md's rule against broad preprocessing. ## Diagnosis steps 1. **Binary search for the failing line.** Extract the body (`\begin{document}` to `\end{document}`), then test pandoc with increasing prefixes to find the first line that causes failure. 2. **Check that line for brace mismatches.** The most common cause is an unbalanced `{` or `}` in the LaTeX source. LaTeX's TeX engine silently tolerates these, but pandoc's structured parser does not. 3. **Fix only the mismatch and re-run `convert-paper`.** A single-character fix (e.g., removing an orphaned `{`) is usually sufficient. Then check that the fix survived, as `references/source-edits.md` describes. ## Example The source `(see, e.g., {\cite{makhlin})` has an unmatched `{`. LaTeX compiles fine but pandoc fails. Fix: remove the stray `{`. -
pandoc-runaway.md 4.4 KB
# Troubleshooting: Conversion Hangs / Runaway Memory (pandoc never returns) A brace-mismatch failure is *fast* — pandoc errors in seconds. A different failure mode is the **hang**: `convert-paper` never returns. `convert_latex.py` bounds pandoc on two axes so this surfaces as a fast error instead of an indefinite hang (both env-overridable): - **Wall-clock timeout** (`PANDOC_TIMEOUT_SECONDS`, default 180s; `ARXIV_PANDOC_TIMEOUT`). This is the *reliable* control — every observed runaway is killed by it. - **RSS watchdog** (`PANDOC_RSS_CAP_MB`, default 8192; `ARXIV_PANDOC_RSS_CAP_MB`). Polls the child's real resident memory and kills early; defense-in-depth for a fast-allocating runaway the timeout alone wouldn't contain in time. Why both, and not the obvious one-liners: observed runaways come in two shapes — a CPU spin at flat memory, and a slow leak (~10 MB/s, reaching tens of GB only after *many minutes*). A timeout catches both, and for the slow leak it also bounds peak memory (≈ timeout × leak-rate, so ~1.8 GB at 180s). A memory cap alone would miss the CPU-spin shape. The naive memory caps were **measured not to work** on this failure: GHC's `pandoc +RTS -M2g -RTS` heap limit did not stop the runaway (major-GC checks don't fire fast enough), and a macOS `RLIMIT_AS` cap (4/8/16 GB) didn't kill it either (enforcement is unreliable and collides with the GHC RTS reserving a huge virtual address space). Hence the watchdog polls **RSS externally** (`ps -o rss=`), which is what actually correlates with swap thrash. You may still hit a hang when driving pandoc manually without these bounds. ## Is it slow or hung? Find the pandoc PID (`ps aux | rg pandoc`) and read its state in one shot: ```bash ps -o pid,etime,time,%cpu,rss,state -p <PID> ``` - the **state** column shows `R` (running) and CPU `time` tracks `etime` → on-CPU (slow or runaway), not deadlocked. - **`rss` climbing into the GB/tens-of-GB** → a parser blowup that will not finish. (A 100-page paper converts in seconds and well under ~1 GB.) - The output `.md` size is **not** a progress signal: pandoc buffers the whole document and writes it only at the end (0 bytes until done). ## Root cause: pandoc reads bundled style `.sty` files pandoc's only channel from a `.sty` is the **macro table** it extracts (there is no per-package special-casing for names like `arxiv`). arXiv source tarballs commonly *bundle* a style file (`arxiv.sty`, conference styles, classicthesis-derived headers) right in the source directory, and pandoc reads any local `.sty` whose name matches a `\usepackage`. The blowup is triggered by a **self-referential macro redefinition that is then invoked**, e.g. the "reduced leading" idiom: ```latex \renewcommand{\normalsize}{\@setfontsize\normalsize\@xpt\@xipt ...} \normalsize % invoking it ``` TeX is fine (`\@setfontsize` consumes `\normalsize` as a non-expanded argument); pandoc does not know `\@setfontsize`, so on the invocation it re-expands `\normalsize` inside its own body without bound. Verified minimal repro: self-reference **+ invocation** blows up; the same definition **without** invocation, or a non-self-referential body, converts instantly. ## Fix: strip the style-only `.sty` (safe, and provably output-neutral here) Move the style `.sty` out of the source directory (reversible) or comment its `\usepackage`, then re-run. Then check that the change survived, as `references/source-edits.md` describes: ```bash mv source/arxiv.sty source/arxiv.sty.bak # pandoc no longer reads it ``` Removing a `.sty` is **not** a blanket no-op, but the impact is decidable: it changes output only on `(commands the .sty defines/redefines) ∩ (commands used in the body)`. For a style-only package that intersection is layout scaffolding — `\section`/`\subsection`/`\maketitle` (which pandoc renders *better* from its built-ins; the `.sty`'s `\@startsection` redefinition actually mangles headings) plus front-matter like `\keywords`. Prose, math, citations, and glossary terms are untouched. Before stripping, confirm the `.sty` defines no **content macro** used in the body (e.g. `\newcommand{\co}{ACME}`); if it does, stripping would lose that text, so strip the `.sty` anyway and also copy that macro's definition from the `.sty` into a `\providecommand` placed before the macro's first use anywhere in the source. For style-only packages the intersection contains no content macro, so stripping is output-equivalent on the substantive content. -
pdf-conversion.md 3.5 KB
# PDF Conversion ## PDF Conversion Scripts `convert-paper` only calls `convert_pdf_simple.py` as a naive fallback. The other scripts below are for manual or agent-driven use when the naive output is insufficient. Iterate by trying different scripts and inspecting results. Replace `SKILL_DIR` in the commands below as SKILL.md's Procedure step 1 says. Relative paths you pass as arguments, such as `paper.pdf` and `output.md`, resolve against your own working directory. ### convert_pdf_simple.py Convert all pages as single-column layout. ```bash uv run "SKILL_DIR/arxiv_doc_builder/convert_pdf_simple.py" paper.pdf -o output.md ``` ### convert_pdf_double_column.py Convert all pages as double-column layout (for academic papers). ```bash uv run "SKILL_DIR/arxiv_doc_builder/convert_pdf_double_column.py" paper.pdf -o output.md ``` ### convert_pdf_extract.py Extract specific pages with optional double-column processing. ```bash # Extract specific pages uv run "SKILL_DIR/arxiv_doc_builder/convert_pdf_extract.py" paper.pdf --pages 1-5,10 -o output.md # Extract with mixed column layouts uv run "SKILL_DIR/arxiv_doc_builder/convert_pdf_extract.py" paper.pdf --pages 1-10 --double-column-pages 3-7 -o output.md ``` **Note:** `--double-column-pages` must be a subset of `--pages`. Invalid page ranges cause immediate error. ## Advanced: Vision-Based PDF Conversion For papers with complex mathematical formulas where text extraction fails, a vision-based approach is available as a manual fallback: ```bash # Generate high-resolution images from PDF uv run "SKILL_DIR/arxiv_doc_builder/convert_pdf_with_vision.py" paper.pdf --dpi 300 --columns 2 ``` This creates page images (with optional column splitting) that can be read manually with Claude's vision capabilities for maximum accuracy. This is NOT part of the automatic workflow—use it only when automatic conversion produces poor results. Without `-o`, the script writes the images to `<stem>/images/` under your working directory, where `<stem>` is the PDF's file name with its last extension removed. The PDF's own directory plays no part in that path. When `<stem>` already exists in your working directory and is not a directory, as it does for a PDF with no extension that you run from its own directory, the script writes nothing and prints an error that tells you to pass `-o DIR`. The PDF that `convert-paper` saved is named `{SAFE_ID}.pdf`, with `{SAFE_ID}` as SKILL.md's Procedure step 1 defines it, so for that PDF the default is the `images/` directory of the paper's directory exactly when your working directory is the one `convert-paper` wrote the paper's directory into. To put the images in any other directory, add `-o DIR` to the command: the images then go into `DIR` itself. The files are named like `page_001_full.png` for a whole page and `page_001_col1.png` for one column of it. ### PDF Conversion Quality PDF conversion is inherently lossy: - Math formulas are not in LaTeX format - Complex layouts (2-column with column-spanning elements) may break reading order - Tables may need manual fixing - References may be malformed PDF conversion is acceptable when no LaTeX source is available and the paper is primarily text. For math-heavy papers, use the vision-based approach above or keep the PDF as the primary reference. **Fallback strategy for complex papers:** 1. Extract structure and text via `convert_pdf_simple.py` 2. Keep PDF link for reference 3. Use vision-based conversion for pages with dense math 4. Focus on readable prose sections -
source-edits.md 1.1 KB
# Hand Edits to the Source The fetch step reuses the cached source, but when the metadata lookup reports a revision that `.arxiv-fetch.json` does not record (a different one, or any at all when the file records none), it usually downloads again, deleting `source/` and every edit made to it. The fetch step's output tells you what happened to the edit: - **It printed `✓ Source already present ...`.** The source was reused and the edit is intact. - **It printed `Fetching source from ...` and its summary lists `✓ LaTeX source available`.** The source was downloaded again and the edit is gone. Apply the edit again and re-run. - **It printed `Fetching source from ...` and its summary does not list `✓ LaTeX source available`.** The source was deleted and not replaced. Re-run once with the same arguments. If the summary then lists `✓ LaTeX source available`, apply the edit again and re-run. If it still does not, stop re-running: the source cannot be downloaded now. When the summary lists `✓ PDF available`, the paper can still be converted from the PDF, as `references/pdf-conversion.md` describes. -
unknown-arity-macros.md 1.9 KB
# Troubleshooting: glossaries / cleveref and other unknown-arity macros Symptom: a fast `unexpected (` / `unexpected [` error, often reported at `\begin{document}` (the real cause is elsewhere — pandoc parsed to a boundary). Cause: a heavily-used package whose commands take **optional arguments** pandoc doesn't know the arity of — most commonly `glossaries` / `glossaries-extra` (`\gls`, `\glspl`, `\glsxtrlong`, …, including the `\gls[prereset]{key}` optional-arg form) and `cleveref` (`\cref`). pandoc mis-counts the braces it should consume and breaks once enough body follows. Fix: inject **arity-correct `\providecommand` stubs** (optional-argument-tolerant) just before `\begin{document}`. `\providecommand` only defines them because pandoc never loaded the real package: ```latex \makeatletter \providecommand{\gls}[2][]{#2}\providecommand{\glspl}[2][]{#2} \providecommand{\Gls}[2][]{#2}\providecommand{\Glspl}[2][]{#2} \providecommand{\glsxtrlong}[2][]{#2}\providecommand{\glsxtrlongpl}[2][]{#2} \providecommand{\glsxtrshort}[2][]{#2}\providecommand{\glsxtrshortpl}[2][]{#2} \providecommand{\glsentryshort}[1]{#1}\providecommand{\glsentrylong}[1]{#1} \providecommand{\glslink}[3][]{#3}\providecommand{\glsadd}[2][]{} \providecommand{\cref}[1]{#1}\providecommand{\Cref}[1]{#1} \makeatother ``` Then re-run `convert-paper` and check that the stubs survived, as `references/source-edits.md` describes. Quality note: this expands `\gls{AF}` to its **key** (`AF`), not the glossary long form ("activation function") — pandoc cannot resolve the glossary database. Keys are usually readable (`ReLU`, `NN`, `tanh`), which is acceptable for an implementation-reference doc; `\cref{sec:x}` likewise renders as the label `sec:x`, not a link. This is a *targeted* exception to SKILL.md's rule against broad preprocessing: stubbing a fixed set of known unknown-arity commands, not rewriting the document.
-
-
tests
-
fixtures
-
datacite
-
0705.1442.json 2 KB
{ "data": { "id": "10.48550/arxiv.0705.1442", "type": "dois", "attributes": { "creators": [ { "name": "Gharibyan, Karlen Garnik", "nameType": "Personal", "givenName": "Karlen Garnik", "familyName": "Gharibyan", "affiliation": [], "nameIdentifiers": [] } ], "dates": [ { "date": "2007-05-10T11:41:26Z", "dateType": "Submitted", "dateInformation": "v1" }, { "date": "2009-12-01T09:10:43Z", "dateType": "Updated", "dateInformation": "v1" }, { "date": "2007-06-30T03:41:18Z", "dateType": "Withdrawn", "dateInformation": "v2; None" }, { "date": "2013-05-07T00:03:54Z", "dateType": "Updated", "dateInformation": "v2" }, { "date": "2007-05", "dateType": "Available", "dateInformation": "v1" }, { "date": "2007", "dateType": "Issued" } ], "descriptions": [ { "description": "This paper has been withdrawn Abstract: This paper has been withdrawn by the author due to the publication.", "descriptionType": "Abstract" } ], "doi": "10.48550/arxiv.0705.1442", "relatedIdentifiers": [], "subjects": [ { "lang": "en", "subject": "Computational Complexity (cs.CC)", "subjectScheme": "arXiv" }, { "subject": "FOS: Computer and information sciences", "subjectScheme": "Fields of Science and Technology (FOS)" }, { "subject": "FOS: Computer and information sciences", "schemeUri": "http://www.oecd.org/science/inno/38235147.pdf", "subjectScheme": "Fields of Science and Technology (FOS)" } ], "titles": [ { "title": "Does P=NP?" } ], "version": "2" } } } -
1207.7214.json 3.2 KB
{ "data": { "id": "10.48550/arxiv.1207.7214", "type": "dois", "attributes": { "doi": "10.48550/arxiv.1207.7214", "titles": [ { "title": "Observation of a new particle in the search for the Standard Model Higgs boson with the ATLAS detector at the LHC" } ], "creators": [ { "name": "The ATLAS Collaboration", "nameType": "Organizational" } ], "dates": [ { "date": "2012-07-31T11:59:59Z", "dateType": "Submitted", "dateInformation": "v1" }, { "date": "2012-08-27T18:31:01Z", "dateType": "Updated", "dateInformation": "v1" }, { "date": "2012-08-31T19:29:54Z", "dateType": "Submitted", "dateInformation": "v2" }, { "date": "2012-09-04T00:04:06Z", "dateType": "Updated", "dateInformation": "v2" }, { "date": "2012-07", "dateType": "Available", "dateInformation": "v1" }, { "date": "2012", "dateType": "Issued" } ], "subjects": [ { "lang": "en", "subject": "High Energy Physics - Experiment (hep-ex)", "subjectScheme": "arXiv" }, { "subject": "FOS: Physical sciences", "subjectScheme": "Fields of Science and Technology (FOS)" }, { "subject": "FOS: Physical sciences", "schemeUri": "http://www.oecd.org/science/inno/38235147.pdf", "subjectScheme": "Fields of Science and Technology (FOS)" } ], "relatedIdentifiers": [ { "relationType": "IsVersionOf", "relatedIdentifier": "10.1016/j.physletb.2012.08.020", "relatedIdentifierType": "DOI" } ], "descriptions": [ { "description": "A search for the Standard Model Higgs boson in proton-proton collisions with the ATLAS detector at the LHC is presented. The datasets used correspond to integrated luminosities of approximately 4.8 fb^-1 collected at sqrt(s) = 7 TeV in 2011 and 5.8 fb^-1 at sqrt(s) = 8 TeV in 2012. Individual searches in the channels H->ZZ^(*)->llll, H->gamma gamma and H->WW->e nu mu nu in the 8 TeV data are combined with previously published results of searches for H->ZZ^(*), WW^(*), bbbar and tau^+tau^- in the 7 TeV data and results from improved analyses of the H->ZZ^(*)->llll and H->gamma gamma channels in the 7 TeV data. Clear evidence for the production of a neutral boson with a measured mass of 126.0 +/- 0.4(stat) +/- 0.4(sys) GeV is presented. This observation, which has a significance of 5.9 standard deviations, corresponding to a background fluctuation probability of 1.7x10^-9, is compatible with the production and decay of the Standard Model Higgs boson.", "descriptionType": "Abstract" }, { "description": "24 pages plus author list (38 pages total), 12 figures, 7 tables, revised author list, matches version to appear in Physics Letters B", "descriptionType": "Other" } ], "version": "2" } } } -
2203.02155.json 6.1 KB
{ "data": { "id": "10.48550/arxiv.2203.02155", "type": "dois", "attributes": { "doi": "10.48550/arxiv.2203.02155", "titles": [ { "title": "Training language models to follow instructions with human feedback" } ], "creators": [ { "name": "Ouyang, Long", "givenName": "Long", "familyName": "Ouyang", "nameType": "Personal" }, { "name": "Wu, Jeff", "givenName": "Jeff", "familyName": "Wu", "nameType": "Personal" }, { "name": "Jiang, Xu", "givenName": "Xu", "familyName": "Jiang", "nameType": "Personal" }, { "name": "Almeida, Diogo", "givenName": "Diogo", "familyName": "Almeida", "nameType": "Personal" }, { "name": "Wainwright, Carroll L.", "givenName": "Carroll L.", "familyName": "Wainwright", "nameType": "Personal" }, { "name": "Mishkin, Pamela", "givenName": "Pamela", "familyName": "Mishkin", "nameType": "Personal" }, { "name": "Zhang, Chong", "givenName": "Chong", "familyName": "Zhang", "nameType": "Personal" }, { "name": "Agarwal, Sandhini", "givenName": "Sandhini", "familyName": "Agarwal", "nameType": "Personal" }, { "name": "Slama, Katarina", "givenName": "Katarina", "familyName": "Slama", "nameType": "Personal" }, { "name": "Ray, Alex", "givenName": "Alex", "familyName": "Ray", "nameType": "Personal" }, { "name": "Schulman, John", "givenName": "John", "familyName": "Schulman", "nameType": "Personal" }, { "name": "Hilton, Jacob", "givenName": "Jacob", "familyName": "Hilton", "nameType": "Personal" }, { "name": "Kelton, Fraser", "givenName": "Fraser", "familyName": "Kelton", "nameType": "Personal" }, { "name": "Miller, Luke", "givenName": "Luke", "familyName": "Miller", "nameType": "Personal" }, { "name": "Simens, Maddie", "givenName": "Maddie", "familyName": "Simens", "nameType": "Personal" }, { "name": "Askell, Amanda", "givenName": "Amanda", "familyName": "Askell", "nameType": "Personal" }, { "name": "Welinder, Peter", "givenName": "Peter", "familyName": "Welinder", "nameType": "Personal" }, { "name": "Christiano, Paul", "givenName": "Paul", "familyName": "Christiano", "nameType": "Personal" }, { "name": "Leike, Jan", "givenName": "Jan", "familyName": "Leike", "nameType": "Personal" }, { "name": "Lowe, Ryan", "givenName": "Ryan", "familyName": "Lowe", "nameType": "Personal" } ], "dates": [ { "date": "2022-03-04T07:04:42Z", "dateType": "Submitted", "dateInformation": "v1" }, { "date": "2022-03-07T01:13:10Z", "dateType": "Updated", "dateInformation": "v1" }, { "date": "2022-03", "dateType": "Available", "dateInformation": "v1" }, { "date": "2022", "dateType": "Issued" } ], "subjects": [ { "lang": "en", "subject": "Computation and Language (cs.CL)", "subjectScheme": "arXiv" }, { "lang": "en", "subject": "Artificial Intelligence (cs.AI)", "subjectScheme": "arXiv" }, { "lang": "en", "subject": "Machine Learning (cs.LG)", "subjectScheme": "arXiv" }, { "subject": "FOS: Computer and information sciences", "subjectScheme": "Fields of Science and Technology (FOS)" }, { "subject": "FOS: Computer and information sciences", "schemeUri": "http://www.oecd.org/science/inno/38235147.pdf", "subjectScheme": "Fields of Science and Technology (FOS)" } ], "relatedIdentifiers": [], "descriptions": [ { "description": "Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback. Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback. We call the resulting models InstructGPT. In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. Moreover, InstructGPT models show improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets. Even though InstructGPT still makes simple mistakes, our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent.", "descriptionType": "Abstract" } ], "version": "1" } } } -
2409.03108.json 4 KB
{ "data": { "id": "10.48550/arxiv.2409.03108", "type": "dois", "attributes": { "doi": "10.48550/arxiv.2409.03108", "titles": [ { "title": "Loop Series Expansions for Tensor Networks" } ], "creators": [ { "name": "Evenbly, Glen", "givenName": "Glen", "familyName": "Evenbly", "nameType": "Personal" }, { "name": "Pancotti, Nicola", "givenName": "Nicola", "familyName": "Pancotti", "nameType": "Personal" }, { "name": "Milsted, Ashley", "givenName": "Ashley", "familyName": "Milsted", "nameType": "Personal" }, { "name": "Gray, Johnnie", "givenName": "Johnnie", "familyName": "Gray", "nameType": "Personal" }, { "name": "Chan, Garnet Kin-Lic", "givenName": "Garnet Kin-Lic", "familyName": "Chan", "nameType": "Personal" } ], "dates": [ { "date": "2024-09-04T22:22:35Z", "dateType": "Submitted", "dateInformation": "v1" }, { "date": "2024-09-06T00:09:56Z", "dateType": "Updated", "dateInformation": "v1" }, { "date": "2025-03-26T20:08:06Z", "dateType": "Submitted", "dateInformation": "v2" }, { "date": "2026-03-09T00:15:09Z", "dateType": "Updated", "dateInformation": "v2" }, { "date": "2024-09", "dateType": "Available", "dateInformation": "v1" }, { "date": "2024", "dateType": "Issued" } ], "subjects": [ { "lang": "en", "subject": "Quantum Physics (quant-ph)", "subjectScheme": "arXiv" }, { "lang": "en", "subject": "Disordered Systems and Neural Networks (cond-mat.dis-nn)", "subjectScheme": "arXiv" }, { "subject": "FOS: Physical sciences", "subjectScheme": "Fields of Science and Technology (FOS)" }, { "subject": "FOS: Physical sciences", "schemeUri": "http://www.oecd.org/science/inno/38235147.pdf", "subjectScheme": "Fields of Science and Technology (FOS)" } ], "relatedIdentifiers": [ { "relationType": "IsVersionOf", "relatedIdentifier": "10.1103/vqks-cr6x", "relatedIdentifierType": "DOI" } ], "descriptions": [ { "description": "Belief propagation (BP) can be a useful tool to approximately contract a tensor network, provided that the contributions from any closed loops in the network are sufficiently weak. In this manuscript we describe how a loop series expansion can be applied to systematically improve the accuracy of a BP approximation to a tensor network contraction, in principle converging arbitrarily close to the exact result. More generally, our result provides a framework for expanding a tensor network as a sum of component networks in a hierarchy of increasing complexity. We benchmark this proposal for the contraction of iPEPS, either representing the ground state of an AKLT model or with randomly defined tensors, where it is shown to improve in accuracy over standard BP by several orders of magnitude whilst incurring only a minor increase in computational cost. These results indicate that the proposed series expansions could be a useful tool to accurately evaluate tensor networks in cases that otherwise exceed the limits of established contraction routines.", "descriptionType": "Abstract" }, { "description": "Main text: 6 pages, 5 figures. Appendices: 8 pages, 7 figures. Revised version contains improved results and an additional appendix entry", "descriptionType": "Other" } ], "version": "2" } } } -
2609.14487.json 2.4 KB
{ "data": { "id": "10.48550/arxiv.2609.14487", "type": "dois", "attributes": { "doi": "10.48550/arxiv.2609.14487", "titles": [ { "title": "Equilibrium Transition and Cartel Formation: A Structural Analysis of Chile's Pharmacy Cartel" } ], "creators": [ { "name": "Yu", "familyName": "Yu", "nameType": "Personal" }, { "name": "Hao", "nameType": "Organizational" } ], "dates": [ { "date": "2026-09-13T12:53:27Z", "dateType": "Submitted", "dateInformation": "v1" }, { "date": "2026-09-15T01:08:38Z", "dateType": "Updated", "dateInformation": "v1" }, { "date": "2026-09", "dateType": "Available", "dateInformation": "v1" } ], "subjects": [ { "lang": "en", "subject": "General Economics (econ.GN)", "subjectScheme": "arXiv" }, { "subject": "FOS: Economics and business", "subjectScheme": "Fields of Science and Technology (FOS)" } ], "relatedIdentifiers": [], "descriptions": [ { "description": "This paper studies how Chile's three largest pharmacy chains moved from the price war to collusion, using court-record daily prices and a structural model. After a court-ordered advertisement ban ended the comparison campaign, the estimated gain from loss-leader pricing fell sharply, weakening the incentive to continue the price war. The chains then used an upstream supplier as intermediary to verify a collusive price-leadership procedure. Once verified, they applied it to restore margins on former loss leaders. They then raised prices further to extract rents. These increases spread more slowly than margin restoration. The Adaptive Confidence specification allows a subjective belief about follower participation to update. It more closely matches the cumulative spread of rent extraction than the Unit Confidence specification. Unit Confidence imposes the rational-belief benchmark by fixing the leader's weight at one.", "descriptionType": "Abstract" }, { "description": "45 pages, 15 figures. Online appendix included as an ancillary file", "descriptionType": "Other" } ], "version": "1" } } } -
hep-th_9711200.json 3.5 KB
{ "data": { "id": "10.48550/arxiv.hep-th/9711200", "type": "dois", "attributes": { "doi": "10.48550/arxiv.hep-th/9711200", "titles": [ { "title": "The Large N Limit of Superconformal Field Theories and Supergravity" } ], "creators": [ { "name": "Maldacena, Juan M.", "givenName": "Juan M.", "familyName": "Maldacena", "nameType": "Personal" } ], "dates": [ { "date": "1997-11-27T23:53:13Z", "dateType": "Submitted", "dateInformation": "v1" }, { "date": "2009-11-30T23:46:14Z", "dateType": "Updated", "dateInformation": "v1" }, { "date": "1997-12-08T18:59:11Z", "dateType": "Submitted", "dateInformation": "v2" }, { "date": "2009-11-30T23:46:14Z", "dateType": "Updated", "dateInformation": "v2" }, { "date": "1998-01-22T15:42:41Z", "dateType": "Submitted", "dateInformation": "v3" }, { "date": "2014-11-17T22:19:27Z", "dateType": "Updated", "dateInformation": "v3" }, { "date": "1997-11", "dateType": "Available", "dateInformation": "v1" }, { "date": "1997", "dateType": "Issued" } ], "subjects": [ { "lang": "en", "subject": "High Energy Physics - Theory (hep-th)", "subjectScheme": "arXiv" }, { "subject": "FOS: Physical sciences", "subjectScheme": "Fields of Science and Technology (FOS)" }, { "subject": "FOS: Physical sciences", "schemeUri": "http://www.oecd.org/science/inno/38235147.pdf", "subjectScheme": "Fields of Science and Technology (FOS)" } ], "relatedIdentifiers": [ { "relationType": "IsVersionOf", "relatedIdentifier": "10.1023/a:1026654312961", "relatedIdentifierType": "DOI" } ], "descriptions": [ { "description": "We show that the large $N$ limit of certain conformal field theories in various dimensions include in their Hilbert space a sector describing supergravity on the product of Anti-deSitter spacetimes, spheres and other compact manifolds. This is shown by taking some branes in the full M/string theory and then taking a low energy limit where the field theory on the brane decouples from the bulk. We observe that, in this limit, we can still trust the near horizon geometry for large $N$. The enhanced supersymmetries of the near horizon geometry correspond to the extra supersymmetry generators present in the superconformal group (as opposed to just the super-Poincare group). The 't Hooft limit of 4-d ${\\cal N} =4$ super-Yang-Mills at the conformal point is shown to contain strings: they are IIB strings. We conjecture that compactifications of M/string theory on various Anti-deSitter spacetimes are dual to various conformal field theories. This leads to a new proposal for a definition of M-theory which could be extended to include five non-compact dimensions.", "descriptionType": "Abstract" }, { "description": "20 pages, harvmac, v2: section on AdS_2 corrected, references added, v3: More references and a sign in eqns 2.8 and 2.9 corrected", "descriptionType": "Other" } ], "version": "3" } } }
-
-
-
conftest.py 8.1 KB
"""Shared test fixtures: metadata-lookup outcomes, a network guard, and a ``convert_paper.main()`` launcher. `test_arxiv_metadata.py` builds its own outcomes, because it tests how `MetadataFetch` is constructed. """ import os import sys from pathlib import Path from types import SimpleNamespace import pytest import arxiv_doc_builder from arxiv_doc_builder import convert_paper from arxiv_doc_builder.arxiv_metadata import ( METADATA_OK, METADATA_SOURCE_ARXIV, METADATA_UNAVAILABLE, ArxivMetadata, MetadataFetch, read_metadata_handoff, ) # Resolved via the installed package, not a test file: tests live under tests/ # while the scripts ship in the package and SKILL.md, references/ and # pyproject.toml sit one level up, at the skill root. PACKAGE_DIR = Path(arxiv_doc_builder.__file__).parent SKILL_DIR = PACKAGE_DIR.parent def read_skill_md() -> str: return (SKILL_DIR / "SKILL.md").read_text(encoding="utf-8") PROBE_ERROR = "OSError: connection reset" PROBE_VERSION = "2409.03108v2" # The environment variable naming the file the network guard records into. NETWORK_RECORD_ENV = "ARXIV_DOC_BUILDER_TEST_NETWORK_RECORD" # Imported at startup by every Python process a test starts, through PYTHONPATH. _NETWORK_GUARD = """\ import os import urllib.request # Chain to the sitecustomize this one displaces. Python imports only the first # on sys.path, so an environment that ships its own would otherwise lose it in # every Python process a test starts, and those processes would behave # differently under pytest than outside it. PathFinder is asked rather than # importlib.util.find_spec, which would answer with this half-imported module. def _chain_displaced_sitecustomize(): import sys from importlib.machinery import PathFinder here = os.path.dirname(os.path.abspath(__file__)) rest = [p for p in sys.path if os.path.abspath(p or os.getcwd()) != here] spec = PathFinder.find_spec("sitecustomize", rest) if spec is not None and spec.loader is not None: from importlib.util import module_from_spec spec.loader.exec_module(module_from_spec(spec)) try: _chain_displaced_sitecustomize() except Exception: # The guard itself must survive whatever the displaced module does. pass def _refuse(url, *args, **kwargs): target = getattr(url, "full_url", url) with open(os.environ["ARXIV_DOC_BUILDER_TEST_NETWORK_RECORD"], "a", encoding="utf-8") as record: record.write(f"{target}\\n") raise OSError(f"network access is blocked in tests: {target}") urllib.request.urlopen = _refuse """ @pytest.fixture(autouse=True) def network_record(tmp_path_factory, monkeypatch): """Refuse and record every ``urlopen`` in the Python processes a test starts. A test that runs a script as a subprocess is out of reach of in-process patches, and a failed lookup only prints a warning, so a real request would pass unnoticed. The guard is a ``sitecustomize`` module on ``PYTHONPATH``, which a child interpreter imports at startup. It does not reach the pytest process itself, where tests replace the lookup or the transport directly, and it leaves non-Python children such as ``uv`` and ``curl`` alone. Yields the record file, and fails the test at teardown if a URL was appended to it. """ guard_dir = tmp_path_factory.mktemp("network-guard") (guard_dir / "sitecustomize.py").write_text(_NETWORK_GUARD, encoding="utf-8") record = guard_dir / "attempts.log" record.write_text("", encoding="utf-8") inherited = os.environ.get("PYTHONPATH") monkeypatch.setenv( "PYTHONPATH", os.pathsep.join(p for p in (str(guard_dir), inherited) if p) ) monkeypatch.setenv(NETWORK_RECORD_ENV, str(record)) yield record # Read with the builtin, not Path.read_text: a test may still have that # method patched when this teardown runs. with open(record, encoding="utf-8") as handle: attempts = handle.read() if attempts: pytest.fail(f"a subprocess tried to reach the network:\n{attempts}") @pytest.fixture def failed_probe() -> MetadataFetch: """A lookup that never reached a record.""" return MetadataFetch(METADATA_UNAVAILABLE, error=PROBE_ERROR) @pytest.fixture def probe_with_version() -> MetadataFetch: """A record that was read and carries a version. Every ``ok`` outcome a lookup produces names the source it was read from, and a handoff written from one that does not is rejected, so these carry one. """ return MetadataFetch( METADATA_OK, metadata=ArxivMetadata(version=PROBE_VERSION, source=METADATA_SOURCE_ARXIV), ) @pytest.fixture def probe_without_version() -> MetadataFetch: """A record that was read but carries no version. The cell that separates "the lookup failed" from "no version to record": both leave the sidecar unwritten, by different routes. """ return MetadataFetch( METADATA_OK, metadata=ArxivMetadata(version=None, source=METADATA_SOURCE_ARXIV), ) @pytest.fixture def patch_fetch(monkeypatch): """Point a module's ``fetch_metadata`` at a fixed outcome. Each caller-level test replaces the same name in one module or the other, so the module is the only thing that varies. """ def install(module, probe: MetadataFetch) -> None: monkeypatch.setattr(module, "fetch_metadata", lambda _id: probe) return install def status_of(document: Path) -> str: """The ``metadata_status`` value the document's frontmatter carries.""" for line in document.read_text(encoding="utf-8").splitlines(): if line.startswith("metadata_status:"): return line.split(":", 1)[1].strip().strip('"') raise AssertionError(f"no metadata_status line in {document}") def seed_cached_source(paper_dir: Path) -> None: """Put a cached LaTeX source under ``paper_dir``. A recorded revision outranks the one a lookup reports only while a source is cached, since that material is what the record speaks for. A test of that rule therefore has to seed one, and one that seeds none is testing the opposite branch. """ source = paper_dir / "source" source.mkdir(parents=True, exist_ok=True) (source / "main.tex").write_text("x", encoding="utf-8") def refuse_lookup(_arxiv_id: str) -> MetadataFetch: """A ``fetch_metadata`` stand-in for steps that must not look anything up.""" raise AssertionError("a step given a metadata handoff looked the record up") # The paper the ``launch`` fixture converts, and the lookup it answers with. LAUNCH_ARXIV_ID = "2409.03108" LAUNCH_LOOKUP = MetadataFetch( METADATA_OK, metadata=ArxivMetadata( title="Handed Over", version="2409.03108v2", source=METADATA_SOURCE_ARXIV ), ) @pytest.fixture def launch(monkeypatch, tmp_path): """Run ``convert_paper.main()`` recording the lookups made and the children started. Each recorded child carries its script name, its handoff path, and what the handoff held when the child started (``None`` when no file was there). ``exit_codes`` maps a script name to the code its child returns. """ state = SimpleNamespace(lookups=[], children=[], output_dir=tmp_path) def fake_fetch(arxiv_id): state.lookups.append(arxiv_id) return LAUNCH_LOOKUP def run(arxiv_id=LAUNCH_ARXIV_ID, *, exit_codes=None): codes = exit_codes or {} def fake_run_script(script_name, args, use_uv=False): handoff = Path(args[args.index("--metadata-handoff") + 1]) state.children.append( SimpleNamespace( script=script_name, handoff=handoff, read=read_metadata_handoff(handoff, LAUNCH_ARXIV_ID) if handoff.is_file() else None, ) ) return codes.get(script_name, 0) monkeypatch.setattr(convert_paper, "fetch_metadata", fake_fetch) monkeypatch.setattr(convert_paper, "run_script", fake_run_script) monkeypatch.setattr( sys, "argv", ["convert_paper.py", arxiv_id, "--output-dir", str(tmp_path)] ) convert_paper.main() state.run = run return state -
test_arxiv_id.py 3.7 KB
"""Contract tests for arXiv ID validation and filesystem helpers. Contract: ``validate_arxiv_id`` accepts IDs that are already in canonical form (and only those); it refuses non-canonical new-style inputs that would otherwise silently resolve to a *different* paper via arXiv's zero-padding redirect (e.g. 2202.1173 → 2202.01173). """ import pytest from arxiv_doc_builder.arxiv_id import safe_arxiv_id, validate_arxiv_id # --- accepted forms ----------------------------------------------------- @pytest.mark.parametrize( "arxiv_id", [ # Post-2015 5-digit IDs (with and without version suffix) "2506.01376", "2506.01376v1", "2506.01376v12", "1501.00001", # first month of 5-digit era "2202.11737", # 2007-04 through 2014-12 4-digit IDs "0704.0001", # first month of new-style scheme "1412.7890", # last 4-digit month "1412.0789v3", # Legacy archive/YYMMNNN IDs — short-uppercase subclass form "hep-th/9901001", "math/0703001", "cond-mat/0601234", "math.AG/0703001", "nlin.AO/0601001", "hep-th/9901001v2", # Legacy boundary months (first and last months of the scheme) "hep-th/9108001", "hep-th/0703999", # Legacy IDs with lowercase / hyphenated subject classes "physics.optics/0501001", "physics.comp-ph/0612001", "cond-mat.str-el/0601234", "cond-mat.stat-mech/0601234v2", ], ) def test_canonical_ids_accepted(arxiv_id): assert validate_arxiv_id(arxiv_id) == arxiv_id # --- rejected forms ----------------------------------------------------- @pytest.mark.parametrize( "arxiv_id", [ # Post-2015 paper with 4 digits: would silently remap to a zero-padded # neighbour. The exact case that motivated the validator. "2202.1173", "1501.0001", # boundary month "2506.1376", "2506.1376v2", # Pre-2015 paper with 5 digits: no such paper exists in that window. "1412.00001", "0704.00001", # Before April 2007 (new-style scheme hadn't started yet). "0701.0001", "0703.0001", # Invalid month. "2013.0001", # month 13 "2500.00001", # month 00 # Structurally malformed. "2506", "2506.", "2506.123", # 3 digits: neither canonical width "2506.1234567", # too many digits "abcd.12345", "", # Semantically impossible: sequences start at 1, not 0. "2506.00000", "1412.0000", "hep-th/9901000", # Semantically impossible: versions start at v1. "2506.01376v0", "1412.1234v0", "hep-th/9901001v0", # Legacy form but YYMM outside the scheme (Apr 2007 onwards is new-style). "hep-th/0704001", "hep-th/0801001", "hep-th/1501001", # Legacy form with pre-arXiv YYMM (before Aug 1991). "hep-th/9101001", "hep-th/9107001", # Legacy form with YY outside any known window. "hep-th/4501001", # Legacy form with invalid month. "hep-th/9913001", ], ) def test_noncanonical_ids_rejected(arxiv_id): with pytest.raises(ValueError): validate_arxiv_id(arxiv_id) # --- filesystem helper -------------------------------------------------- @pytest.mark.parametrize( "arxiv_id, expected", [ ("2506.01376", "2506.01376"), # new-style: no change ("hep-th/9901001", "hep-th_9901001"), # legacy: slash → underscore ("math.AG/0703001v2", "math.AG_0703001v2"), ], ) def test_safe_arxiv_id(arxiv_id, expected): assert safe_arxiv_id(arxiv_id) == expected -
test_arxiv_metadata.py 59.9 KB
"""Contract tests for the metadata lookup and the frontmatter producer. The schema is total (every key always present; an unsupplied value renders as YAML null, or ``[]`` for ``categories``) and the text is valid YAML. Round-trips use PyYAML as an independent oracle, skipped when it is not installed; the structural assertions need no dependency. The lookup tests replace the HTTP transport and read DataCite responses from ``fixtures/datacite``, reduced from real responses to the attributes the mapping reads (each entry kept as DataCite serves it) plus the record's DOI, type and version count, so nothing here reaches the network. """ import io import json import math import subprocess import sys import textwrap import threading import time import types import urllib.error from decimal import Decimal from fractions import Fraction from pathlib import Path from typing import Optional import pytest import arxiv_doc_builder.arxiv_metadata as arxiv_metadata from arxiv_doc_builder.arxiv_metadata import ( METADATA_NOT_REQUESTED, METADATA_OK, METADATA_STATUSES, METADATA_UNAVAILABLE, ArxivMetadata, MetadataFetch, build_frontmatter, fetch_metadata, format_unavailable_warning, read_metadata_handoff, write_metadata_handoff, ) # The total key set every frontmatter must carry, in any state. FRONTMATTER_KEYS = { "title", "authors", "arxiv_id", "version", "published", "primary_category", "categories", "doi", "journal", "source_type", "metadata_status", "conversion_date", "abstract", } _FIXTURES = Path(__file__).parent / "fixtures" / "datacite" _FULL = ArxivMetadata( title="A Study of Things", authors=["Chiara Capecci", "C. Balázs", "C. -P. Yuan"], version="2606.09995v1", published="2026-06-08", primary_category="quant-ph", categories=["quant-ph", "cond-mat.str-el"], doi="10.1103/PhysRevD.76.013009", journal="Phys. Rev. D 76, 013009 (2007)", abstract="Line one.\n wrapped with odd spacing\nand a colon: here.", source=arxiv_metadata.METADATA_SOURCE_ARXIV, ) def _fm(meta: Optional[ArxivMetadata] = None, **kwargs) -> str: """``build_frontmatter`` with the provenance arguments defaulted. Every call needs an id, a source type, a date and a status. Defaulting them keeps each test's arguments down to what it asserts about. """ kwargs.setdefault("arxiv_id", "2606.09995") kwargs.setdefault("source_type", "latex") kwargs.setdefault("conversion_date", "d") kwargs.setdefault("metadata_status", METADATA_OK) return build_frontmatter(meta, **kwargs) def _parse(frontmatter: str): """Parse the inner YAML of a ``---``-fenced frontmatter block via PyYAML.""" yaml = pytest.importorskip("yaml") inner = frontmatter.split("---\n", 1)[1].rsplit("\n---", 1)[0] return yaml.safe_load(inner) def _record(arxiv_id: str) -> bytes: """The stored DataCite response for ``arxiv_id``.""" return (_FIXTURES / (arxiv_id.replace("/", "_") + ".json")).read_bytes() # --- structural contract (no third-party dependency) ---------------------- def test_all_keys_present_in_full_metadata(): fm = _fm(_FULL, conversion_date="2026-06-11T10:00:00") for key in FRONTMATTER_KEYS: assert f"{key}:" in fm, f"missing key {key!r} in frontmatter" assert fm.startswith("---\n") assert "\n---\n" in fm def test_absent_doi_journal_render_as_bare_null_keys(): # A record with no DOI must emit `doi:` (null), distinct from omitting the # key. A bare `key:` line, not `key: ""`, is what a parser reads as None. meta = ArxivMetadata(title="T", doi=None, journal=None) fm = _fm(meta) assert "\ndoi:\n" in fm assert "\njournal:\n" in fm assert 'doi: ""' not in fm def test_the_document_does_not_name_which_source_answered(): # Both sources render the same key set, and none of its keys names the # source. A reader that had to know which record answered would need the # converter to say what a null implies about the paper, which is the # bibliographic question it does not answer. from_arxiv = _parse(_fm(_FULL)) from_datacite = _parse( _fm( ArxivMetadata( title="T", version="2606.09995v1", source=arxiv_metadata.METADATA_SOURCE_DATACITE, ) ) ) assert set(from_arxiv) == FRONTMATTER_KEYS assert set(from_datacite) == FRONTMATTER_KEYS assert "metadata_source" not in FRONTMATTER_KEYS for parsed in (from_arxiv, from_datacite): assert arxiv_metadata.METADATA_SOURCE_ARXIV not in parsed.values() assert arxiv_metadata.METADATA_SOURCE_DATACITE not in parsed.values() def test_absent_arxiv_id_renders_as_null(): # Manual PDF scripts invoke the converter without an id. fm = _fm( None, arxiv_id=None, source_type="pdf", metadata_status=METADATA_NOT_REQUESTED, fallback_title="From PDF", ) assert "\narxiv_id:\n" in fm assert "From PDF" in fm # --- round-trip contract (PyYAML oracle) ---------------------------------- def test_full_metadata_round_trips(): fm = _fm(_FULL, conversion_date="2026-06-11T10:00:00") parsed = _parse(fm) assert set(parsed.keys()) == FRONTMATTER_KEYS assert parsed["title"] == "A Study of Things" # Authors are joined into a single citation-friendly string; Unicode is # preserved through the double-quoted scalar. assert parsed["authors"] == "Chiara Capecci, C. Balázs, C. -P. Yuan" assert parsed["arxiv_id"] == "2606.09995" assert parsed["version"] == "2606.09995v1" assert parsed["published"] == "2026-06-08" assert parsed["primary_category"] == "quant-ph" assert parsed["categories"] == ["quant-ph", "cond-mat.str-el"] assert parsed["doi"] == "10.1103/PhysRevD.76.013009" assert parsed["journal"] == "Phys. Rev. D 76, 013009 (2007)" assert parsed["source_type"] == "latex" # Abstract is whitespace-normalized to a single paragraph. assert parsed["abstract"] == "Line one. wrapped with odd spacing and a colon: here." def test_absent_fields_parse_to_none(): meta = ArxivMetadata(title="T", doi=None, journal=None, abstract=None) parsed = _parse(_fm(meta)) assert "doi" in parsed and parsed["doi"] is None assert "journal" in parsed and parsed["journal"] is None assert "abstract" in parsed and parsed["abstract"] is None assert parsed["categories"] == [] assert parsed["authors"] is None def test_no_metadata_keeps_title_null_not_fabricated(): # A PDF with no embedded title and no arXiv id keeps the title null rather # than fabricating one from the file name — "unknown stays unknown". parsed = _parse( _fm( None, arxiv_id=None, source_type="pdf", metadata_status=METADATA_NOT_REQUESTED, ) ) assert parsed["title"] is None assert parsed["authors"] is None assert parsed["arxiv_id"] is None assert parsed["source_type"] == "pdf" def test_offline_metadata_keeps_total_schema_with_fallback_title(): parsed = _parse( _fm( None, metadata_status=METADATA_UNAVAILABLE, fallback_title="LaTeX Title", ) ) assert set(parsed.keys()) == FRONTMATTER_KEYS assert parsed["title"] == "LaTeX Title" assert parsed["version"] is None assert parsed["arxiv_id"] == "2606.09995" def test_tricky_title_round_trips(): meta = ArxivMetadata(title='Tricky: "quotes", colon: and \\backslash') parsed = _parse(_fm(meta, arxiv_id="x")) assert parsed["title"] == 'Tricky: "quotes", colon: and \\backslash' def test_pdf_style_raw_author_with_newline_stays_valid_yaml(): # The PDF fallback path builds ArxivMetadata from raw embedded metadata, # which bypasses the record-side normalization. An author carrying an # embedded newline (common in malformed PDF /Author fields) must not corrupt # the YAML; build_frontmatter normalizes it to a single line. meta = ArxivMetadata(title="T", authors=["Jane Doe\n--- affiliation"]) parsed = _parse(_fm(meta, arxiv_id="x", source_type="pdf")) assert parsed["authors"] == "Jane Doe --- affiliation" def test_non_printable_characters_are_stripped_and_yaml_stays_valid(): # All reachable: raw C1 controls in arXiv's XML, \u escapes in DataCite's # JSON, U+FFFF from pypdf decoding \xff\xff. The abstract's block scalar # cannot escape them, so they are dropped. controls = "".join(chr(c) for c in (0x07, 0x1B, 0x80, 0x9F, 0x7F, 0xFFFE, 0xFFFF)) meta = ArxivMetadata( title="A" + controls + "B", authors=["Jo" + controls + "hn"], abstract="Clean" + controls + "Abstract", ) parsed = _parse(_fm(meta, arxiv_id="x", source_type="pdf")) assert parsed["title"] == "AB" assert parsed["authors"] == "John" assert parsed["abstract"] == "CleanAbstract" # --- lookup: the transport, arXiv's feed and DataCite's records ------------- _ATOM_ENTRY = b"""<?xml version="1.0" encoding="UTF-8"?> <feed xmlns="http://www.w3.org/2005/Atom" xmlns:arxiv="http://arxiv.org/schemas/atom"> <entry> <id>http://arxiv.org/abs/2606.09995v2</id> <title>A Study of Things</title> <published>2026-06-08T00:00:00Z</published> <summary>An abstract.</summary> <author><name>Chiara Capecci</name></author> <arxiv:primary_category term="quant-ph"/> <category term="quant-ph"/> <category term="cond-mat.str-el"/> <arxiv:doi>10.1103/PhysRevD.76.013009</arxiv:doi> <arxiv:journal_ref>Phys. Rev. D 76, 013009 (2007)</arxiv:journal_ref> </entry> </feed> """ _ATOM_NO_ENTRY = b"""<?xml version="1.0" encoding="UTF-8"?> <feed xmlns="http://www.w3.org/2005/Atom"></feed> """ _ATOM_ERROR_ENTRY = b"""<?xml version="1.0" encoding="UTF-8"?> <feed xmlns="http://www.w3.org/2005/Atom"> <entry> <id>http://arxiv.org/api/errors#incorrect_id_format</id> <title>Error</title> <summary>incorrect id format for nonsense</summary> </entry> </feed> """ @pytest.fixture def transport(monkeypatch): """Replace the HTTP transport so no test reaches arXiv or DataCite. ``install`` takes the DataCite leg's outcome and, optionally, the arXiv leg's: bytes (a body the lookup parses), an exception instance (raised in place of the request), or a callable returning a response. A test that installs no arXiv outcome gets a failing arXiv leg, which is what puts the DataCite leg under test. It returns the list the requested URLs are appended to, arXiv's first. """ requested: list[str] = [] def install(datacite, arxiv=None): if arxiv is None: arxiv = OSError("this test installed no arXiv outcome") def respond(outcome): if isinstance(outcome, BaseException): raise outcome if callable(outcome): return outcome() return io.BytesIO(outcome) def fake_urlopen(request, timeout=None): # The lookup passes a Request, so that it can name this client. url = getattr(request, "full_url", request) requested.append(url) on_arxiv = url.startswith(arxiv_metadata._ARXIV_API_URL) return respond(arxiv if on_arxiv else datacite) monkeypatch.setattr(arxiv_metadata.urllib.request, "urlopen", fake_urlopen) return requested return install def _datacite_record(result: MetadataFetch) -> ArxivMetadata: """The record ``result`` carries, asserting that DataCite answered.""" assert result.status == METADATA_OK assert result.metadata is not None assert result.metadata.source == arxiv_metadata.METADATA_SOURCE_DATACITE return result.metadata def _datacite_requests(requested: list[str]) -> list[str]: return [url for url in requested if url.startswith(arxiv_metadata._DATACITE_URL)] def _http_error( code: int, *, url: str = "https://api.datacite.org", fp: Optional[io.BytesIO] = None ) -> urllib.error.HTTPError: """An HTTP failure as urlopen raises it, its body read from ``fp``.""" from email.message import Message return urllib.error.HTTPError( url, code, "msg", Message(), fp if fp is not None else io.BytesIO() ) # --- lookup: arXiv's feed, and the fallback to DataCite --------------------- def test_the_arxiv_record_maps_onto_the_frontmatter_fields(transport): requested = transport(b"unused", arxiv=_ATOM_ENTRY) result = fetch_metadata("2606.09995") assert requested == [arxiv_metadata._ARXIV_API_URL + "?id_list=2606.09995"], ( "arXiv is asked first, and its answer ends the lookup" ) assert result.status == METADATA_OK assert result.metadata is not None assert result.metadata.title == "A Study of Things" assert result.metadata.authors == ["Chiara Capecci"] assert result.metadata.version == "2606.09995v2" assert result.metadata.published == "2026-06-08" assert result.metadata.primary_category == "quant-ph" assert result.metadata.categories == ["quant-ph", "cond-mat.str-el"] assert result.metadata.doi == "10.1103/PhysRevD.76.013009" assert result.metadata.journal == "Phys. Rev. D 76, 013009 (2007)" assert result.metadata.abstract == "An abstract." assert result.metadata.source == arxiv_metadata.METADATA_SOURCE_ARXIV # What arXiv can do instead of answering, and the cause its leg reports. _ARXIV_FAILURES = pytest.mark.parametrize( ("arxiv_outcome", "cause"), [ (_http_error(429), "arXiv rate-limited the request (HTTP 429)"), (_http_error(503), "arXiv server error (HTTP 503)"), (_http_error(403), "HTTP 403"), (OSError("connection reset"), "OSError: connection reset"), (b"<feed>not closed", "arXiv returned malformed XML: "), (_ATOM_NO_ENTRY, "arXiv returned no record for this id"), (_ATOM_ERROR_ENTRY, "incorrect id format for nonsense"), ], ids=["429", "5xx", "other-http", "os-error", "malformed-xml", "no-entry", "error"], ) @_ARXIV_FAILURES def test_datacite_answers_when_arxiv_does_not(transport, arxiv_outcome, cause): requested = transport(_record("2409.03108"), arxiv=arxiv_outcome) result = fetch_metadata("2409.03108") assert _datacite_requests(requested) == [ arxiv_metadata._DATACITE_URL + "2409.03108" ] _datacite_record(result) @_ARXIV_FAILURES def test_each_arxiv_failure_is_named_when_datacite_fails_too( transport, arxiv_outcome, cause ): transport(_http_error(404), arxiv=arxiv_outcome) result = fetch_metadata("2409.03108") assert result.status == METADATA_UNAVAILABLE assert result.failure_cause.startswith(f"arXiv: {cause}") def test_a_failure_on_both_sources_names_what_each_one_said(transport): transport(_http_error(404), arxiv=_http_error(429)) result = fetch_metadata("2606.09995") assert result.status == METADATA_UNAVAILABLE assert result.failure_cause == ( "arXiv: arXiv rate-limited the request (HTTP 429); " "DataCite: DataCite has no record for 10.48550/arXiv.2606.09995 (HTTP 404)" ) @pytest.mark.parametrize( ("stored", "decoded"), [ ("a > b & c", "a > b & c"), ("x ∉ S", "x \u2209 S"), ("AB", "AB"), # Legacy forms without the semicolon are literal text here. ("¬ation and ¶", "¬ation and ¶"), ("&bogus;", "&bogus;"), ], ids=["named", "longer-name", "numeric", "no-semicolon", "unknown-name"], ) def test_datacite_prose_decodes_only_complete_entity_references(stored, decoded): assert arxiv_metadata._prose(stored) == decoded @pytest.mark.parametrize( ("id_url", "version"), [ ("http://arxiv.org/abs/2409.03108v2", "2409.03108v2"), ("http://arxiv.org/abs/hep-th/9901001v3", "hep-th/9901001v3"), ("http://[unclosed", "[unclosed"), ("", None), (None, None), ], ids=["canonical", "legacy", "unparseable-url", "empty", "absent"], ) def test_the_version_comes_from_the_entry_id_tail(id_url, version): assert arxiv_metadata.parse_version_from_id(id_url) == version # --- lookup: the feed has to identify the paper before it is read ----------- _ATOM_ANOTHER_PAPER = _ATOM_ENTRY.replace(b"2606.09995v2", b"2606.09996v2") _ATOM_UNPARSEABLE_ID = _ATOM_ENTRY.replace( b"<id>http://arxiv.org/abs/2606.09995v2</id>", b"<id>http://[unclosed</id>" ) _ATOM_WITHOUT_AN_ID = _ATOM_ENTRY.replace( b" <id>http://arxiv.org/abs/2606.09995v2</id>\n", b"" ) @pytest.mark.parametrize( ("feed", "arxiv_id", "cause"), [ ( _ATOM_ANOTHER_PAPER, "2606.09995", "arXiv answered with a record for 2606.09996, not 2606.09995", ), ( _ATOM_ENTRY, "2606.09995v1", "arXiv answered with 2606.09995v2, not the requested 2606.09995v1", ), ( _ATOM_UNPARSEABLE_ID, "2606.09995", "arXiv's entry id names no revision: [unclosed", ), ( _ATOM_WITHOUT_AN_ID, "2606.09995", "arXiv's entry carries no id to identify the paper by", ), ], ids=["another-paper", "another-revision", "no-revision", "no-id"], ) def test_a_feed_that_does_not_identify_the_paper_is_not_read_as_its_record( transport, monkeypatch, feed, arxiv_id, cause ): # Refused whole, before its other fields are read: an `ok` would put # another paper's fields in the document and keep the fallback unasked. parsed: list[object] = [] monkeypatch.setattr(arxiv_metadata, "_parse_entry", parsed.append) transport(_http_error(404), arxiv=feed) result = fetch_metadata(arxiv_id) assert result.status == METADATA_UNAVAILABLE assert result.metadata is None assert result.failure_cause.startswith(f"arXiv: {cause}") assert parsed == [] def test_a_feed_naming_another_paper_leaves_the_fallback_to_answer(transport): # The other half of the rule above: the refusal is what lets DataCite run, # so the paper still gets a record rather than the run recording none. requested = transport(_record("2409.03108"), arxiv=_ATOM_ANOTHER_PAPER) result = fetch_metadata("2409.03108") assert _datacite_requests(requested) == [ arxiv_metadata._DATACITE_URL + "2409.03108" ] assert _datacite_record(result).title != "A Study of Things" def test_an_entry_id_the_url_parser_rejects_still_identifies_the_paper(transport): # The entry id is untrusted text, and an unclosed IPv6 bracket makes # urlparse raise. The tail split recovers an id that does name this # paper's revision, so the record is read rather than refused — the # identity check tests what the id says, not how it parsed. feed = _ATOM_ENTRY.replace( b"<id>http://arxiv.org/abs/2606.09995v2</id>", b"<id>http://[::1/abs/2606.09995v2</id>", ) transport(b"unused", arxiv=feed) result = fetch_metadata("2606.09995") assert result.status == METADATA_OK assert result.metadata is not None assert result.metadata.version == "2606.09995v2" # --- lookup: DataCite's records --------------------------------------------- @pytest.mark.parametrize( ("arxiv_id", "expected"), [ ( "2409.03108", { "title": "Loop Series Expansions for Tensor Networks", "authors": [ "Glen Evenbly", "Nicola Pancotti", "Ashley Milsted", "Johnnie Gray", "Garnet Kin-Lic Chan", ], "version": "2409.03108v2", "published": "2024-09-04", "primary_category": "quant-ph", "categories": ["quant-ph", "cond-mat.dis-nn"], "doi": "10.1103/vqks-cr6x", }, ), ( # A collaboration: DataCite gives only `name`. "1207.7214", { "authors": ["The ATLAS Collaboration"], "version": "1207.7214v2", "published": "2012-07-31", "primary_category": "hep-ex", "categories": ["hep-ex"], "doi": "10.1016/j.physletb.2012.08.020", }, ), ( # "Yu" is Personal without a givenName and "Hao" a single name tagged # Organizational; `name` holds each as arXiv lists them. "2609.14487", { "authors": ["Yu", "Hao"], "version": "2609.14487v1", "primary_category": "econ.GN", "doi": None, }, ), ( # A legacy id keeps its slash, and DataCite stores the DOI lowercased. "hep-th/9711200", { "authors": ["Juan M. Maldacena"], "version": "hep-th/9711200v3", "published": "1997-11-27", "categories": ["hep-th"], "doi": "10.1023/a:1026654312961", }, ), ( # The primary category comes first even though it is not the # alphabetically first code. "2203.02155", { "version": "2203.02155v1", "published": "2022-03-04", "primary_category": "cs.CL", "categories": ["cs.CL", "cs.AI", "cs.LG"], "doi": None, }, ), ], ids=lambda value: value if isinstance(value, str) else "", ) def test_a_datacite_record_maps_onto_the_frontmatter_fields( transport, arxiv_id, expected ): requested = transport(_record(arxiv_id)) result = fetch_metadata(arxiv_id) assert _datacite_requests(requested) == [arxiv_metadata._DATACITE_URL + arxiv_id] assert result.status == METADATA_OK assert result.metadata is not None for field_name, value in expected.items(): assert getattr(result.metadata, field_name) == value, field_name assert result.metadata.title assert result.metadata.abstract def test_the_journal_key_is_null_on_a_record_that_carries_no_such_field(transport): # Only arXiv's feed has a journal reference. DataCite's registration has no # such field, so the key is null on its records — which the document states # as a property of the answering record, not of the paper. transport(_record("2409.03108")) result = fetch_metadata("2409.03108") assert result.metadata is not None assert result.metadata.journal is None def test_the_abstract_is_the_abstract_description_not_the_comments(transport): transport(_record("2409.03108")) result = fetch_metadata("2409.03108") assert result.metadata is not None assert result.metadata.abstract is not None assert result.metadata.abstract.startswith("Belief propagation (BP)") assert "Main text: 6 pages" not in result.metadata.abstract def test_html_entities_in_record_prose_are_decoded(transport): # DataCite's abstract for 1207.7214 stores "H->ZZ" as "H->ZZ". transport(_record("1207.7214")) result = fetch_metadata("1207.7214") assert result.metadata is not None assert result.metadata.abstract is not None assert "H->ZZ" in result.metadata.abstract assert ">" not in result.metadata.abstract def test_html_entities_in_titles_and_author_names_are_decoded(): attributes = { "titles": [{"title": "A & B"}], "creators": [ {"givenName": "Jörg", "familyName": "Müller"}, {"name": "R&D Group"}, ], } meta = arxiv_metadata._parse_record("2001.00001", attributes) assert meta.title == "A & B" assert meta.authors == ["Jörg Müller", "R&D Group"] def test_an_atom_title_is_not_decoded_twice(transport): # The XML parser has already decoded the entities, so running the prose # decoder over them again would turn a title that spells ">" literally # into ">". feed = _ATOM_ENTRY.replace( b"<title>A Study of Things</title>", b"<title>Writing &gt; in a title</title>", ) transport(b"unused", arxiv=feed) result = fetch_metadata("2606.09995") assert result.metadata is not None assert result.metadata.title == "Writing > in a title" @pytest.mark.parametrize("revision", [1, 2]) def test_a_requested_revision_is_recorded_and_looked_up_by_the_bare_id( transport, revision ): requested = transport(_record("2409.03108")) result = fetch_metadata(f"2409.03108v{revision}") assert _datacite_requests(requested) == [ arxiv_metadata._DATACITE_URL + "2409.03108" ] assert result.status == METADATA_OK assert result.metadata is not None assert result.metadata.version == f"2409.03108v{revision}" def test_the_arxiv_request_keeps_a_requested_revision(transport): # The two legs treat a pinned id oppositely, which is why the frontmatter # documents them apart. DataCite holds one record per paper, so the test # above asks it by the bare id and the record follows the latest revision. # arXiv answers per revision, so it is asked for the one requested and its # record describes that revision rather than the paper's latest. feed = _ATOM_ENTRY.replace(b"2606.09995v2", b"2606.09995v1") requested = transport(b"unused", arxiv=feed) result = fetch_metadata("2606.09995v1") assert requested == [arxiv_metadata._ARXIV_API_URL + "?id_list=2606.09995v1"] assert result.status == METADATA_OK assert result.metadata is not None assert result.metadata.version == "2606.09995v1" @pytest.mark.parametrize( ("arxiv_id", "doi_id"), [ ("math.GT/0309136", "math/0309136"), ("cond-mat.str-el/0601234v2", "cond-mat/0601234"), ("hep-th/9711200", "hep-th/9711200"), ], ) def test_a_legacy_subject_class_is_left_out_of_the_doi_looked_up( transport, arxiv_id, doi_id ): # DataCite answers 404 for 10.48550/arXiv.math.GT/0309136 and 200 for # 10.48550/arXiv.math/0309136. requested = transport(_http_error(404)) result = fetch_metadata(arxiv_id) assert _datacite_requests(requested) == [arxiv_metadata._DATACITE_URL + doi_id] assert result.failure_cause.endswith( f"DataCite: DataCite has no record for 10.48550/arXiv.{doi_id} (HTTP 404)" ) @pytest.mark.parametrize( ("arxiv_id", "version"), [ ("math.GT/0309136", "math/0309136v2"), ("math.GT/0309136v1", "math/0309136v1"), ], ) def test_both_sources_spell_a_legacy_version_the_way_arxiv_does( transport, arxiv_id, version ): # The version-drift check compares this string with the sidecar's. A # DataCite answer spelling it with the subject class would read as another # paper, and the fetch step would delete the cached source over it. The # stored record lists v1 and v2; it is relabelled as this paper's, since # DataCite names the DOI it answers for and the lookup checks it. document = json.loads(_record("2409.03108")) document["data"]["id"] = "10.48550/arxiv.math/0309136" transport(json.dumps(document).encode()) result = fetch_metadata(arxiv_id) assert result.metadata is not None assert result.metadata.version == version def test_the_arxiv_request_drops_a_legacy_subject_class(transport): # arXiv answers a legacy id only under the archive alone; the subject-class # spelling returns an empty feed, which would send every such paper to the # fallback and lose the journal reference only arXiv's record carries. # The entry comes back spelled the way arXiv spells it, which is also the # spelling the identity check has to compare against. feed = _ATOM_ENTRY.replace(b"2606.09995v2", b"math/0309136v2") requested = transport(b"unused", arxiv=feed) result = fetch_metadata("math.GT/0309136") assert requested == [arxiv_metadata._ARXIV_API_URL + "?id_list=math%2F0309136"] assert result.status == METADATA_OK assert result.metadata is not None assert result.metadata.version == "math/0309136v2" def test_a_rejected_id_is_reported_by_what_arxiv_said(transport): # arXiv answers a malformed id with HTTP 400 whose body holds an error # entry. The entry says what was wrong with the id; the status does not. rejected = _http_error( 400, url=arxiv_metadata._ARXIV_API_URL, fp=io.BytesIO(_ATOM_ERROR_ENTRY) ) transport(_http_error(404), arxiv=rejected) result = fetch_metadata("2606.09995") assert result.failure_cause.startswith( "arXiv: arXiv rejected the id: incorrect id format for nonsense" ) def test_a_rejected_id_without_a_summary_is_reported_by_its_status(transport): feed = b"""<?xml version="1.0" encoding="UTF-8"?> <feed xmlns="http://www.w3.org/2005/Atom"> <entry><id>https://arxiv.org/api/errors#unspecified</id></entry> </feed> """ rejected = _http_error(400, url=arxiv_metadata._ARXIV_API_URL, fp=io.BytesIO(feed)) transport(_http_error(404), arxiv=rejected) result = fetch_metadata("2606.09995") assert result.failure_cause.startswith("arXiv: HTTP 400;") def test_an_error_entry_without_a_summary_still_names_a_cause(transport): # The warning prints the cause verbatim. An error entry can arrive without # a summary, and the fallback literal is what keeps that line from being # blank. feed = b"""<?xml version="1.0" encoding="UTF-8"?> <feed xmlns="http://www.w3.org/2005/Atom"> <entry> <id>https://arxiv.org/api/errors#unspecified</id> <title>Error</title> </entry> </feed> """ transport(_http_error(404), arxiv=feed) result = fetch_metadata("2606.09995") assert result.failure_cause.startswith("arXiv: arXiv reported an error for this id") def test_a_withdrawn_revision_counts_as_the_latest(transport): # DataCite lists a withdrawal as `Withdrawn`, labelled "v2; None", while # arXiv's own record names v2 as the paper's latest. Counting only # `Submitted` entries would have the two sources disagree, and the # version-drift check reads a disagreement as a new revision. transport(_record("0705.1442")) result = fetch_metadata("0705.1442") assert result.metadata is not None assert result.metadata.version == "0705.1442v2" # `published` stays the first revision's date, as arXiv's record reports it. assert result.metadata.published == "2007-05-10" def test_both_requests_name_this_client(monkeypatch): # arXiv asks API clients to identify themselves, and an unidentified one is # the first a public API throttles — the failure this lookup exists to # survive. agents: list[Optional[str]] = [] def fake_urlopen(request, timeout=None): agents.append(request.get_header("User-agent")) raise OSError("offline") monkeypatch.setattr(arxiv_metadata.urllib.request, "urlopen", fake_urlopen) fetch_metadata("2409.03108") assert agents == [arxiv_metadata._USER_AGENT, arxiv_metadata._USER_AGENT] def test_a_trickling_arxiv_leg_leaves_the_fallback_its_time(transport): # urlopen's timeout bounds each socket operation, not the request, so a # response arriving in slow pieces — what a rate-limited service often # does — would spend the whole deadline and the fallback would never be # asked, in exactly the case the fallback was added for. def trickle(): time.sleep(1.0) return io.BytesIO(_ATOM_ENTRY) transport(_record("2409.03108"), arxiv=trickle) started = time.monotonic() result = fetch_metadata("2409.03108", deadline=0.4) elapsed = time.monotonic() - started _datacite_record(result) assert elapsed < 0.9, "the arXiv leg was abandoned, not waited out" def test_an_error_response_is_closed_on_each_leg(transport): # An HTTPError is an open response holding a socket, and urlopen raises it # before any `with` can bind it, so each leg closes it by hand. arxiv_body = io.BytesIO(_ATOM_ERROR_ENTRY) datacite_body = io.BytesIO(b"{}") transport( _http_error(404, url=arxiv_metadata._DATACITE_URL, fp=datacite_body), arxiv=_http_error(400, url=arxiv_metadata._ARXIV_API_URL, fp=arxiv_body), ) fetch_metadata("2606.09995") assert arxiv_body.closed assert datacite_body.closed def test_an_unclassified_datacite_failure_keeps_the_arxiv_cause(transport): # The mirror image of the arXiv leg's guard: an exception DataCite's leg # does not classify must not report itself alone, dropping what arXiv said. import http.client transport(http.client.IncompleteRead(b"half"), arxiv=_http_error(429)) result = fetch_metadata("2409.03108") assert result.status == METADATA_UNAVAILABLE assert result.failure_cause.startswith( "arXiv: arXiv rate-limited the request (HTTP 429); DataCite: IncompleteRead" ) def test_an_unclassified_arxiv_failure_still_reaches_the_fallback(transport): # http.client's exceptions are not OSError, so a response dropped midway — # which is what an overloaded arXiv does — must not skip DataCite. import http.client requested = transport( _record("2409.03108"), arxiv=http.client.IncompleteRead(b"half") ) result = fetch_metadata("2409.03108") assert _datacite_requests(requested) == [ arxiv_metadata._DATACITE_URL + "2409.03108" ] _datacite_record(result) def test_a_requested_revision_with_no_listed_revisions_is_ok(transport): # With no Submitted dates there is no latest revision to compare against, so # the requested one is not rejected. transport( b'{"data": {"id": "10.48550/arxiv.2001.00001",' b' "attributes": {"titles": [{"title": "T"}]}}}' ) result = fetch_metadata("2001.00001v3") assert result.status == METADATA_OK assert result.metadata is not None assert result.metadata.version == "2001.00001v3" def test_a_requested_revision_beyond_the_record_is_unavailable(transport): transport(_record("2409.03108")) result = fetch_metadata("2409.03108v3") assert result.status == METADATA_UNAVAILABLE assert result.failure_cause.endswith( "DataCite: 2409.03108v3 is later than v2, the latest revision DataCite " "lists for 2409.03108" ) # --- lookup: mapping rules on shapes the stored records do not cover -------- def _submitted(*labels: str) -> list[dict]: return [ { "date": f"2020-01-{i + 1:02d}T00:00:00Z", "dateType": "Submitted", "dateInformation": label, } for i, label in enumerate(labels) ] @pytest.mark.parametrize( ("document", "named"), [ (_record("2409.03108"), "'10.48550/arxiv.2409.03108'"), (b'{"data": {"attributes": {"titles": [{"title": "T"}]}}}', "None"), ], ids=["another-paper", "no-doi"], ) def test_a_record_not_naming_the_doi_is_refused_unread( transport, monkeypatch, document, named ): # The fields of another paper's record would otherwise reach the document # as this paper's, under metadata_status: ok. The record is refused before # any of them is read. parsed: list[object] = [] monkeypatch.setattr( arxiv_metadata, "_parse_record", lambda *args: parsed.append(args) ) transport(document) result = fetch_metadata("2001.00001") assert result.status == METADATA_UNAVAILABLE assert f"DataCite answered with a record for {named}" in result.failure_cause assert parsed == [] def test_a_record_naming_the_doi_in_another_case_is_read(transport): # DOIs are case-insensitive, so a record spelling the requested one in # upper case names the same paper. document = json.loads(_record("2409.03108")) document["data"]["id"] = "10.48550/ARXIV.2409.03108" transport(json.dumps(document).encode()) result = fetch_metadata("2409.03108") assert result.status == METADATA_OK @pytest.mark.parametrize("date_type", ["Submitted", "Withdrawn"]) def test_a_listed_revision_without_a_usable_date_still_counts_as_latest(date_type): attributes = { "dates": _submitted("v1") + [{"dateType": date_type, "dateInformation": "v2", "date": None}] } meta = arxiv_metadata._parse_record("2001.00001", attributes) assert meta.version == "2001.00001v2" assert meta.published == "2020-01-01" def test_the_latest_revision_compares_numerically(): attributes = {"dates": _submitted("v1", "v9", "v10")} assert ( arxiv_metadata._parse_record("2001.00001", attributes).version == "2001.00001v10" ) def test_no_submitted_date_leaves_version_and_published_null(): attributes = {"dates": [{"date": "2020", "dateType": "Issued"}]} meta = arxiv_metadata._parse_record("2001.00001", attributes) assert meta.version is None assert meta.published is None @pytest.mark.parametrize( ("date", "published"), [ ("2020-01-02T03:04:05Z", "2020-01-02"), ("2020-01-02", "2020-01-02"), ("2020-01", None), ("2020", None), ], ) def test_published_is_null_unless_the_v1_date_carries_a_calendar_date(date, published): attributes = { "dates": [{"date": date, "dateType": "Submitted", "dateInformation": "v1"}] } assert arxiv_metadata._parse_record("2001.00001", attributes).published == published def test_a_record_with_wrongly_shaped_fields_parses_to_empty_fields(): # DataCite is an external producer; a null or mistyped field must not raise. attributes = { "titles": "not a list", "creators": None, "dates": {"v1": "2020"}, "subjects": [5, None], "relatedIdentifiers": 7, "descriptions": [["Abstract"]], } meta = arxiv_metadata._parse_record("2001.00001", attributes) assert meta == ArxivMetadata(source=arxiv_metadata.METADATA_SOURCE_DATACITE) def test_the_first_non_empty_title_is_taken(): attributes = {"titles": [{"title": " "}, {"title": "Second"}]} assert arxiv_metadata._parse_record("2001.00001", attributes).title == "Second" def test_a_creator_without_both_name_parts_falls_back_to_name(): attributes = { "creators": [ {"name": "Doe, Jane", "givenName": "Jane", "familyName": "Doe"}, {"name": "Plato", "givenName": "Plato"}, {"name": "Nobody Given", "familyName": "Given"}, {"givenName": "", "familyName": ""}, ] } authors = arxiv_metadata._parse_record("2001.00001", attributes).authors assert authors == ["Jane Doe", "Plato", "Nobody Given"] def test_only_arxiv_subjects_with_a_code_become_categories(): attributes = { "subjects": [ {"subject": "Mathematical Physics (math-ph)", "subjectScheme": "arXiv"}, { "subject": "FOS: Physical sciences", "subjectScheme": "Fields of Science and Technology (FOS)", }, {"subject": "No code here", "subjectScheme": "arXiv"}, {"subject": "Quantum Physics (quant-ph)", "subjectScheme": "arXiv"}, ] } meta = arxiv_metadata._parse_record("2001.00001", attributes) assert meta.categories == ["math-ph", "quant-ph"] assert meta.primary_category == "math-ph" def test_only_version_of_dois_are_transcribed_and_several_are_joined(): attributes = { "relatedIdentifiers": [ { "relationType": "IsVersionOf", "relatedIdentifierType": "DOI", "relatedIdentifier": "10.1/a", }, { "relationType": "IsVersionOf", "relatedIdentifierType": "URL", "relatedIdentifier": "https://x", }, { "relationType": "IsCitedBy", "relatedIdentifierType": "DOI", "relatedIdentifier": "10.1/c", }, { "relationType": "IsVersionOf", "relatedIdentifierType": "DOI", "relatedIdentifier": "10.1/B", }, ] } assert arxiv_metadata._parse_record("2001.00001", attributes).doi == "10.1/a 10.1/B" def test_a_record_without_an_abstract_description_has_a_null_abstract(): attributes = { "descriptions": [{"description": "5 pages", "descriptionType": "Other"}] } assert arxiv_metadata._parse_record("2001.00001", attributes).abstract is None # --- lookup: failures, each with its own cause ------------------------------ @pytest.mark.parametrize( ("outcome", "cause"), [ ( _http_error(404), "DataCite has no record for 10.48550/arXiv.2606.09995 (HTTP 404)", ), (_http_error(429), "DataCite rate-limited the request (HTTP 429)"), (_http_error(503), "DataCite server error (HTTP 503)"), (_http_error(403), "HTTP 403"), ( urllib.error.URLError("Name or service not known"), "URLError: Name or service not known", ), (OSError("connection reset"), "OSError: connection reset"), (b'{"data": {"attributes": null}}', "DataCite response has no data.attributes"), (b"[]", "DataCite response has no data.attributes"), (b'{"errors": []}', "DataCite response has no data.attributes"), ], ids=[ "404", "429", "5xx", "other-http", "url-error", "os-error", "null-attributes", "top-level-list", "no-data", ], ) def test_each_failure_is_unavailable_with_its_own_cause(transport, outcome, cause): transport(outcome) result = fetch_metadata("2606.09995") assert result.status == METADATA_UNAVAILABLE assert result.metadata is None assert result.failure_cause.endswith(f"DataCite: {cause}") @pytest.mark.parametrize("body", [b"{not json", b"\xff\xfe\xfa"], ids=["json", "utf-8"]) def test_an_unparseable_body_is_unavailable(transport, body): transport(body) result = fetch_metadata("2606.09995") assert result.status == METADATA_UNAVAILABLE assert "DataCite: DataCite returned malformed JSON: " in result.failure_cause def test_an_unanticipated_exception_in_the_lookup_is_still_unavailable( transport, monkeypatch ): transport(_record("2409.03108")) def boom(*_args): raise RuntimeError("boom") monkeypatch.setattr(arxiv_metadata, "_parse_record", boom) result = fetch_metadata("2409.03108") assert result.status == METADATA_UNAVAILABLE # Each leg is guarded, so an exception neither of them anticipated is still # reported alongside what the other source said. assert result.failure_cause.endswith("DataCite: RuntimeError: boom") def test_surrogates_in_codes_and_dois_are_dropped_before_they_reach_a_handoff(tmp_path): # A JSON string can carry a lone surrogate through a \u escape, which no # UTF-8 file can hold. attributes = arxiv_metadata.json.loads( r""" { "subjects": [{"subject": "Broken (\ud800)", "subjectScheme": "arXiv"}], "relatedIdentifiers": [ { "relationType": "IsVersionOf", "relatedIdentifierType": "DOI", "relatedIdentifier": "10.1/a\udfff" } ] } """ ) meta = arxiv_metadata._parse_record("2001.00001", attributes) assert meta.categories == [] assert meta.doi == "10.1/a" path = tmp_path / "handoff.json" fetch = MetadataFetch(METADATA_OK, metadata=meta) write_metadata_handoff(path, "2001.00001", fetch) assert read_metadata_handoff(path, "2001.00001") == fetch def test_a_worker_that_cannot_start_is_unavailable(monkeypatch): class Unstartable: def __init__(self, *args, **kwargs): pass def start(self): raise RuntimeError("can't start new thread") monkeypatch.setattr(arxiv_metadata.threading, "Thread", Unstartable) result = fetch_metadata("2409.03108") assert result.status == METADATA_UNAVAILABLE assert result.error == ( "arXiv: RuntimeError: can't start new thread; " "DataCite: RuntimeError: can't start new thread" ) # The test ends the worker with an uncaught SystemExit on purpose, which pytest # reports as an unhandled thread exception. @pytest.mark.filterwarnings("ignore::pytest.PytestUnhandledThreadExceptionWarning") def test_a_worker_that_ends_without_a_result_is_not_reported_as_late(monkeypatch): def exit_thread(*_args): raise SystemExit monkeypatch.setattr(arxiv_metadata, "_lookup_arxiv", exit_thread) monkeypatch.setattr(arxiv_metadata, "_lookup_datacite", exit_thread) result = fetch_metadata("2409.03108", deadline=5) assert result.status == METADATA_UNAVAILABLE assert result.error == ( "arXiv: the arXiv lookup ended without a result; " "DataCite: the DataCite lookup ended without a result" ) # --- lookup: the deadline ---------------------------------------------------- def test_the_deadline_bounds_how_long_the_call_waits(transport): def slow(): time.sleep(2) return io.BytesIO(_record("2409.03108")) transport(slow) started = time.monotonic() result = fetch_metadata("2409.03108", deadline=0.5) elapsed = time.monotonic() - started assert result.status == METADATA_UNAVAILABLE assert elapsed < 1.5 # Running out of time does not cost the caller what each source said. assert result.failure_cause.startswith("arXiv: ") assert "DataCite: the DataCite lookup exceeded" in result.failure_cause def test_the_process_exits_without_waiting_for_an_abandoned_lookup(): # A worker the interpreter joins at exit would hold the process until the # request ended, which is what a daemon thread avoids. program = textwrap.dedent( """ import time, urllib.request import arxiv_doc_builder.arxiv_metadata as m def hang(url, timeout=None): time.sleep(30) urllib.request.urlopen = hang print(m.fetch_metadata("2409.03108", deadline=0.2).status) """ ) started = time.monotonic() result = subprocess.run( [sys.executable, "-c", program], capture_output=True, text=True, timeout=60 ) elapsed = time.monotonic() - started assert result.returncode == 0, result.stderr assert result.stdout.strip() == METADATA_UNAVAILABLE assert elapsed < 10 def test_an_arxiv_leg_that_leaves_no_time_does_not_ask_datacite(transport, monkeypatch): # The arXiv leg may spend only half the deadline, so this takes a leg that # overruns its bound, as a stalled scheduler can make it; a stub clock # stands in for that. What arXiv said must still reach the caller. asked = transport(OSError("offline"), arxiv=OSError("offline")) clock = iter([0.0, 10.0]) monkeypatch.setattr( arxiv_metadata, "time", types.SimpleNamespace(monotonic=lambda: next(clock)) ) result = fetch_metadata("2409.03108", deadline=5) assert result.failure_cause == ( "arXiv: OSError: offline; no time left to ask DataCite" ) assert len(asked) == 1 @pytest.mark.parametrize( "deadline", [0, -1.0, math.nan, math.inf, threading.TIMEOUT_MAX * 2, 10**400], ids=["zero", "negative", "nan", "inf", "above-timeout-max", "huge-int"], ) def test_a_deadline_that_cannot_bound_the_lookup_is_rejected(deadline, monkeypatch): # A finite value above threading.TIMEOUT_MAX would otherwise start the # request and then make worker.join raise OverflowError. started: list[object] = [] monkeypatch.setattr(arxiv_metadata, "_lookup", lambda *a: started.append(a)) with pytest.raises(ValueError): fetch_metadata("2409.03108", deadline=deadline) assert started == [] def test_the_largest_accepted_deadline_is_timeout_max(transport): transport(OSError("offline")) result = fetch_metadata("2409.03108", deadline=threading.TIMEOUT_MAX) assert result.failure_cause.endswith("DataCite: OSError: offline") def test_each_leg_is_given_its_share_of_the_deadline_as_a_socket_timeout(monkeypatch): # A request that stalls outright ends on its own only because urlopen # receives a timeout. The arXiv leg gets a share of the deadline so a slow # failure there cannot spend what the DataCite leg needs. timeouts: list[float] = [] def fake_urlopen(url, timeout=None): assert timeout is not None, "every leg passes its budget as the timeout" timeouts.append(timeout) raise OSError("offline") monkeypatch.setattr(arxiv_metadata.urllib.request, "urlopen", fake_urlopen) fetch_metadata("2409.03108", deadline=0.75) assert timeouts[0] == 0.75 * arxiv_metadata._ARXIV_DEADLINE_SHARE # The DataCite leg gets the time the arXiv leg left, measured rather than # assumed, so a fast failure hands it nearly the whole deadline and it never # gets more than that. assert 0 < timeouts[1] <= 0.75 assert len(timeouts) == 2 @pytest.mark.parametrize( "deadline", [Decimal("5"), Fraction(5), True, "5", None], ids=["decimal", "fraction", "bool", "str", "none"], ) def test_a_deadline_that_is_not_an_int_or_float_is_rejected_before_the_lookup( deadline, monkeypatch ): # Decimal and Fraction pass the range check, so without the type check the # request would start and worker.join would raise afterwards. started: list[object] = [] monkeypatch.setattr(arxiv_metadata, "_lookup", lambda *a: started.append(a)) with pytest.raises(TypeError): fetch_metadata("2409.03108", deadline=deadline) assert started == [] def test_the_lookup_takes_exactly_one_positional_argument(): # The suite's one-argument stubs keep it off the network; a second # positional parameter would make them raise instead. import inspect positional = [ parameter for parameter in inspect.signature(fetch_metadata).parameters.values() if parameter.kind in (parameter.POSITIONAL_ONLY, parameter.POSITIONAL_OR_KEYWORD) ] assert [parameter.name for parameter in positional] == ["arxiv_id"] # --- handoff between processes ----------------------------------------------- @pytest.mark.parametrize( "fetch", [ MetadataFetch(METADATA_OK, metadata=_FULL), MetadataFetch( METADATA_OK, metadata=ArxivMetadata(source=arxiv_metadata.METADATA_SOURCE_ARXIV), ), MetadataFetch( METADATA_OK, metadata=ArxivMetadata( title='q"uote \\ back\nline é \x85', authors=[""], source=arxiv_metadata.METADATA_SOURCE_ARXIV, ), ), MetadataFetch( METADATA_OK, metadata=ArxivMetadata( title=" ", categories=["\ud800"], doi="10.1/\udfff", source=arxiv_metadata.METADATA_SOURCE_DATACITE, ), ), MetadataFetch(METADATA_UNAVAILABLE, error="URLError: timed out"), ], ids=[ "full", "empty-record", "string-edges", "blank-and-unencodable", "unavailable", ], ) def test_a_handoff_reads_back_equal_to_what_was_written(tmp_path, fetch): path = tmp_path / "handoff.json" write_metadata_handoff(path, "2606.09995", fetch) assert read_metadata_handoff(path, "2606.09995") == fetch def test_the_writer_refuses_what_the_reader_would_reject(tmp_path): # Refused in the parent, rather than read back as unavailable in each child. path = tmp_path / "handoff.json" fetch = MetadataFetch(METADATA_OK, metadata=ArxivMetadata(title="T")) with pytest.raises(ValueError): write_metadata_handoff(path, "2606.09995", fetch) assert not path.exists() def test_a_record_naming_a_source_outside_the_vocabulary_is_refused(): # fetch_paper keeps a recorded revision over the record's only when the # fallback answered, so a record naming anything else would decide that # rule by a value no lookup produces. with pytest.raises(ValueError): ArxivMetadata(source="bogus") def test_a_handoff_path_given_as_a_string_is_read(tmp_path): path = tmp_path / "handoff.json" fetch = MetadataFetch(METADATA_OK, metadata=_FULL) write_metadata_handoff(path, "2606.09995", fetch) assert read_metadata_handoff(str(path), "2606.09995") == fetch # type: ignore[arg-type] _GOOD_METADATA = { "title": "T", "authors": ["A"], "version": None, "published": None, "primary_category": None, "categories": [], "doi": None, "journal": None, "abstract": None, "source": arxiv_metadata.METADATA_SOURCE_ARXIV, } def _handoff(**overrides) -> dict: fetch = {"status": METADATA_OK, "metadata": dict(_GOOD_METADATA), "error": None} fetch.update(overrides.pop("fetch", {})) payload = {"arxiv_id": "2606.09995", "fetch": fetch} payload.update(overrides) return payload def _with_metadata(**fields) -> dict: return _handoff(fetch={"metadata": {**_GOOD_METADATA, **fields}}) @pytest.mark.parametrize( "content", [ "{not json", "[]", _handoff(extra=1), _handoff(arxiv_id="2606.09995v1"), _handoff(fetch={"status": 5}), _handoff(fetch={"status": METADATA_UNAVAILABLE, "metadata": None, "error": 5}), _handoff( fetch={"status": METADATA_UNAVAILABLE, "metadata": None, "error": " "} ), _handoff(fetch={"status": METADATA_OK, "metadata": None}), _handoff(fetch={"metadata": ["T"]}), { "arxiv_id": "2606.09995", "fetch": {"status": METADATA_OK, "metadata": _GOOD_METADATA}, }, _with_metadata(authors="A"), _with_metadata(authors=["A", 1]), _with_metadata(title=7), _with_metadata(title=True), _with_metadata(source="bogus"), _with_metadata(source=None), { **_handoff(), "fetch": { **_handoff()["fetch"], "metadata": {**_GOOD_METADATA, "bogus": None}, }, }, ], ids=[ "not-json", "not-an-object", "extra-top-level-key", "other-id", "status-not-a-string", "error-not-a-string", "blank-error", "ok-without-record", "metadata-not-an-object", "missing-error-key", "authors-a-string", "authors-with-a-number", "title-a-number", "title-a-bool", "source-outside-the-vocabulary", "ok-without-a-source", "unknown-metadata-key", ], ) def test_an_untrustworthy_handoff_is_unavailable_with_a_cause(tmp_path, content): # Two cases are the source (``source-outside-the-vocabulary`` and # ``ok-without-a-source``): ``fetch_paper`` keeps a recorded # revision over the record's only when the fallback answered, so a handoff # whose source is missing or outside the vocabulary would decide that rule # by a value no lookup produces. path = tmp_path / "handoff.json" path.write_text( content if isinstance(content, str) else json.dumps(content), encoding="utf-8" ) result = read_metadata_handoff(path, "2606.09995") assert result.status == METADATA_UNAVAILABLE assert result.failure_cause.startswith("metadata handoff unreadable: ") @pytest.mark.parametrize("kind", ["missing", "directory"]) def test_a_handoff_path_that_is_not_a_readable_file_is_unavailable(tmp_path, kind): path = tmp_path / "handoff.json" if kind == "directory": path.mkdir() result = read_metadata_handoff(path, "2606.09995") assert result.status == METADATA_UNAVAILABLE assert result.failure_cause.startswith("metadata handoff unreadable: ") # --- fetch outcome: status, cause, and what each one licenses --------------- @pytest.mark.parametrize( "kwargs", [ {"status": "bogus"}, {"status": METADATA_NOT_REQUESTED, "error": "e"}, {"status": METADATA_OK}, {"status": METADATA_OK, "metadata": ArxivMetadata(), "error": "e"}, {"status": METADATA_UNAVAILABLE}, {"status": METADATA_UNAVAILABLE, "error": ""}, {"status": METADATA_UNAVAILABLE, "error": " "}, {"status": METADATA_UNAVAILABLE, "metadata": ArxivMetadata(), "error": "e"}, ], ids=[ "unknown-token", "not-a-fetch-outcome", "ok-without-record", "ok-with-error", "failed-without-cause", "failed-with-empty-cause", "failed-with-blank-cause", "failed-with-record", ], ) def test_incoherent_fetch_outcomes_are_rejected(kwargs): with pytest.raises(ValueError): MetadataFetch(**kwargs) def test_warning_names_the_paper_the_cause_and_what_the_document_records(): text = format_unavailable_warning("2606.09995", cause="OSError: connection reset") assert "2606.09995" in text assert "OSError: connection reset" in text assert "unknown" in text assert METADATA_UNAVAILABLE in text def test_the_warning_names_no_single_source_as_the_one_that_failed(): # Two sources are asked, and the cause already says what each of them # said. A message naming one of them would contradict the cause beside it # whenever the other was the one that answered. text = format_unavailable_warning("2606.09995", cause="arXiv: x; DataCite: y") assert "arXiv record" not in text assert "DataCite record" not in text # --- metadata_status in the frontmatter ------------------------------------ @pytest.mark.parametrize( ("status", "arxiv_id"), [ (METADATA_OK, "2606.09995"), (METADATA_UNAVAILABLE, "2606.09995"), (METADATA_NOT_REQUESTED, None), ], ids=list(METADATA_STATUSES), ) def test_every_status_token_round_trips_into_the_document(status, arxiv_id): # Each token needs the id state that can produce it. The guard rejects # every other pairing. Taking the ids from METADATA_STATUSES while # spelling the cases out means a token added there and not here leaves the # two lengths unequal, and pytest fails at collection. parsed = _parse(_fm(_FULL, arxiv_id=arxiv_id, metadata_status=status)) assert parsed["metadata_status"] == status assert set(parsed.keys()) == FRONTMATTER_KEYS def test_unknown_status_token_is_rejected_rather_than_rendered(): # A consumer branches on this key. An unrecognized token would render as # one more state to handle. with pytest.raises(ValueError): _fm(_FULL, metadata_status="degraded") def test_a_populated_record_does_not_imply_a_record_was_read(): # The PDF path builds a record from the PDF's own title when the lookup # failed. A present record is not evidence that either source was reached. # Only metadata_status answers that. from_pdf = ArxivMetadata(title="From The PDF", authors=["Jane Doe"]) parsed = _parse( _fm(from_pdf, source_type="pdf", metadata_status=METADATA_UNAVAILABLE) ) assert parsed["title"] == "From The PDF" assert parsed["metadata_status"] == METADATA_UNAVAILABLE assert parsed["doi"] is None def test_failure_cause_is_a_plain_string_for_every_failed_outcome(): # Callers formatting a warning need the cause unconditionally. Reading it # off the optional field instead would have each of them add a fallback # for a state __post_init__ already rules out. probe = MetadataFetch(METADATA_UNAVAILABLE, error="OSError: connection reset") assert probe.failure_cause == "OSError: connection reset" def test_failure_cause_refuses_an_outcome_that_did_not_fail(): with -
test_cli_contracts.py 10.3 KB
"""Cross-script CLI-level contracts. Promoted from review findings that revealed untested specifications: - Exit-code disjointness: exit code 2 is reserved for the "ambiguous main .tex" signal propagated through convert_latex → convert_paper so wrappers may retry with --tex-file. Any other failure — including validator rejection — must use a different code (currently 1) so a plain ID typo is never interpreted as a source-selection problem. - Default-path safe-normalization: when a CLI builds a default "<id>/..." path for a legacy ID like hep-th/9901001, the slash must be normalized away by safe_arxiv_id before becoming a Path component, otherwise the resulting tree (hep-th/9901001/...) does not match the fetch-side cache (hep-th_9901001/). """ import subprocess import sys from pathlib import Path import pytest from arxiv_doc_builder import convert_latex from conftest import PACKAGE_DIR def _run(args: list, cwd: Path) -> subprocess.CompletedProcess: """Run the current interpreter with ``args`` in ``cwd``, capturing its text.""" return subprocess.run( [sys.executable, *args], capture_output=True, text=True, cwd=str(cwd) ) @pytest.mark.parametrize( "script", ["fetch_paper.py", "convert_paper.py", "convert_latex.py"], ) def test_validator_failure_exits_1_leaving_2_reserved(script, tmp_path): # "2506.1376" is a structurally-valid-but-non-canonical new-style ID # (4-digit sequence on a post-2015 paper). Every CLI entry point # routes this through validate_arxiv_id at argparse time and must # exit 1, never 2 — exit 2 belongs to the ambiguous-main-tex channel. result = _run([str(PACKAGE_DIR / script), "2506.1376"], tmp_path) assert result.returncode == 1, ( f"{script}: validator failure exited {result.returncode}, " f"expected 1 (exit 2 is reserved for ambiguous main .tex).\n" f"stderr: {result.stderr}" ) assert "Error:" in result.stderr def test_convert_latex_default_path_normalizes_slash_for_legacy_id(tmp_path): # Legacy IDs contain a slash. Running convert_latex with no # --source-dir must build the default path via safe_arxiv_id so # the directory name has an underscore, not a slash. We verify by # observing the "source directory not found" diagnostic, which # echoes the constructed path. result = _run([str(PACKAGE_DIR / "convert_latex.py"), "hep-th/9901001"], tmp_path) # Expected failure: the working directory holds no source tree. assert result.returncode != 0 combined = result.stdout + result.stderr # The whole default, so a directory put in front of it shows up too. assert f"not found: {Path('hep-th_9901001') / 'source'}\n" in combined, combined assert "hep-th_9901001" in combined, ( "Expected the safe-normalized legacy-ID directory in the path " "message, so that convert_latex matches the fetch-side cache.\n" f"stdout: {result.stdout}\nstderr: {result.stderr}" ) # The unnormalized form would only appear if a slash leaked into a # Path component. Assert its absence to pin down the contract. assert "hep-th/9901001" not in combined, ( "Unnormalized legacy ID leaked into a path component; " "safe_arxiv_id must run before Path construction.\n" f"stdout: {result.stdout}\nstderr: {result.stderr}" ) def test_convert_latex_defaults_are_the_paper_directory_under_cwd( monkeypatch, tmp_path ): # Observed without pandoc: main() checks for it, converts, post-processes # and copies figures, and each of those is replaced. The last two record # the paths main() builds when neither --source-dir nor --output is given. monkeypatch.chdir(tmp_path) source = Path("hep-th_9901001") / "source" source.mkdir(parents=True) (source / "main.tex").write_text("\\documentclass{article}\n", encoding="utf-8") received = {} monkeypatch.setattr( convert_latex.subprocess, "run", lambda args, **kwargs: subprocess.CompletedProcess(args, 0), ) monkeypatch.setattr(convert_latex, "convert_with_pandoc", lambda tex, out: True) monkeypatch.setattr( convert_latex, "post_process_markdown", lambda md_file, *args, **kwargs: received.update(markdown=md_file), ) monkeypatch.setattr( convert_latex, "copy_figures", lambda source_dir, output_dir: received.update(source=source_dir), ) monkeypatch.setattr(sys, "argv", ["convert_latex.py", "hep-th/9901001"]) convert_latex.main() assert received == { "markdown": Path("hep-th_9901001") / "hep-th_9901001.md", "source": source, } @pytest.mark.parametrize( ("module", "function", "subdirectory"), [ ("convert_pdf_with_vision", "convert_pdf_to_images", "images"), ("convert_pdf_split_columns", "convert_pdf_split_columns", "images_split"), ], ) @pytest.mark.parametrize( "pdf_location", ["sub/2409.03108.pdf", "2409.03108/pdf/2409.03108.pdf"] ) def test_image_script_default_directory_is_named_after_the_pdf_under_cwd( module, function, subdirectory, pdf_location, tmp_path ): # The PDF's name has two dots and it sits below the working directory, so a # default built from the PDF's directory, or from a stem cut at the first # dot, would print a different path. The second location is where # convert-paper saves the PDF, which makes the stem an existing directory. # Each script imports its converter by name, so the stand-in replaces the # name bound in the script module. pdf = tmp_path / pdf_location pdf.parent.mkdir(parents=True) pdf.write_bytes(b"%PDF-stub") program = f""" import sys sys.path.insert(0, {str(PACKAGE_DIR)!r}) import {module} def record(pdf_path, output_dir, *args, **kwargs): print(output_dir) {module}.{function} = record sys.argv = [{module + ".py"!r}, {pdf_location!r}] {module}.main() """ result = _run(["-c", program], tmp_path) assert result.returncode == 0, result.stderr assert result.stdout.strip() == str(Path("2409.03108") / subdirectory) @pytest.mark.parametrize( "script", ["convert_pdf_with_vision.py", "convert_pdf_split_columns.py"] ) @pytest.mark.parametrize("obstacle", ["the pdf itself", "a dangling symlink"]) def test_image_script_refuses_a_default_directory_under_a_non_directory( script, obstacle, tmp_path ): # A PDF with no extension is its own stem, so run from its directory the # default would be a directory under the PDF itself. A symlink to nothing # at the stem blocks the directory the same way while not existing as a # file. The script says so and names the option, before the converter is # reached. if obstacle == "the pdf itself": pdf_location = "9901001" else: pdf_location = "sub/9901001.pdf" (tmp_path / "sub").mkdir() (tmp_path / "9901001").symlink_to("missing") (tmp_path / pdf_location).write_bytes(b"%PDF-stub") result = _run([str(PACKAGE_DIR / script), pdf_location], tmp_path) combined = result.stdout + result.stderr assert result.returncode == 1, combined assert "-o DIR" in combined assert "Traceback" not in combined @pytest.mark.parametrize( "script", ["fetch_paper.py", "convert_latex.py", "convert_pdf_simple.py"], ) def test_the_metadata_handoff_option_is_not_advertised(script, tmp_path): # The option carries convert_paper's lookup between its own steps. It is not # an interface for users, so --help leaves it out. result = _run([str(PACKAGE_DIR / script), "--help"], tmp_path) assert result.returncode == 0, result.stderr assert "--metadata-handoff" not in result.stdout def test_argparse_accepts_a_metadata_handoff_path_without_exit_2(tmp_path): # The option converts its value with Path, which accepts any string argparse # hands to it. A converter that rejected a value would make argparse exit 2, # which convert_paper would report as an ambiguous main .tex. The run stops # at the empty source directory, before the handoff is read. source_dir = tmp_path / "source" source_dir.mkdir() result = _run( [ str(PACKAGE_DIR / "convert_latex.py"), "2409.03108", "--source-dir", str(source_dir), "--metadata-handoff", "not/a/real/handoff.json", ], tmp_path, ) # No main .tex in an empty source directory is a generic failure. assert result.returncode == 1, f"stdout: {result.stdout}\nstderr: {result.stderr}" def test_convert_pdf_simple_refuses_a_handoff_with_no_arxiv_id(tmp_path): # --arxiv-id is optional there, so this pairing is reachable from the # command line. Refused at the entry the user invoked, naming the option, # rather than as a ValueError out of the converter — and with exit 1, since # exit 2 is reserved. pdf = tmp_path / "p.pdf" pdf.write_bytes(b"%PDF-stub") result = _run( [ str(PACKAGE_DIR / "convert_pdf_simple.py"), str(pdf), "--metadata-handoff", str(tmp_path / "handoff.json"), ], tmp_path, ) combined = result.stdout + result.stderr assert result.returncode == 1, combined # On stderr, where the other scripts report their argument errors, so a # wrapper that reports a failed child from its stderr has the reason. assert "--metadata-handoff" in result.stderr assert "Traceback" not in combined def test_convert_pdf_simple_passes_the_handoff_path_through(tmp_path): # The PDF step runs under uv with its dependencies, so the forwarding is # observed by replacing the converter inside a plain child interpreter. pdf = tmp_path / "p.pdf" pdf.write_bytes(b"%PDF-stub") handoff = tmp_path / "handoff.json" program = f""" import sys sys.path.insert(0, {str(PACKAGE_DIR)!r}) import convert_pdf_simple def record(**kwargs): print(repr(kwargs["arxiv_id"]), repr(str(kwargs["metadata_handoff"]))) convert_pdf_simple.convert_pdf_to_markdown = record sys.argv = ["convert_pdf_simple.py", {str(pdf)!r}, "--arxiv-id", "2409.03108", "--metadata-handoff", {str(handoff)!r}] convert_pdf_simple.main() """ result = _run(["-c", program], tmp_path) assert result.returncode == 0, result.stderr assert result.stdout.strip() == f"'2409.03108' {str(handoff)!r}" -
test_convert_pandoc_bounds.py 9.3 KB
"""Contract tests for the pandoc runaway bounds in convert_with_pandoc. Contract: convert_with_pandoc must never hang on a runaway pandoc. It bounds execution by a wall-clock timeout and an RSS watchdog; on either trip it kills the child and returns False (it does not raise and does not block past the timeout). _process_rss_mb must parse `ps -o rss=` output defensively, returning None rather than raising when the output is unparseable or `ps` fails. These use fakes for subprocess.Popen / subprocess.run / time.monotonic so the control flow is exercised without depending on a real pandoc or on timing. """ import subprocess from pathlib import Path from typing import NoReturn import pytest from arxiv_doc_builder import convert_latex from arxiv_doc_builder.convert_latex import _process_rss_mb, convert_with_pandoc @pytest.fixture(autouse=True) def _which_resolves_fake(monkeypatch: pytest.MonkeyPatch) -> None: """Resolve any binary to a fake absolute path so the suite does not depend on the host actually having pandoc / ps installed (convert_with_pandoc and _process_rss_mb now go through shutil.which before launching a subprocess).""" monkeypatch.setattr(convert_latex.shutil, "which", lambda name: f"/usr/bin/{name}") class _FakeCompleted: """Stand-in for subprocess.run's CompletedProcess (stdout only).""" def __init__(self, stdout: str) -> None: self.stdout = stdout class _FakeProc: """Stand-in for the Popen handle convert_with_pandoc drives. ``finishes`` controls whether ``wait(timeout=...)`` returns or raises TimeoutExpired. ``kill()`` flips it to finished so the subsequent bounded ``wait()`` returns instead of hanging the test. """ def __init__(self, *, finishes: bool, returncode: int = 0) -> None: self._finishes = finishes self.returncode = returncode self.pid = 4321 self.killed = False def wait(self, timeout: float | None = None) -> int: if self._finishes: return self.returncode # The bounded-wait path is only ever driven with a real float timeout; # narrow it so TimeoutExpired (which requires a float) typechecks. assert timeout is not None raise subprocess.TimeoutExpired(cmd="pandoc", timeout=timeout) def kill(self) -> None: self.killed = True self._finishes = True def _patch_popen(monkeypatch: pytest.MonkeyPatch, proc: _FakeProc) -> None: monkeypatch.setattr(convert_latex.subprocess, "Popen", lambda *a, **k: proc) # ---- _process_rss_mb ---- def test_process_rss_mb_parses_kb_to_mb(monkeypatch: pytest.MonkeyPatch) -> None: monkeypatch.setattr( convert_latex.subprocess, "run", lambda *a, **k: _FakeCompleted("2048\n") ) assert _process_rss_mb(1234) == 2 # 2048 KB -> 2 MB def test_process_rss_mb_none_on_unparseable(monkeypatch: pytest.MonkeyPatch) -> None: monkeypatch.setattr( convert_latex.subprocess, "run", lambda *a, **k: _FakeCompleted("") ) assert _process_rss_mb(1234) is None def test_process_rss_mb_none_on_ps_failure(monkeypatch: pytest.MonkeyPatch) -> None: def boom(*a: object, **k: object) -> NoReturn: raise OSError("ps unavailable") monkeypatch.setattr(convert_latex.subprocess, "run", boom) assert _process_rss_mb(1234) is None # ---- convert_with_pandoc bounds ---- def test_timeout_kills_and_returns_false(monkeypatch: pytest.MonkeyPatch) -> None: proc = _FakeProc(finishes=False) _patch_popen(monkeypatch, proc) # start = 0.0, then the loop's clock reads past the timeout on first check. times = iter([0.0, 1000.0, 1000.0]) monkeypatch.setattr(convert_latex.time, "monotonic", lambda: next(times)) # RSS stays low so only the timeout can trip. monkeypatch.setattr(convert_latex, "_process_rss_mb", lambda pid: 1) ok = convert_with_pandoc(Path("x.tex"), Path("out.md"), timeout=5, rss_cap_mb=8192) assert ok is False assert proc.killed def test_rss_watchdog_kills_and_returns_false(monkeypatch: pytest.MonkeyPatch) -> None: proc = _FakeProc(finishes=False) _patch_popen(monkeypatch, proc) # Clock never advances past the (large) timeout, so only RSS can trip. monkeypatch.setattr(convert_latex.time, "monotonic", lambda: 0.0) monkeypatch.setattr(convert_latex, "_process_rss_mb", lambda pid: 9000) ok = convert_with_pandoc( Path("x.tex"), Path("out.md"), timeout=300, rss_cap_mb=8192 ) assert ok is False assert proc.killed def test_returns_false_when_pandoc_missing(monkeypatch: pytest.MonkeyPatch) -> None: # Fail closed (not hang, not crash) when pandoc cannot be resolved on PATH. monkeypatch.setattr(convert_latex.shutil, "which", lambda _name: None) def _no_popen(*_a: object, **_k: object) -> NoReturn: raise AssertionError("Popen must not run when pandoc is unresolved") monkeypatch.setattr(convert_latex.subprocess, "Popen", _no_popen) assert convert_with_pandoc(Path("x.tex"), Path("out.md")) is False def test_returns_false_when_pandoc_path_relative( monkeypatch: pytest.MonkeyPatch, ) -> None: # A relative which() result (relative PATH entry) must be rejected, since it # would re-resolve against the untrusted cwd. monkeypatch.setattr(convert_latex.shutil, "which", lambda _name: "pandoc") def _no_popen(*_a: object, **_k: object) -> NoReturn: raise AssertionError("Popen must not run for a relative pandoc path") monkeypatch.setattr(convert_latex.subprocess, "Popen", _no_popen) assert convert_with_pandoc(Path("x.tex"), Path("out.md")) is False def test_process_rss_mb_none_when_ps_missing(monkeypatch: pytest.MonkeyPatch) -> None: monkeypatch.setattr(convert_latex.shutil, "which", lambda _name: None) def _no_run(*_a: object, **_k: object) -> NoReturn: raise AssertionError("subprocess.run must not run when ps is unresolved") monkeypatch.setattr(convert_latex.subprocess, "run", _no_run) assert _process_rss_mb(1234) is None def test_process_rss_mb_none_when_ps_path_relative( monkeypatch: pytest.MonkeyPatch, ) -> None: monkeypatch.setattr(convert_latex.shutil, "which", lambda _name: "ps") # The relative path must be rejected BEFORE any exec, so run must not fire. def _no_run(*_a: object, **_k: object) -> NoReturn: raise AssertionError("subprocess.run must not run for a relative ps path") monkeypatch.setattr(convert_latex.subprocess, "run", _no_run) assert _process_rss_mb(1234) is None def test_success_returns_true_without_kill(monkeypatch: pytest.MonkeyPatch) -> None: proc = _FakeProc(finishes=True, returncode=0) _patch_popen(monkeypatch, proc) monkeypatch.setattr(convert_latex.time, "monotonic", lambda: 0.0) ok = convert_with_pandoc(Path("x.tex"), Path("out.md")) assert ok is True assert not proc.killed def test_nonzero_exit_returns_false_without_kill( monkeypatch: pytest.MonkeyPatch, ) -> None: proc = _FakeProc(finishes=True, returncode=1) _patch_popen(monkeypatch, proc) monkeypatch.setattr(convert_latex.time, "monotonic", lambda: 0.0) ok = convert_with_pandoc(Path("x.tex"), Path("out.md")) assert ok is False assert not proc.killed def test_post_kill_wait_timeout_does_not_rehang( monkeypatch: pytest.MonkeyPatch, ) -> None: # A child that does not reap within the post-kill grace window must not # re-hang convert_with_pandoc: the bounded wait swallows TimeoutExpired. class _StuckProc(_FakeProc): def kill(self) -> None: self.killed = True # but leave _finishes False so wait() keeps raising proc = _StuckProc(finishes=False) _patch_popen(monkeypatch, proc) times = iter([0.0, 1000.0, 1000.0]) monkeypatch.setattr(convert_latex.time, "monotonic", lambda: next(times)) monkeypatch.setattr(convert_latex, "_process_rss_mb", lambda pid: 1) ok = convert_with_pandoc(Path("x.tex"), Path("out.md"), timeout=5) assert ok is False assert proc.killed # ---- _int_env (env-override parsing) ---- def test_int_env_uses_default_when_unset(monkeypatch: pytest.MonkeyPatch) -> None: monkeypatch.delenv("ARXIV_TEST_INT", raising=False) assert convert_latex._int_env("ARXIV_TEST_INT", 7) == 7 def test_int_env_parses_valid_value(monkeypatch: pytest.MonkeyPatch) -> None: monkeypatch.setenv("ARXIV_TEST_INT", "42") assert convert_latex._int_env("ARXIV_TEST_INT", 7) == 42 def test_int_env_falls_back_on_non_numeric( monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str] ) -> None: monkeypatch.setenv("ARXIV_TEST_INT", "3m") assert convert_latex._int_env("ARXIV_TEST_INT", 7) == 7 # A malformed override warns on stderr rather than raising (would crash import). assert "Ignoring non-integer" in capsys.readouterr().err def test_int_env_falls_back_on_zero( monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str] ) -> None: # 0 is syntactically valid but a zero timeout / RSS cap would trip instantly. monkeypatch.setenv("ARXIV_TEST_INT", "0") assert convert_latex._int_env("ARXIV_TEST_INT", 7) == 7 assert "Ignoring non-positive" in capsys.readouterr().err def test_int_env_falls_back_on_negative( monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str] ) -> None: monkeypatch.setenv("ARXIV_TEST_INT", "-5") assert convert_latex._int_env("ARXIV_TEST_INT", 7) == 7 assert "Ignoring non-positive" in capsys.readouterr().err -
test_convert_paper_handoff.py 2.1 KB
"""convert_paper's single metadata lookup, and how it reaches each step. The steps run as child processes, so the lookup is handed over in a file. These drive ``main()`` with the lookup and the child launcher replaced, and record what each child would have received. """ from pathlib import Path import pytest from conftest import LAUNCH_ARXIV_ID, LAUNCH_LOOKUP def _seed_latex(output_dir: Path) -> None: source = output_dir / LAUNCH_ARXIV_ID / "source" source.mkdir(parents=True) (source / "main.tex").write_text("\\documentclass{article}\n", encoding="utf-8") def _seed_pdf(output_dir: Path) -> None: pdf_dir = output_dir / LAUNCH_ARXIV_ID / "pdf" pdf_dir.mkdir(parents=True) (pdf_dir / f"{LAUNCH_ARXIV_ID}.pdf").write_bytes(b"%PDF-stub") @pytest.mark.parametrize( ("seed", "converter"), [(_seed_latex, "convert_latex.py"), (_seed_pdf, "convert_pdf_simple.py")], ids=["latex", "pdf"], ) def test_one_lookup_reaches_every_step_through_one_file(launch, seed, converter): seed(launch.output_dir) launch.run() assert launch.lookups == [LAUNCH_ARXIV_ID] assert [child.script for child in launch.children] == ["fetch_paper.py", converter] assert len({child.handoff for child in launch.children}) == 1 assert all(child.read == LAUNCH_LOOKUP for child in launch.children) assert not launch.children[0].handoff.exists() @pytest.mark.parametrize( ("exit_codes", "expected_exit"), [({"fetch_paper.py": 1}, 1), ({"convert_latex.py": 2}, 2)], ids=["fetch-fails", "ambiguous-tex"], ) def test_the_handoff_is_removed_when_a_step_ends_the_run( launch, exit_codes, expected_exit ): _seed_latex(launch.output_dir) with pytest.raises(SystemExit) as exit_info: launch.run(exit_codes=exit_codes) assert exit_info.value.code == expected_exit assert launch.children assert not launch.children[0].handoff.exists() def test_an_invalid_id_is_rejected_before_any_lookup(launch): with pytest.raises(SystemExit) as exit_info: launch.run("2506.1376") assert exit_info.value.code == 1 assert launch.lookups == [] assert launch.children == [] -
test_convert_paper_routing.py 5.1 KB
"""Contract tests for convert-paper routing decisions. Contract: an explicit --tex-file override must always route to the LaTeX conversion path, regardless of what the top-level source/*.tex auto-detection finds. Silently falling through to the PDF branch (or any other path) would make the explicit override unusable for source layouts where the real entrypoint lives in a subdirectory, or where source/ contains no top-level .tex files at all. Contract: the PDF fallback converts only the file the fetch step writes, ``{SAFE_ID}/pdf/{SAFE_ID}.pdf``. When that file is absent, a PDF elsewhere in the paper's directory is not converted, and the run exits 1 naming the path it looked for. The --tex-file test runs ``main()`` in the test process with only its own metadata lookup replaced. The steps it starts are real child processes, which read that lookup from the handoff file. The PDF-location test replaces the step launcher as well, since what it tests is ``main()``'s own check for the PDF. Nothing here reaches the network. """ import json import shutil import sys from types import SimpleNamespace import pytest from conftest import LAUNCH_ARXIV_ID from arxiv_doc_builder import convert_paper from arxiv_doc_builder.arxiv_metadata import ( METADATA_OK, METADATA_SOURCE_ARXIV, ArxivMetadata, MetadataFetch, ) _SENTINEL_TITLE = "Routing Sentinel Title" def test_tex_file_forces_latex_path_even_with_no_top_level_tex( tmp_path, monkeypatch, capfd ): # Build a source tree that has NO top-level .tex under source/ — # only a subdirectory entrypoint. Without the override guard, # auto-detection would fall through to the PDF branch and exit with # "PDF file not found", ignoring the user's explicit --tex-file. # Use a canonical-form placeholder so validate_arxiv_id accepts it; # the routing contract does not depend on the specific ID. arxiv_id = "2409.03108" version = f"{arxiv_id}v2" paper_dir = tmp_path / arxiv_id source_dir = paper_dir / "source" subdir = source_dir / "sub" subdir.mkdir(parents=True) tex_file = subdir / "main.tex" tex_file.write_text( "\\documentclass{article}\n\\begin{document}\nhi\n\\end{document}\n", encoding="utf-8", ) # Seed a non-empty PDF so fetch_pdf() reuses it without a network request. # It does so only while the drift record below matches the lookup's version. pdf_dir = paper_dir / "pdf" pdf_dir.mkdir(parents=True, exist_ok=True) (pdf_dir / f"{arxiv_id}.pdf").write_bytes(b"%PDF-stub") # Seed the drift record with the lookup's version. Without it any reported # version reads as drift, and the fetch step discards the seeded source and # downloads again. (paper_dir / ".arxiv-fetch.json").write_text( json.dumps({"version": version}), encoding="utf-8" ) lookup = MetadataFetch( METADATA_OK, metadata=ArxivMetadata( title=_SENTINEL_TITLE, version=version, source=METADATA_SOURCE_ARXIV ), ) monkeypatch.setattr(convert_paper, "fetch_metadata", lambda _id: lookup) monkeypatch.setattr( sys, "argv", [ "convert_paper.py", arxiv_id, "--output-dir", str(tmp_path), "--tex-file", str(tex_file), ], ) if shutil.which("pandoc"): convert_paper.main() out = capfd.readouterr().out # Only a LaTeX child that read the handoff can have written this title. document = (paper_dir / f"{arxiv_id}.md").read_text(encoding="utf-8") assert f'title: "{_SENTINEL_TITLE}"' in document else: # Without pandoc the LaTeX child exits before it reads metadata. Step 2 # is printed only after the fetch child exited 0, and that child reads # the handoff at its drift check, so a child looking the record up for # itself would still be observable here. with pytest.raises(SystemExit) as exit_info: convert_paper.main() assert exit_info.value.code == 1 out = capfd.readouterr().out assert "Step 2: Converting to Markdown" in out # Positive assertion: the explicit-override routing marker must be # printed, proving the LaTeX branch was entered. assert "Using explicit --tex-file" in out, out # Negative assertion: the PDF branch must NOT have been taken. Its # own marker string would indicate a routing regression. assert "falling back to naive PDF conversion" not in out, out def test_a_pdf_outside_pdf_dir_is_not_converted( launch: SimpleNamespace, capsys: pytest.CaptureFixture[str] ) -> None: paper_dir = launch.output_dir / LAUNCH_ARXIV_ID paper_dir.mkdir() # Directly in the paper's directory, where no step writes a PDF. (paper_dir / f"{LAUNCH_ARXIV_ID}.pdf").write_bytes(b"%PDF-stub") with pytest.raises(SystemExit) as exit_info: launch.run() assert exit_info.value.code == 1 # Only the fetch step was started: no converter ran on the stray file. assert [child.script for child in launch.children] == ["fetch_paper.py"] assert str(paper_dir / "pdf" / f"{LAUNCH_ARXIV_ID}.pdf") in capsys.readouterr().out -
test_dual_import_parity.py 4.6 KB
"""Both ways of loading the dual-import modules keep working. Some modules load two ways, as members of the installed package and as bare script files whose siblings sit beside them on ``sys.path``. Each tries the package-qualified import of its siblings first and falls back, on a ``ModuleNotFoundError`` naming the package, to a bare sibling import. Every other test here loads them the first way, which leaves the fallback branch unexercised. Two failures live in it. A name present in one branch and missing from the other raises ``NameError`` when that name is first used. A wrong sibling module name raises ``ImportError``, and only when the file runs as a script. The static check below catches the first and the executing check the second. """ import ast import subprocess import sys import pytest import arxiv_doc_builder from conftest import PACKAGE_DIR _PACKAGE = arxiv_doc_builder.__name__ def _import_pair(tree: ast.Module): """The (package-branch, fallback-branch) imports of a dual-import ``try``. Returns ``None`` for a module that carries no such pair, so the discovery below stays a scan instead of a hardcoded list. A module that grows the pattern later is covered without editing this file. """ for node in tree.body: if not isinstance(node, ast.Try): continue primary = [ n for n in node.body if isinstance(n, ast.ImportFrom) and (n.module or "").startswith(f"{_PACKAGE}.") ] if not primary: continue fallback = [ n for h in node.handlers for n in h.body if isinstance(n, ast.ImportFrom) ] return primary, fallback return None def _modules_with_import_pair(): found = [] for path in sorted(PACKAGE_DIR.glob("*.py")): pair = _import_pair(ast.parse(path.read_text(encoding="utf-8"))) if pair is not None: found.append(pytest.param(path, pair, id=path.stem)) return found _DUAL_IMPORT_MODULES = _modules_with_import_pair() def test_the_scan_found_the_dual_import_modules(): # A scan that silently matched nothing would make every parametrized test # below vacuous. assert _DUAL_IMPORT_MODULES @pytest.mark.parametrize(("path", "pair"), _DUAL_IMPORT_MODULES) def test_both_branches_import_the_same_names_from_the_same_modules(path, pair): primary, fallback = pair def names(imports): return {(alias.name, alias.asname) for node in imports for alias in node.names} def modules(imports, strip_package: bool): return { (node.module or "").removeprefix(f"{_PACKAGE}.") if strip_package else (node.module or "") for node in imports } assert names(primary) == names(fallback), ( f"{path.name}: the package and script import branches name different " f"symbols, so one configuration would fail on a name the other has" ) assert modules(primary, strip_package=True) == modules(fallback, False), ( f"{path.name}: the script branch imports from a different module than " f"the package branch" ) # Blocks ``arxiv_doc_builder`` in the child exactly as its genuine absence # would: the exception names the top-level package, which is what the modules' # own ``_exc.name`` guard re-raises on when it does not match. Blocking by any # route that names the submodule instead would trip that guard and prove # nothing about the fallback. Filtering ``sys.path`` would also work, but it # removes the directory the third-party dependencies live in whenever the # package is installed there, so the fallback would fail for the wrong reason. _BLOCK_PACKAGE = f''' import sys class _Block: def find_spec(self, name, path=None, target=None): if name == "{_PACKAGE}" or name.startswith("{_PACKAGE}."): raise ModuleNotFoundError(f"blocked {{name}}", name="{_PACKAGE}") return None sys.meta_path.insert(0, _Block()) sys.path.insert(0, {str(PACKAGE_DIR)!r}) ''' @pytest.mark.parametrize(("path", "pair"), _DUAL_IMPORT_MODULES) def test_the_fallback_branch_actually_resolves(path, pair, tmp_path): # Importing the module is the whole check: the fallback branch resolves or # it does not. Whether it names the same symbols as the package branch is # the static check's job above, not something an import can observe. program = _BLOCK_PACKAGE + f"import {path.stem}\n" result = subprocess.run( [sys.executable, "-c", program], capture_output=True, text=True, cwd=str(tmp_path), ) assert result.returncode == 0, ( f"{path.name} does not import with {_PACKAGE} absent:\n{result.stderr}" ) -
test_extract_title.py 1.7 KB
"""The LaTeX fallback title, used when arXiv supplies none. Whatever survives this extraction is written verbatim into a YAML title a reader sees. Leftover markup lands there as wrong text. """ import pytest from arxiv_doc_builder.convert_latex import extract_title_from_latex def _tex(tmp_path, body: str): path = tmp_path / "main.tex" path.write_text(body, encoding="utf-8") return path @pytest.mark.parametrize( ("source", "expected"), [ (r"\title{Plain Title}", "Plain Title"), # LaTeX's non-breaking space. Left alone it glues the two words it was # written to keep on one line. (r"\title{Koopman-Mori-Zwanzig~formalism}", "Koopman-Mori-Zwanzig formalism"), # Consecutive tildes collapse with the surrounding whitespace rather # than leaving a run of spaces behind. (r"\title{A~~B}", "A B"), (r"\title{Multi~word~title}", "Multi word title"), # A tilde adjacent to a command keeps the single separating space. (r"\title{Section~\ref{intro} overview}", "Section intro overview"), ], ) def test_spacing_markup_does_not_reach_the_title(tmp_path, source, expected): assert extract_title_from_latex(_tex(tmp_path, source)) == expected def test_the_tilde_accent_is_left_as_it_is(tmp_path): # ``\~`` is the tilde accent, not a spacing token, so the normalization # must not touch it. Its markup is not resolved here either. This pins # that the case stays exactly where it already stood. assert extract_title_from_latex(_tex(tmp_path, r"\title{A\~{n}o study}")) == ( r"A\~no study" ) def test_no_title_stays_none(tmp_path): assert extract_title_from_latex(_tex(tmp_path, "no title here\n")) is None -
test_fetch_paper_main.py 7.8 KB
"""When the fetch entry point warns that the drift record stood still. `test_version_drift.py` covers the parts. `main()` holds only their composition, and an inverted operand there passes every test of either part. These drive `main()` with the lookup and both downloads replaced. `test_without_output_dir_the_paper_directory_is_under_the_working_directory` uses the same fixture for a different question: where `main()` writes when `--output-dir` is left off. """ import inspect import sys import pytest from conftest import PROBE_ERROR, PROBE_VERSION, refuse_lookup, seed_cached_source from arxiv_doc_builder import fetch_paper from arxiv_doc_builder.arxiv_metadata import ( METADATA_OK, METADATA_SOURCE_DATACITE, ArxivMetadata, MetadataFetch, write_metadata_handoff, ) @pytest.fixture def run_main(monkeypatch, tmp_path): """Invoke ``main()`` with the lookup and both downloads replaced. The lookup outcome comes either from ``probe``, standing in for the lookup ``main()`` makes itself, or from a handoff file written with ``handoff``, in which case a lookup would fail the test. Each download reports ``has_source`` / ``has_pdf``, or runs ``download`` when one is given, so a caller can see what was asked for. ``output_dir=False`` leaves ``--output-dir`` off the command line and runs with ``tmp_path`` as the working directory. Returns the paper directory the run wrote into, so a caller can check whether the sidecar landed. """ def run( *, probe=None, handoff=None, has_source=True, has_pdf=True, download=None, output_dir=True, ): argv = ["fetch_paper.py", "2409.03108"] if output_dir: argv += ["--output-dir", str(tmp_path)] else: monkeypatch.chdir(tmp_path) if handoff is not None: handoff_file = tmp_path / "handoff.json" write_metadata_handoff(handoff_file, "2409.03108", handoff) argv += ["--metadata-handoff", str(handoff_file)] monkeypatch.setattr(fetch_paper, "_probe_metadata", refuse_lookup) else: monkeypatch.setattr(fetch_paper, "_probe_metadata", lambda _id: probe) monkeypatch.setattr(sys, "argv", argv) monkeypatch.setattr( fetch_paper, "fetch_source", download or (lambda *a, **k: has_source) ) monkeypatch.setattr( fetch_paper, "fetch_pdf", download or (lambda *a, **k: has_pdf) ) fetch_paper.main() return tmp_path / "2409.03108" return run def test_the_probe_takes_exactly_one_positional_argument(): # Every stand-in for it across the suite is a one-argument callable, and # those stand-ins are what keep this module off the network. A second # positional parameter would keep the production call working while each # of them raised. positional = [ parameter for parameter in inspect.signature( fetch_paper._probe_metadata ).parameters.values() if parameter.kind in (parameter.POSITIONAL_ONLY, parameter.POSITIONAL_OR_KEYWORD) ] assert [parameter.name for parameter in positional] == ["arxiv_id"] @pytest.mark.parametrize( ("probe_name", "downloaded"), [("probe_with_version", PROBE_VERSION), ("failed_probe", "2409.03108")], ) def test_the_downloads_name_the_revision_the_sidecar_records( run_main, request, probe_name, downloaded ): # While the fallback has yet to list a revision arXiv already serves, an # unversioned download would put that newer revision on disk under the # older version the sidecar records. probe = request.getfixturevalue(probe_name) ids: list[str] = [] def record(arxiv_id, *_args, **_kwargs): ids.append(arxiv_id) return True run_main(probe=probe, download=record) assert ids == [downloaded, downloaded] def test_a_record_of_a_later_revision_is_kept_while_the_lookup_lags( run_main, tmp_path, capsys ): # The fallback reports PROBE_VERSION while the sidecar already records a # later revision. Refreshing would delete the cached source, edits # included, and record the older revision over newer material. Only the # fallback's record can trail arXiv, which is why the probe names it. probe_with_version = MetadataFetch( METADATA_OK, metadata=ArxivMetadata(version=PROBE_VERSION, source=METADATA_SOURCE_DATACITE), ) paper_dir = tmp_path / "2409.03108" seed_cached_source(paper_dir) later = PROBE_VERSION.rsplit("v", 1)[0] + "v99" fetch_paper._write_cached_version(paper_dir, later) calls: list[tuple[str, bool]] = [] def record(arxiv_id, *_args, refresh, **_kwargs): calls.append((arxiv_id, refresh)) return True run_main(probe=probe_with_version, download=record) assert calls == [(later, False), (later, False)] assert fetch_paper._read_cached_version(paper_dir) == later # The document will name PROBE_VERSION while the body is `later`, so the # run says which it kept. err = capsys.readouterr().err assert PROBE_VERSION in err assert later in err def test_material_without_a_version_warns_and_writes_no_sidecar( run_main, capsys, failed_probe ): paper_dir = run_main(probe=failed_probe, has_source=True, has_pdf=False) assert not (paper_dir / fetch_paper._METADATA_FILE).exists() assert fetch_paper._METADATA_FILE in capsys.readouterr().err def test_material_with_a_version_records_it_and_stays_silent( run_main, capsys, probe_with_version ): # The other side of the composed condition: warning here would fire on # every ordinary successful run. paper_dir = run_main(probe=probe_with_version, has_source=True, has_pdf=True) assert fetch_paper._read_cached_version(paper_dir) == PROBE_VERSION assert capsys.readouterr().err == "" def test_a_run_that_obtained_nothing_fails_instead_of_warning( run_main, capsys, failed_probe ): # No material means no successful-looking result to qualify, so the run # exits non-zero on its own and the drift warning would only add noise. with pytest.raises(SystemExit) as exit_info: run_main(probe=failed_probe, has_source=False, has_pdf=False) assert exit_info.value.code == 1 assert capsys.readouterr().err == "" def test_a_read_record_without_a_version_warns_the_same_way( run_main, capsys, probe_without_version ): # The cell that separates branching on the version from branching on the # status. The lookup succeeded here, so a warning gated on `unavailable` # would stay silent while the sidecar still went unwritten. paper_dir = run_main(probe=probe_without_version, has_source=True, has_pdf=False) assert not (paper_dir / fetch_paper._METADATA_FILE).exists() assert fetch_paper._METADATA_FILE in capsys.readouterr().err def test_a_handed_over_record_drives_the_sidecar_without_a_lookup( run_main, capsys, probe_with_version ): paper_dir = run_main(handoff=probe_with_version, has_source=True, has_pdf=True) assert fetch_paper._read_cached_version(paper_dir) == PROBE_VERSION assert capsys.readouterr().err == "" def test_a_handed_over_failure_is_quoted_in_the_warning(run_main, capsys, failed_probe): paper_dir = run_main(handoff=failed_probe, has_source=True, has_pdf=False) assert not (paper_dir / fetch_paper._METADATA_FILE).exists() assert PROBE_ERROR in capsys.readouterr().err def test_without_output_dir_the_paper_directory_is_under_the_working_directory( run_main, tmp_path, probe_with_version ): # convert-paper's own default, so a fetch run by hand beside it writes into # the directory convert-paper reads. paper_dir = run_main(probe=probe_with_version, output_dir=False) assert fetch_paper._read_cached_version(paper_dir) == PROBE_VERSION assert [entry.name for entry in tmp_path.iterdir()] == ["2409.03108"] -
test_fetch_source_host.py 882 B
"""Which host the source download goes to. `fetch_source` must download from `export.arxiv.org`; the reason is stated where it builds `source_url`. Both hosts serve the same archive, so nothing else notices if the URL drifts to `arxiv.org`; only this pin does. """ import subprocess from arxiv_doc_builder import fetch_paper def test_source_is_downloaded_from_the_export_host(monkeypatch, tmp_path): calls: list[list[str]] = [] def fake_run(argv, **_kwargs): calls.append(argv) # Fail the download so fetch_source returns before extraction. return subprocess.CompletedProcess(argv, returncode=22) monkeypatch.setattr(fetch_paper.subprocess, "run", fake_run) assert not fetch_paper.fetch_source("2409.03108", tmp_path, "2409.03108") assert calls[0][0] == "curl" assert calls[0][-1] == "https://export.arxiv.org/src/2409.03108" -
test_find_main_tex.py 4.4 KB
"""Contract tests for main .tex discovery. Contract: find_main_tex must never silently guess between multiple \\documentclass files. Conventional names (main.tex, paper.tex, ms.tex, article.tex) take precedence; if none match and more than one \\documentclass file is present, it must raise AmbiguousMainTexError so the caller can fail explicitly and require --tex-file on re-run. """ import subprocess import sys from pathlib import Path import pytest from conftest import PACKAGE_DIR from arxiv_doc_builder.convert_latex import ( AmbiguousMainTexError, extract_title_from_latex, find_main_tex, ) def _write_tex(dir_: Path, name: str, with_documentclass: bool = True) -> Path: path = dir_ / name body = "\\documentclass{article}\n" if with_documentclass else "" body += "\\begin{document}\nhello\n\\end{document}\n" path.write_text(body, encoding="utf-8") return path def test_known_name_wins_over_ambiguity(tmp_path): # Even with multiple \documentclass files, a conventional main.tex # must short-circuit the ambiguity check. main = _write_tex(tmp_path, "main.tex") _write_tex(tmp_path, "supplement.tex") _write_tex(tmp_path, "appendix.tex") assert find_main_tex(tmp_path) == main def test_single_documentclass_returned(tmp_path): only = _write_tex(tmp_path, "foo.tex") # A file without \documentclass must not count as a candidate. _write_tex(tmp_path, "bar.tex", with_documentclass=False) assert find_main_tex(tmp_path) == only def test_multiple_documentclass_raises(tmp_path): a = _write_tex(tmp_path, "alpha.tex") b = _write_tex(tmp_path, "beta.tex") with pytest.raises(AmbiguousMainTexError) as exc_info: find_main_tex(tmp_path) # All candidates must be reported so the caller can surface them. assert set(exc_info.value.candidates) == {a, b} def test_no_tex_files_returns_none(tmp_path): assert find_main_tex(tmp_path) is None def test_cli_exits_2_on_ambiguity_with_candidates_in_stderr(tmp_path): # End-to-end CLI contract: invoking convert_latex.py against an # ambiguous source dir must exit 2 and report every candidate on stderr. # This guards against future refactors that swallow the exception. _write_tex(tmp_path, "alpha.tex") _write_tex(tmp_path, "beta.tex") script = PACKAGE_DIR / "convert_latex.py" result = subprocess.run( [ sys.executable, str(script), # Canonical placeholder so validate_arxiv_id accepts it; the # ambiguity contract is independent of the specific ID. "2409.03108", "--source-dir", str(tmp_path), "--output", str(tmp_path / "out.md"), ], capture_output=True, text=True, ) assert result.returncode == 2, ( f"expected exit 2, got {result.returncode}\n" f"stdout: {result.stdout}\nstderr: {result.stderr}" ) assert "alpha.tex" in result.stderr assert "beta.tex" in result.stderr assert "--tex-file" in result.stderr def test_extract_title_uses_selected_tex_not_arbitrary_sibling(tmp_path): # Contract: the fallback title must come from the file actually converted, # not an arbitrary .tex picked from the source directory. (tmp_path / "main.tex").write_text(r"\title{Correct Main Title}", encoding="utf-8") (tmp_path / "supplement.tex").write_text( r"\title{Wrong Supplement Title}", encoding="utf-8" ) assert extract_title_from_latex(tmp_path / "main.tex") == "Correct Main Title" assert ( extract_title_from_latex(tmp_path / "supplement.tex") == "Wrong Supplement Title" ) def test_extract_title_falls_back_to_sibling_when_selected_has_none(tmp_path): # The \title may sit in an included preamble; fall back to siblings rather # than reporting no title. (tmp_path / "main.tex").write_text(r"\documentclass{article}", encoding="utf-8") (tmp_path / "preamble.tex").write_text( r"\title{Title In Preamble}", encoding="utf-8" ) assert extract_title_from_latex(tmp_path / "main.tex") == "Title In Preamble" def test_extract_title_returns_none_when_absent(tmp_path): # No \title anywhere -> None, so the frontmatter title stays null rather than # a fabricated placeholder. (tmp_path / "main.tex").write_text(r"\documentclass{article}", encoding="utf-8") assert extract_title_from_latex(tmp_path / "main.tex") is None -
test_metadata_status_latex.py 4.1 KB
"""The LaTeX path's recorded status, and when it warns. ``test_arxiv_metadata.py`` covers the frontmatter producer. Handing it a token cannot catch a caller that always reports ``ok`` or never warns, which is why these drive the caller and read the status back out of the document it wrote. Kept apart from the PDF-path tests to stay importable without the optional PDF dependencies. Each test replaces the caller's ``fetch_metadata`` or hands it a lookup, so nothing here reaches the network. """ import subprocess import sys from conftest import PROBE_ERROR, refuse_lookup, status_of from arxiv_doc_builder import convert_latex from arxiv_doc_builder.arxiv_metadata import ( METADATA_OK, METADATA_UNAVAILABLE, write_metadata_handoff, ) import pytest @pytest.fixture def latex_inputs(tmp_path): """A converted-markdown file and the ``.tex`` its title falls back to.""" tex = tmp_path / "main.tex" tex.write_text(r"\title{Fallback~Title}" + "\n", encoding="utf-8") md = tmp_path / "2606.09995.md" md.write_text("body\n", encoding="utf-8") return md, tex def test_latex_path_records_unavailable_and_warns( patch_fetch, capsys, latex_inputs, failed_probe ): md, tex = latex_inputs patch_fetch(convert_latex, failed_probe) convert_latex.post_process_markdown(md, "2606.09995", tex) assert status_of(md) == METADATA_UNAVAILABLE err = capsys.readouterr().err assert "2606.09995" in err assert PROBE_ERROR in err # The title still comes from the LaTeX source, and reaches the document # with the non-breaking space normalized. assert 'title: "Fallback Title"' in md.read_text(encoding="utf-8") def test_latex_path_records_ok_and_stays_silent( patch_fetch, capsys, latex_inputs, probe_with_version ): # Without this, a path that warns unconditionally would pass the test above. md, tex = latex_inputs patch_fetch(convert_latex, probe_with_version) convert_latex.post_process_markdown(md, "2606.09995", tex) assert status_of(md) == METADATA_OK assert capsys.readouterr().err == "" @pytest.mark.parametrize( ("probe_name", "status"), [("probe_with_version", METADATA_OK), ("failed_probe", METADATA_UNAVAILABLE)], ) def test_latex_path_given_a_handoff_uses_it_without_looking_up( monkeypatch, request, capsys, tmp_path, latex_inputs, probe_name, status ): md, tex = latex_inputs handoff = tmp_path / "handoff.json" write_metadata_handoff(handoff, "2606.09995", request.getfixturevalue(probe_name)) monkeypatch.setattr(convert_latex, "fetch_metadata", refuse_lookup) convert_latex.post_process_markdown(md, "2606.09995", tex, metadata_handoff=handoff) assert status_of(md) == status # A handed-over failure still reaches the user with its cause. err = capsys.readouterr().err assert (PROBE_ERROR in err) == (status == METADATA_UNAVAILABLE) def test_latex_main_hands_its_handoff_path_to_post_processing(monkeypatch, tmp_path): # Observed without pandoc: conversion and figure copying are replaced, so # the only thing left to check is what main() passes on. source = tmp_path / "source" source.mkdir() (source / "main.tex").write_text("\\documentclass{article}\n", encoding="utf-8") handoff = tmp_path / "handoff.json" received = {} def record(md_file, arxiv_id, tex_file, **kwargs): received.update(kwargs, arxiv_id=arxiv_id) monkeypatch.setattr( convert_latex.subprocess, "run", lambda args, **kwargs: subprocess.CompletedProcess(args, 0), ) monkeypatch.setattr(convert_latex, "convert_with_pandoc", lambda tex, out: True) monkeypatch.setattr(convert_latex, "post_process_markdown", record) monkeypatch.setattr(convert_latex, "copy_figures", lambda *args: None) monkeypatch.setattr( sys, "argv", [ "convert_latex.py", "2606.09995", "--source-dir", str(source), "--output", str(tmp_path / "out.md"), "--metadata-handoff", str(handoff), ], ) convert_latex.main() assert received == {"arxiv_id": "2606.09995", "metadata_handoff": handoff} -
test_metadata_status_pdf.py 4.5 KB
"""The PDF path's recorded status, and when it warns. Counterpart to the LaTeX-path module, kept apart because building a fixture document here needs the optional PDF dependencies. Nothing reaches the network. Tests that supply an arXiv id replace the caller's ``fetch_metadata`` or hand it a lookup, and the rest look nothing up. """ from pathlib import Path import pytest from conftest import PROBE_ERROR, refuse_lookup, status_of from pypdf import PdfWriter from arxiv_doc_builder import pdf_converter_lib from arxiv_doc_builder.arxiv_metadata import ( METADATA_NOT_REQUESTED, METADATA_OK, METADATA_UNAVAILABLE, write_metadata_handoff, ) def _blank_pdf(path: Path, *, title: str, author: str) -> Path: writer = PdfWriter() writer.add_blank_page(width=612, height=792) writer.add_metadata({"/Title": title, "/Author": author}) with path.open("wb") as fh: writer.write(fh) return path @pytest.fixture def pdf_inputs(tmp_path): """A PDF carrying an embedded title and author, and the output path.""" pdf = _blank_pdf(tmp_path / "p.pdf", title="Embedded Title", author="Jane Doe") return pdf, tmp_path / "p.md" def test_pdf_path_records_unavailable_and_warns( patch_fetch, capsys, pdf_inputs, failed_probe ): pdf, out = pdf_inputs patch_fetch(pdf_converter_lib, failed_probe) pdf_converter_lib.convert_pdf_to_markdown(pdf, out, arxiv_id="2606.09995") assert status_of(out) == METADATA_UNAVAILABLE assert PROBE_ERROR in capsys.readouterr().err # Both fields the PDF can supply are filled from it, which is why neither # is evidence that arXiv was reached. The status is. text = out.read_text(encoding="utf-8") assert 'title: "Embedded Title"' in text assert 'authors: "Jane Doe"' in text assert "\ndoi:\n" in text def test_pdf_path_records_ok_when_the_record_was_read( patch_fetch, capsys, pdf_inputs, probe_with_version ): pdf, out = pdf_inputs patch_fetch(pdf_converter_lib, probe_with_version) pdf_converter_lib.convert_pdf_to_markdown(pdf, out, arxiv_id="2606.09995") assert status_of(out) == METADATA_OK assert capsys.readouterr().err == "" def test_pdf_path_records_not_requested_and_stays_silent(capsys, pdf_inputs): # The manual PDF scripts run this way on every invocation, so treating it # as degraded would make the ordinary case look like an outage. pdf, out = pdf_inputs pdf_converter_lib.convert_pdf_to_markdown(pdf, out, arxiv_id=None) assert status_of(out) == METADATA_NOT_REQUESTED assert capsys.readouterr().err == "" def test_an_empty_arxiv_id_is_treated_as_no_id_at_all(capsys, pdf_inputs): # argparse yields None for an omitted --arxiv-id, but a caller can pass an # empty string. Normalizing it at entry is what keeps the frontmatter # builder's id/status equivalence from rejecting the document: without it # the status says no id was supplied while the id argument says one was. pdf, out = pdf_inputs pdf_converter_lib.convert_pdf_to_markdown(pdf, out, arxiv_id="") assert status_of(out) == METADATA_NOT_REQUESTED assert "arxiv_id:\n" in out.read_text(encoding="utf-8") assert capsys.readouterr().err == "" @pytest.mark.parametrize( ("probe_name", "status"), [("probe_with_version", METADATA_OK), ("failed_probe", METADATA_UNAVAILABLE)], ) def test_pdf_path_given_a_handoff_uses_it_without_looking_up( monkeypatch, request, capsys, tmp_path, pdf_inputs, probe_name, status ): pdf, out = pdf_inputs handoff = tmp_path / "handoff.json" write_metadata_handoff(handoff, "2606.09995", request.getfixturevalue(probe_name)) monkeypatch.setattr(pdf_converter_lib, "fetch_metadata", refuse_lookup) pdf_converter_lib.convert_pdf_to_markdown( pdf, out, arxiv_id="2606.09995", metadata_handoff=handoff ) assert status_of(out) == status # A handed-over failure still reaches the user with its cause. err = capsys.readouterr().err assert (PROBE_ERROR in err) == (status == METADATA_UNAVAILABLE) @pytest.mark.parametrize("arxiv_id", [None, ""]) def test_a_handoff_without_an_arxiv_id_is_rejected_before_any_output( tmp_path, pdf_inputs, arxiv_id ): # A handoff describes one id's lookup, so a document with no id cannot use # one, and the refusal comes before anything is written. pdf, out = pdf_inputs with pytest.raises(ValueError): pdf_converter_lib.convert_pdf_to_markdown( pdf, out, arxiv_id=arxiv_id, metadata_handoff=tmp_path / "handoff.json" ) assert not out.exists() -
test_network_guard.py 2 KB
"""The conftest network guard refuses and records a child's request. Every other test relies on the guard being silent, so a guard that stopped loading would leave them all green. This drives it directly. """ import os import subprocess import sys import textwrap _URL = "https://api.datacite.org/dois/10.48550/arxiv.2409.03108" def test_a_child_request_is_refused_and_recorded(network_record): program = textwrap.dedent( f""" import urllib.request try: urllib.request.urlopen({_URL!r}, timeout=5) except OSError as exc: print("refused:", exc) """ ) result = subprocess.run( [sys.executable, "-c", program], capture_output=True, text=True ) assert result.returncode == 0, result.stderr assert result.stdout.startswith("refused: network access is blocked in tests") assert network_record.read_text(encoding="utf-8") == _URL + "\n" # Emptied so this test's own teardown check passes. network_record.write_text("", encoding="utf-8") def test_the_guard_chains_to_the_sitecustomize_it_displaces( tmp_path, monkeypatch, network_record ): # Python imports only the first sitecustomize on sys.path. Without the # chain, an environment shipping its own would lose it in every Python # process a test starts, silently and only under pytest. displaced = tmp_path / "displaced" displaced.mkdir() marker = tmp_path / "marker.txt" (displaced / "sitecustomize.py").write_text( f"open({str(marker)!r}, 'w', encoding='utf-8').write('loaded')\n", encoding="utf-8", ) # After the guard's own directory, which the fixture put first. monkeypatch.setenv( "PYTHONPATH", os.environ["PYTHONPATH"] + os.pathsep + str(displaced) ) result = subprocess.run( [sys.executable, "-c", "import urllib.request"], capture_output=True, text=True, ) assert result.returncode == 0, result.stderr assert marker.read_text(encoding="utf-8") == "loaded" -
test_packaging.py 1.6 KB
"""Guards for the dependency-declaration contract. The core install must stay dependency-free (the LaTeX happy path pulls nothing), and the heavy PDF stack must remain an opt-in extra. Encoding this as a test — rather than a comment in pyproject.toml — makes a regression (a stray runtime dep added to the core, or the pdf extra losing a member) fail CI loudly. """ import re import tomllib from conftest import SKILL_DIR # A PEP 508 dependency string starts with the distribution name, followed by # optional extras / version specifiers / environment markers. The leading run of # name characters is the name regardless of which specifier operator (>=, ==, ~=, # <, !=, …), marker, or extra follows, so match that rather than splitting on a # hand-picked subset of operators. _PEP508_NAME = re.compile(r"[A-Za-z0-9._-]+") _PYPROJECT = SKILL_DIR / "pyproject.toml" def _load() -> dict: with _PYPROJECT.open("rb") as fh: return tomllib.load(fh) def _dist_name(dep: str) -> str: """Return the distribution name of a PEP 508 dependency string.""" m = _PEP508_NAME.match(dep.strip()) assert m is not None, f"no PEP 508 name in {dep!r}" return m.group() def test_core_is_dependency_free(): assert _load()["project"]["dependencies"] == [] def test_pdf_extra_lists_the_heavy_stack(): extras = _load()["project"]["optional-dependencies"]["pdf"] # Compare on package names only, tolerant of any future version pins / markers. names = {_dist_name(dep) for dep in extras} assert names == {"pdfplumber", "pdf2image", "pypdf", "pillow"} -
test_pdf_converter_lib.py 4.2 KB
"""Tests for the pdfplumber-based converter library. These import the tier as a package member, which only resolves when the `pdf` extra is installed (the dev group pulls it in). The pure helpers (parse_page_ranges / clean_text / header / footer detection) need no PDF at all; extract_metadata and convert_pdf_to_markdown drive a pypdf-written PDF so the text-extraction path is exercised without poppler. """ from pathlib import Path import pytest from pypdf import PdfWriter from arxiv_doc_builder.pdf_converter_lib import ( clean_text, convert_pdf_to_markdown, extract_metadata, is_likely_footer, is_likely_header, parse_page_ranges, ) def _write_pdf( path: Path, *, title: str | None = None, author: str | None = None, pages: int = 1 ) -> Path: """Write a blank-page PDF with optional embedded metadata. Blank pages carry no text layer, so pdfplumber's extract_text returns None — enough to exercise the no-text branch and the metadata/frontmatter paths without needing a real typeset document or the poppler binary. """ writer = PdfWriter() for _ in range(pages): writer.add_blank_page(width=612, height=792) embedded = {} if title is not None: embedded["/Title"] = title if author is not None: embedded["/Author"] = author if embedded: writer.add_metadata(embedded) with path.open("wb") as fh: writer.write(fh) return path class TestParsePageRanges: def test_mixed_ranges_and_singles(self): assert parse_page_ranges("1-3,5,7-9") == {1, 2, 3, 5, 7, 8, 9} def test_whitespace_tolerated(self): assert parse_page_ranges(" 1 - 2 , 4 ") == {1, 2, 4} def test_single_page(self): assert parse_page_ranges("7") == {7} @pytest.mark.parametrize("empty", [None, ""]) def test_empty_is_empty_set(self, empty): assert parse_page_ranges(empty) == set() class TestCleanText: def test_collapses_whitespace(self): assert clean_text("a b\t c\n d") == "a b c d" def test_fixes_ligatures(self): assert clean_text("fi fl ff") == "fi fl ff" def test_empty_returns_empty(self): assert clean_text("") == "" class TestHeaderFooterDetection: def test_header_matches_journal_banner_case_insensitively(self): assert is_likely_header("physical review b") is True assert is_likely_header("VOLUME 12") is True def test_header_rejects_plain_text_and_empty(self): assert is_likely_header("Introduction") is False assert is_likely_header("") is False def test_footer_matches_page_number_only(self): assert is_likely_footer(" 42 ") is True assert is_likely_footer("12 and text") is False def test_footer_matches_copyright_markers(self): assert is_likely_footer("© 2024 The American Physical Society") is True assert is_likely_footer("Copyright notice") is True def test_footer_rejects_empty(self): assert is_likely_footer("") is False class TestExtractMetadata: def test_returns_embedded_title_and_author(self, tmp_path): pdf = _write_pdf(tmp_path / "m.pdf", title="A Title", author="An Author") meta = extract_metadata(pdf) assert meta["title"] == "A Title" assert meta["author"] == "An Author" def test_missing_title_author_stay_none(self, tmp_path): # No synthesis: absent title/author must surface as None so the # frontmatter renders null rather than a filename-derived guess. pdf = _write_pdf(tmp_path / "bare.pdf") meta = extract_metadata(pdf) assert meta["title"] is None assert meta["author"] is None class TestConvertPdfToMarkdown: def test_blank_pdf_emits_frontmatter_and_no_text_marker(self, tmp_path): pdf = _write_pdf(tmp_path / "doc.pdf", title="Probe Paper") out = tmp_path / "doc.md" # arxiv_id=None keeps it offline: the PDF's embedded metadata drives the # frontmatter and no arXiv fetch is attempted. convert_pdf_to_markdown(pdf, out, arxiv_id=None) text = out.read_text(encoding="utf-8") assert text.startswith("---") assert "source_type: " in text assert "pdf" in text assert "Probe Paper" in text assert "No text extracted" in text -
test_pdf_image_lib.py 3.8 KB
"""Tests for the pdf2image-based image library. The pure kernel (split_image_columns) and the pypdf metadata helper run with no external binary. The full convert_pdf_to_images path needs poppler (pdf2image.convert_from_path shells out to pdftoppm), so its smoke test is skipped when poppler is absent rather than failing the suite. """ import importlib import shutil from pathlib import Path import pytest from PIL import Image from pypdf import PdfWriter from arxiv_doc_builder.pdf_image_lib import ( convert_pdf_to_images, extract_metadata, split_image_columns, ) _NO_POPPLER = shutil.which("pdftoppm") is None def _write_pdf(path: Path, *, title: str | None = None, pages: int = 1) -> Path: writer = PdfWriter() for _ in range(pages): writer.add_blank_page(width=612, height=792) if title is not None: writer.add_metadata({"/Title": title}) with path.open("wb") as fh: writer.write(fh) return path class TestSplitImageColumns: @pytest.mark.parametrize("width", [100, 101]) def test_two_columns_tile_width_exactly(self, width): # Boundaries derived in the plan: i=0 -> [0, W//2), i=1 (last) -> [W//2, W), # an exact gapless non-overlapping cover. The right column is one pixel # wider for odd W. img = Image.new("RGB", (width, 40)) cols = split_image_columns(img, num_columns=2) assert [c.size for c in cols] == [(width // 2, 40), (width - width // 2, 40)] assert sum(c.size[0] for c in cols) == width def test_three_columns_tile_width_exactly(self): width = 101 img = Image.new("RGB", (width, 30)) cols = split_image_columns(img, num_columns=3) # Last column absorbs the division remainder; widths still sum to W. assert sum(c.size[0] for c in cols) == width assert len(cols) == 3 class TestExtractMetadata: def test_title_from_embedded_metadata(self, tmp_path): pdf = _write_pdf(tmp_path / "m.pdf", title="Vision Paper", pages=2) meta = extract_metadata(pdf) assert meta["title"] == "Vision Paper" assert meta["total_pages"] == 2 def test_missing_title_falls_back_to_stem(self, tmp_path): # Display-oriented variant: unlike the converter lib, a missing title # falls back to the file stem (and author to "Unknown") for direct print. pdf = _write_pdf(tmp_path / "fallback.pdf") meta = extract_metadata(pdf) assert meta["title"] == "fallback" assert meta["author"] == "Unknown" class TestShimsImportAsPackageMembers: @pytest.mark.parametrize( "module", [ "arxiv_doc_builder.convert_pdf_with_vision", "arxiv_doc_builder.convert_pdf_split_columns", ], ) def test_shim_resolves_shared_lib_under_package_import(self, module): # The shims must resolve their pdf_image_lib import when loaded as # package members (the `python -m arxiv_doc_builder.<shim>` path), not # only as bare scripts under uv. A regression to a bare # `from pdf_image_lib import ...` raises ModuleNotFoundError here, since # pdf_image_lib is not a top-level module. Importing runs the module's # top level but not main() (guarded by __name__ == "__main__"). importlib.import_module(module) @pytest.mark.skipif(_NO_POPPLER, reason="poppler (pdftoppm) not installed") class TestConvertPdfToImages: def test_single_page_produces_full_plus_two_columns(self, tmp_path): pdf = _write_pdf(tmp_path / "doc.pdf", title="T") out_dir = tmp_path / "images" image_paths, metadata = convert_pdf_to_images(pdf, out_dir, dpi=72) # One full-page image + two column images for the single page. assert len(image_paths) == 3 assert metadata["total_pages"] == 1 assert all(p.exists() for p in image_paths) -
test_skill_placeholders.py 1.7 KB
"""SKILL.md's `{SAFE_ID}` examples are what `safe_arxiv_id` returns. The agent builds paths under the output directory from `{SAFE_ID}` as SKILL.md defines it, while `convert-paper` builds them with `safe_arxiv_id`. A change to the replacement would update that function's own tests and leave the SKILL.md example stale, so the example is checked against the function here. """ import re from arxiv_doc_builder.arxiv_id import safe_arxiv_id, validate_arxiv_id from conftest import read_skill_md _OUTPUT_DIR_BULLET = re.compile(r"^(\s*)- `--output-dir`:") # The failure section uses the same arrow to route an output to a reference # file, so pairs are taken from the `--output-dir` bullet alone. _EXAMPLE = re.compile(r"`([^`]+)` → `([^`]+)`") def _output_dir_bullet(text: str) -> str: """The `--output-dir` bullet with its continuation lines.""" lines = text.splitlines() for i, line in enumerate(lines): match = _OUTPUT_DIR_BULLET.match(line) if match: indent = len(match.group(1)) bullet = [line] for rest in lines[i + 1 :]: if len(rest) - len(rest.lstrip()) <= indent: break bullet.append(rest) return "\n".join(bullet) raise AssertionError("SKILL.md has no `--output-dir` bullet") def test_safe_id_examples_match_safe_arxiv_id() -> None: examples = _EXAMPLE.findall(_output_dir_bullet(read_skill_md())) # An ID without `/` maps to itself under any replacement, so only a # slash-bearing example exercises it. assert any("/" in arxiv_id for arxiv_id, _ in examples), examples for arxiv_id, safe_id in examples: validate_arxiv_id(arxiv_id) assert safe_arxiv_id(arxiv_id) == safe_id -
test_skill_references.py 8.6 KB
"""Every `references/*.md` path the agent is sent to names a file that exists. SKILL.md keeps only the condition under which a procedure applies and the reference file that holds it, so a pointer naming a missing file leaves the agent with the condition and no procedure. The test scans SKILL.md, every `.md` file in references/, the source of every module in the package (comments and messages alike), and the evaluated remedy `convert_latex.py` prints when it kills a runaway pandoc. The same holds for the package scripts the agent is told to run. When a command in SKILL.md or references/ names a script, that script exists, and the command names it by a path the agent can run from its own working directory. For the reference files the converse holds too. Every `.md` file in references/ is named by SKILL.md or by a reference file that is itself reached from SKILL.md, so none sits where no pointer leads. That check reads the text only: it shows a pointer exists, not that an agent follows it. """ import re import pytest from arxiv_doc_builder import convert_latex from conftest import PACKAGE_DIR, SKILL_DIR, read_skill_md _REFERENCE = re.compile(r"references/[\w.-]+\.md") # The procedures SKILL.md sends the agent to when a situation it names holds. # Each must be named in SKILL.md itself, since SKILL.md is what the agent has # loaded. _CONDITIONAL_REFERENCES = { "references/multiple-documentclass.md", "references/pandoc-failures.md", "references/unknown-arity-macros.md", "references/pandoc-runaway.md", "references/pdf-conversion.md", "references/source-edits.md", } def _markdown_sources() -> list[tuple[str, str]]: sources = [("SKILL.md", read_skill_md())] reference_files = sorted((SKILL_DIR / "references").glob("*.md")) assert reference_files, "no reference files found" sources += [ (f"references/{p.name}", p.read_text(encoding="utf-8")) for p in reference_files ] return sources def _pointer_sources() -> list[tuple[str, str]]: sources = _markdown_sources() # Module sources cover comments and single-literal messages; the evaluated # remedy covers the string the user sees, which its source splits across # adjacent literals. modules = sorted(PACKAGE_DIR.glob("*.py")) sources += [(p.name, p.read_text(encoding="utf-8")) for p in modules] sources.append(("_RUNAWAY_REMEDY", convert_latex._RUNAWAY_REMEDY)) return sources _SOURCES = _pointer_sources() @pytest.mark.parametrize(("name", "text"), _SOURCES, ids=[name for name, _ in _SOURCES]) def test_every_reference_path_exists(name: str, text: str) -> None: missing = [ ref for ref in _REFERENCE.findall(text) if not (SKILL_DIR / ref).is_file() ] assert not missing, f"{name} points at missing files: {missing}" def test_skill_md_points_at_every_conditional_procedure() -> None: skill_md = read_skill_md() assert _CONDITIONAL_REFERENCES <= set(_REFERENCE.findall(skill_md)) def test_runaway_remedy_names_a_reference_file() -> None: # Without a token, the existence check above passes on nothing. assert "references/pandoc-runaway.md" in _REFERENCE.findall( convert_latex._RUNAWAY_REMEDY ) def _unreachable_references(skill_md: str, references: dict[str, str]) -> list[str]: """The reference files that no chain of pointers starting in SKILL.md names. ``references`` maps each file's ``references/<name>.md`` path to its text. A pointer at a path with no file leads nowhere and is skipped; the existence check above is what reports it. """ reached: set[str] = set() pending = set(_REFERENCE.findall(skill_md)) while pending: ref = pending.pop() reached.add(ref) pending |= set(_REFERENCE.findall(references.get(ref, ""))) - reached return sorted(set(references) - reached) def test_unreachable_references_follows_pointers_between_files() -> None: references = { "references/direct.md": "Then read `references/chained.md`.", "references/chained.md": "Go back to `references/direct.md`, or on to " "`references/deep.md`.", "references/deep.md": "Nothing further.", "references/orphan-a.md": "See `references/orphan-b.md`.", "references/orphan-b.md": "See `references/orphan-a.md`.", } skill_md = "Read `references/direct.md`, or `references/missing.md`." assert _unreachable_references(skill_md, references) == [ "references/orphan-a.md", "references/orphan-b.md", ] def test_every_reference_file_is_reachable_from_skill_md() -> None: references = { name: text for name, text in _markdown_sources() if name != "SKILL.md" } unreachable = _unreachable_references(read_skill_md(), references) assert not unreachable, f"no pointer chain from SKILL.md names: {unreachable}" # A command the agent runs names a package script by a path starting at the # skill root. The agent runs it from its own working directory, where that path # resolves only if it starts with the literal placeholder `SKILL_DIR/`, which # SKILL.md tells the agent to replace with the skill root's absolute path. The # whole path sits in double quotes, so a skill root containing spaces stays one # argument. Prose outside fenced blocks may name a module without being a # command. _FENCE = re.compile(r"^\s*```") _SCRIPT_PATH = re.compile(r'(\S*?)arxiv_doc_builder/([\w.-]+\.py)("?)') _QUOTED_PREFIX = '"SKILL_DIR/' # The files that carry such commands; without them the check below could pass # on nothing, for instance if an indented fence stopped being recognized. _FILES_WITH_SCRIPT_COMMANDS = {"SKILL.md", "references/pdf-conversion.md"} def _fenced_lines(text: str) -> list[str]: """The lines inside fenced code blocks, fences excluded.""" lines = [] inside = False for line in text.splitlines(): if _FENCE.match(line): inside = not inside elif inside: lines.append(line) return lines def _script_commands(text: str) -> list[tuple[str, str, str]]: """Each script a fenced line names, as (prefix, name, closing quote). The prefix is the non-space text right before ``arxiv_doc_builder/``; the closing quote is ``'"'`` when one follows the name, else ``""``. """ return [c for line in _fenced_lines(text) for c in _SCRIPT_PATH.findall(line)] def test_script_commands_scan_fenced_lines_only() -> None: text = ( "See `arxiv_doc_builder/prose.py`.\n" " ```bash\n" " uv run arxiv_doc_builder/bare.py x\n" ' uv run --project "SKILL_DIR" "SKILL_DIR/arxiv_doc_builder/quoted.py" x\n' " ```\n" "arxiv_doc_builder/after.py\n" ) assert _script_commands(text) == [ ("", "bare.py", ""), (_QUOTED_PREFIX, "quoted.py", '"'), ] _MARKDOWN = _markdown_sources() @pytest.mark.parametrize( ("name", "text"), _MARKDOWN, ids=[name for name, _ in _MARKDOWN] ) def test_script_commands_start_at_skill_dir(name: str, text: str) -> None: commands = _script_commands(text) unquoted = [ script for prefix, script, closing in commands if prefix != _QUOTED_PREFIX or closing != '"' ] assert not unquoted, f'{name} runs scripts not written "SKILL_DIR/...": {unquoted}' missing = [s for _, s, _ in commands if not (PACKAGE_DIR / s).is_file()] assert not missing, f"{name} runs scripts that do not exist: {missing}" def test_script_commands_are_found() -> None: carrying = {name for name, text in _MARKDOWN if _script_commands(text)} assert _FILES_WITH_SCRIPT_COMMANDS <= carrying # A script with a PEP 723 header gets an environment built from that header. # One without it runs in whatever project uv picks, which older uv releases # take from the working directory; there the skill's `requires-python` and # dependencies are not guaranteed. Such a script's command therefore names # the skill's own project. _PEP_723_HEADER = "# /// script" _SKILL_PROJECT = '--project "SKILL_DIR"' def _has_inline_metadata(script: str) -> bool: return _PEP_723_HEADER in (PACKAGE_DIR / script).read_text(encoding="utf-8") def test_scripts_without_inline_metadata_run_in_the_skill_project() -> None: headerless = [ (name, line.strip()) for name, text in _MARKDOWN for line in _fenced_lines(text) for _, script, _ in _SCRIPT_PATH.findall(line) if not _has_inline_metadata(script) ] # convert_paper.py is such a script, so the check has something to cover. assert headerless, "no command runs a script without a PEP 723 header" lacking = [(name, line) for name, line in headerless if _SKILL_PROJECT not in line] assert not lacking, f"commands without {_SKILL_PROJECT}: {lacking}" -
test_version.py 8.5 KB
"""Tests for the --version emitter and its source-tree fallback. The skill is normally run straight from the checkout (no install), so the fallback path — parsing pyproject.toml — is the *primary* path here, not a rare edge. These tests pin that behavior and the CLI contract. """ import importlib import subprocess import sys import tomllib from importlib import metadata import pytest from conftest import PACKAGE_DIR, SKILL_DIR from arxiv_doc_builder import _version as version_module from arxiv_doc_builder import convert_paper from arxiv_doc_builder._version import ( _DIST_NAME, _version_from_pyproject, read_version, ) _PYPROJECT = SKILL_DIR / "pyproject.toml" def _expected_version() -> str: with _PYPROJECT.open("rb") as f: return tomllib.load(f)["project"]["version"] def _raiser(error: Exception): """Return a stand-in callable that raises ``error`` whatever it is passed.""" def _raise(*args, **kwargs): raise error return _raise @pytest.fixture def not_installed(monkeypatch): """Make the distribution look uninstalled, so the pyproject fallback runs.""" monkeypatch.setattr( metadata, "version", _raiser(metadata.PackageNotFoundError(_DIST_NAME)) ) @pytest.fixture def sibling_pyproject(tmp_path, monkeypatch): """Point the fallback at ``tmp_path / "pyproject.toml"``; return its writer. The fallback finds pyproject.toml from the module's ``__file__``, so that is what gets patched. Until the writer is called, the file does not exist. """ monkeypatch.setattr( version_module, "__file__", str(tmp_path / "pkg" / "_version.py") ) def _write(content: bytes) -> None: (tmp_path / "pyproject.toml").write_bytes(content) return _write def test_dist_name_is_hyphenated(): # The lookup name must be the distribution name, which intentionally # differs from the underscored import name. A regression to # "arxiv_doc_builder" would silently miss installed metadata. assert _DIST_NAME == "arxiv-doc-builder" assert _DIST_NAME != PACKAGE_DIR.name def test_read_version_matches_pyproject_ssot(): # In the source tree (uninstalled), read_version resolves via the # pyproject fallback and must equal the [project] version SSOT. assert read_version() == _expected_version() def test_version_from_pyproject_matches_ssot(): assert _version_from_pyproject() == _expected_version() def test_read_version_prefers_installed_metadata(monkeypatch): # When dist-info exists, read_version must return the metadata version # (the installed-CLI SSOT) and query it under the hyphenated dist name — # NOT silently fall through to the pyproject parse. The sentinel differs # from the pyproject version so a fall-through would fail the assertion. seen = {} def _fake_version(dist): seen["dist"] = dist return "9.9.9-installed" monkeypatch.setattr(metadata, "version", _fake_version) assert read_version() == "9.9.9-installed" assert seen["dist"] == _DIST_NAME @pytest.mark.parametrize( ("owner", "attr", "stand_in"), [ # [project].version is absent, so the lookup raises KeyError pytest.param(tomllib, "load", lambda f: {}, id="missing_key"), pytest.param( tomllib, "load", _raiser(tomllib.TOMLDecodeError("malformed")), id="decode_error", ), # the pyproject path exists but cannot be opened pytest.param( version_module.Path, "open", _raiser(OSError("unreadable")), id="os_error" ), # the path to the pyproject cannot be built pytest.param( version_module.Path, "resolve", _raiser(RuntimeError("symlink loop")), id="resolve_error", ), ], ) def test_version_from_pyproject_degrades_to_unknown(monkeypatch, owner, attr, stand_in): # A failure at each step of the fallback — locating the pyproject, # opening it, parsing it, and looking up [project].version — must degrade # to "unknown" instead of propagating. monkeypatch.setattr(owner, attr, stand_in) assert _version_from_pyproject() == "unknown" @pytest.mark.parametrize( "content", [ pytest.param( b'[project]\nname = "\xff"\nversion = "1.2.3"\n', id="undecodable_byte" ), pytest.param(b'project = "x"\n', id="project_not_a_table"), pytest.param(b"[project]\nversion = 1\n", id="integer_version"), pytest.param(b"[project]\nversion = 2026-10-04\n", id="date_version"), pytest.param(b"[tool]\nx = 1\n", id="no_project_table"), ], ) def test_malformed_pyproject_file_degrades_to_unknown(sibling_pyproject, content): # Real files through the real read, parse and lookup, so what each file # provokes is not an outcome a stub chose. sibling_pyproject(content) assert _version_from_pyproject() == "unknown" def test_absent_pyproject_file_degrades_to_unknown(sibling_pyproject): assert _version_from_pyproject() == "unknown" def test_well_formed_pyproject_file_yields_its_version(sibling_pyproject): # The control for the two tests above: the value differs from the # repository's own [project] version, so it can only have come from the # redirected file. sibling_pyproject(b'[project]\nversion = "1.2.3"\n') assert _version_from_pyproject() == "1.2.3" def test_read_version_falls_back_when_not_installed(not_installed): # PackageNotFoundError (no dist-info — the source-tree run mode) must # route to the pyproject fallback rather than propagate. assert read_version() == _expected_version() def test_read_version_degrades_to_unknown_on_malformed_fallback_file( not_installed, sibling_pyproject ): # The route --version takes from a checkout: not installed, then a # pyproject the parser rejects. sibling_pyproject(b'[project]\nname = "\xff"\nversion = "1.2.3"\n') assert read_version() == "unknown" def test_read_version_degrades_to_unknown_on_corrupt_metadata(monkeypatch): # Unexpected metadata failures must degrade too, not just # PackageNotFoundError: a corrupt/unparseable installed distribution # makes metadata.version raise a generic error, which must degrade to # "unknown" rather than propagate. monkeypatch.setattr(metadata, "version", _raiser(RuntimeError("corrupt metadata"))) assert read_version() == "unknown" def test_read_version_degrades_to_unknown_on_non_string_metadata(monkeypatch): # None is what metadata.version was observed to return, without raising, # for a distribution whose METADATA carries no Version field. monkeypatch.setattr(metadata, "version", lambda dist: None) assert read_version() == "unknown" def test_read_version_degrades_to_unknown_when_metadata_cannot_be_imported( monkeypatch, ): # A None entry in sys.modules makes the import raise ImportError. The # attribute is removed first: `from importlib import metadata` would # otherwise pick the already-imported submodule off the package. monkeypatch.delattr(importlib, "metadata") monkeypatch.setitem(sys.modules, "importlib.metadata", None) assert read_version() == "unknown" @pytest.mark.parametrize("version", ["1%", "%(prog)s"]) def test_cli_version_prints_percent_signs_literally(monkeypatch, capsys, version): # argparse %-formats the whole version string, so an unescaped "%" in the # resolved version either raises ("1%") or is expanded ("%(prog)s"). monkeypatch.setattr(convert_paper, "read_version", lambda: version) monkeypatch.setattr(sys, "argv", ["convert_paper.py", "--version"]) with pytest.raises(SystemExit) as exit_info: convert_paper.main() assert exit_info.value.code == 0 assert capsys.readouterr().out == f"convert_paper.py {version}\n" def test_cli_version_flag_emits_version(): # `--version` must print and exit 0 without requiring the positional # arxiv_id (action="version" is eager). result = subprocess.run( [sys.executable, str(PACKAGE_DIR / "convert_paper.py"), "--version"], capture_output=True, text=True, ) assert result.returncode == 0, result.stderr assert _expected_version() in result.stdout @pytest.mark.parametrize("flag", ["-V", "--version"]) def test_cli_version_short_and_long(flag): result = subprocess.run( [sys.executable, str(PACKAGE_DIR / "convert_paper.py"), flag], capture_output=True, text=True, ) assert result.returncode == 0, result.stderr assert _expected_version() in result.stdout -
test_version_drift.py 10.9 KB
"""Tests for version drift detection logic. They cover the pure decisions the drift check rests on, namely whether to re-fetch, which revision the run should hold, what version the lookup reports, whether to write the record, and what to say when a run fetched material without advancing it. The cache read and write helpers underneath are covered too. Nothing here touches the network. """ import pytest from conftest import PROBE_ERROR, PROBE_VERSION, seed_cached_source from arxiv_doc_builder.arxiv_metadata import ( METADATA_SOURCE_ARXIV, METADATA_SOURCE_DATACITE, ) from arxiv_doc_builder.fetch_paper import ( _format_sidecar_skip_warning, _latest_version, _needs_refresh, _read_cached_version, _record_version, _target_version, _write_cached_version, _METADATA_FILE, ) def test_needs_refresh_no_cache_with_latest(): """No recorded version → re-fetch to establish version record.""" assert _needs_refresh(None, "2409.03108v2") is True def test_needs_refresh_no_cache_lookup_failed(): """No record and no version from the lookup → trust cache (no re-fetch).""" assert _needs_refresh(None, None) is False def test_needs_refresh_version_matches(): """Cached version matches latest → skip.""" assert _needs_refresh("2409.03108v2", "2409.03108v2") is False def test_needs_refresh_version_differs(): """Cached v1, latest v2 → re-fetch.""" assert _needs_refresh("2409.03108v1", "2409.03108v2") is True def test_needs_refresh_lookup_failed_with_cache(): """Lookup reports no version but cache exists → trust cache.""" assert _needs_refresh("2409.03108v1", None) is False def test_write_then_read_roundtrip(tmp_path): """Write and read back the version string.""" _write_cached_version(tmp_path, "2409.03108v2") assert _read_cached_version(tmp_path) == "2409.03108v2" def test_read_missing_file(tmp_path): """No metadata file → None.""" assert _read_cached_version(tmp_path) is None def test_read_corrupt_file(tmp_path): """Corrupt metadata → None (graceful fallback).""" (tmp_path / _METADATA_FILE).write_text("not json", encoding="utf-8") assert _read_cached_version(tmp_path) is None @pytest.mark.parametrize( "content", ['{"version": 2}', '{"version": null}', '["2409.03108v2"]'], ids=["number", "null", "not-an-object"], ) def test_a_version_that_is_not_text_reads_as_absent(tmp_path, content): # The readers downstream treat this value as text, which raises or # misfires on anything else. A hand-edited sidecar must not end the run. (tmp_path / _METADATA_FILE).write_text(content, encoding="utf-8") cached = _read_cached_version(tmp_path) assert cached is None assert _needs_refresh(cached, "2409.03108v2") is True assert ( _target_version( tmp_path, "2409.03108v2", cached, pinned=False, source=METADATA_SOURCE_DATACITE, ) == "2409.03108v2" ) def test_write_overwrites(tmp_path): """Second write overwrites the first.""" _write_cached_version(tmp_path, "2409.03108v1") _write_cached_version(tmp_path, "2409.03108v2") assert _read_cached_version(tmp_path) == "2409.03108v2" def test_a_pinned_id_after_a_record_of_another_revision_refreshes_once(tmp_path): """A pinned id reports its own revision; a sidecar naming another one re-fetches once, which downloads the same pinned revision, and is stable afterwards.""" _write_cached_version(tmp_path, "2409.03108v2") assert _needs_refresh(_read_cached_version(tmp_path), "2409.03108v1") is True assert _record_version(tmp_path, "2409.03108v1", fetched=True) is True assert _needs_refresh(_read_cached_version(tmp_path), "2409.03108v1") is False # --- which revision the run should hold ------------------------------------- @pytest.mark.parametrize( ("cached", "latest", "pinned", "target"), [ (None, "2409.03108v2", False, "2409.03108v2"), ("2409.03108v1", "2409.03108v2", False, "2409.03108v2"), ("2409.03108v2", "2409.03108v2", False, "2409.03108v2"), # The fallback can lag arXiv, and an earlier run may already hold the # later revision. Going back would delete the source only to fetch it # again once that record catches up. ("2409.03108v3", "2409.03108v2", False, "2409.03108v3"), ("2409.03108v10", "2409.03108v9", False, "2409.03108v10"), # A requested revision is what the user asked for, whatever is cached. ("2409.03108v3", "2409.03108v2", True, "2409.03108v2"), # A record of another paper says nothing about this lookup's revision. ("2409.03109v3", "2409.03108v2", False, "2409.03108v2"), # A hand edit can leave a trailing newline, and that value would go # into the download URL if it won. ("2409.03108v3\n", "2409.03108v2", False, "2409.03108v2"), ("math/0309136v3", "math/0309136v2", False, "math/0309136v3"), ("2409.03108v3", None, False, None), ], ids=[ "no-record", "record-older", "record-equal", "record-newer", "record-newer-numerically", "pinned-overrides-newer-record", "record-of-another-paper", "record-with-a-trailing-newline", "legacy-record-newer", "no-version-from-lookup", ], ) def test_target_version_keeps_a_later_recorded_revision_of_an_unpinned_id( tmp_path, cached, latest, pinned, target ): # The recorded revision speaks for material on disk, so these cases seed # some; the case without it is its own test below. They read as the # fallback answering, the only source whose record can trail arXiv. seed_cached_source(tmp_path) target_version = _target_version( tmp_path, latest, cached, pinned=pinned, source=METADATA_SOURCE_DATACITE ) assert target_version == target def test_a_record_ahead_of_arxivs_own_answer_does_not_win(tmp_path): # arXiv's record is authoritative about its own revisions, so a cached # revision ahead of it is not a lag. Letting it win would hold the paper # at that revision for as long as the file stayed — no lookup could move # it, since the comparison would keep going the same way. seed_cached_source(tmp_path) cached = "2409.03108v99" from_arxiv = _target_version( tmp_path, "2409.03108v2", cached, pinned=False, source=METADATA_SOURCE_ARXIV ) assert from_arxiv == "2409.03108v2" from_datacite = _target_version( tmp_path, "2409.03108v2", cached, pinned=False, source=METADATA_SOURCE_DATACITE, ) assert from_datacite == "2409.03108v99" def test_a_record_without_a_cached_source_cannot_outvote_the_lookup(tmp_path): # A sidecar naming a revision that no longer exists would otherwise be # re-confirmed on every run, while the source download for that revision # failed on every run and the LaTeX path never came back. lagging = { "cached": "2409.03108v99", "pinned": False, "source": METADATA_SOURCE_DATACITE, } assert _target_version(tmp_path, "2409.03108v2", **lagging) == "2409.03108v2" # A cached PDF does not change that: the source would still be fetched at # the recorded revision, which is the download that fails. (tmp_path / "pdf").mkdir() (tmp_path / "pdf" / "2409.03108.pdf").write_bytes(b"%PDF-stub") assert _target_version(tmp_path, "2409.03108v2", **lagging) == "2409.03108v2" # A cached source is what the record speaks for, so it wins there. seed_cached_source(tmp_path) assert _target_version(tmp_path, "2409.03108v2", **lagging) == "2409.03108v99" # --- what the lookup yields, and when the sidecar advances ------------------ def test_latest_version_reads_the_version_off_a_successful_probe(probe_with_version): assert _latest_version(probe_with_version) == PROBE_VERSION def test_latest_version_is_none_when_the_probe_failed(failed_probe): assert _latest_version(failed_probe) is None def test_latest_version_is_none_when_the_record_carried_no_version( probe_without_version, ): # A record can parse and still yield no version, reaching the same # decision as a failed lookup by a different route. That is why the # sidecar branches on the version and not on the status. assert _latest_version(probe_without_version) is None @pytest.mark.parametrize( ("latest", "fetched", "expected"), [ ("2409.03108v2", True, True), ("2409.03108v2", False, False), (None, True, False), (None, False, False), ], ids=[ "version-and-material", "version-no-material", "no-version-but-material", "neither", ], ) def test_record_version_writes_only_with_both_a_version_and_material( tmp_path, latest, fetched, expected ): assert _record_version(tmp_path, latest, fetched=fetched) is expected assert _read_cached_version(tmp_path) == (latest if expected else None) def test_record_version_leaves_an_existing_record_alone_when_it_declines(tmp_path): # Declining must not clear what a previous run established, or the next # drift check would re-fetch a paper whose version it already knew. _write_cached_version(tmp_path, "2409.03108v1") assert _record_version(tmp_path, None, fetched=True) is False assert _read_cached_version(tmp_path) == "2409.03108v1" def test_sidecar_skip_warning_reports_a_failed_lookup_by_its_cause(failed_probe): text = _format_sidecar_skip_warning("2409.03108", failed_probe) assert PROBE_ERROR in text assert _METADATA_FILE in text def test_sidecar_skip_warning_names_the_missing_version_when_the_record_was_read( probe_without_version, ): # This is the cell a status-only branch would leave silent: the lookup # succeeded, so there is no error to quote, yet the sidecar still did not # advance and the user still needs to hear it. text = _format_sidecar_skip_warning("2409.03108", probe_without_version) assert "2409.03108" in text assert "no version" in text assert _METADATA_FILE in text # A usable record *was* read here, so the warning must not say otherwise. # A message shared with the conversion paths would, since theirs opens by # reporting that no usable record was read. assert "no usable" not in text def test_sidecar_skip_warning_stays_off_the_frontmatter(failed_probe): # This step writes no document, so it has no null fields to explain. # Borrowing the conversion paths' wording would describe a surface the # fetch step never touches. text = _format_sidecar_skip_warning("2409.03108", failed_probe) assert "null" not in text assert "frontmatter" not in text def test_sidecar_skip_warning_refuses_a_probe_that_did_report_a_version( probe_with_version, ): # Its whole text asserts that no version was available. Called on a probe # that supplied one, every sentence it returns would be false. with pytest.raises(ValueError): _format_sidecar_skip_warning("2409.03108", probe_with_version)
-
-
pyproject.toml 1.7 KB
[project] name = "arxiv-doc-builder" version = "0.4.4" requires-python = ">=3.11" dependencies = [] # The PDF converter tier is heavy and optional. Declaring it as an extra keeps # the core install dependency-free (consumers who only need the LaTeX happy path # pay nothing) while `pip install arxiv-doc-builder[pdf]` pulls the stack on # demand. The standalone convert_pdf_*.py scripts still carry their own PEP 723 # inline deps for `uv run --no-project` subprocess execution; this extra is what # lets the tier import under pytest and pyright. [project.optional-dependencies] pdf = ["pdfplumber", "pdf2image", "pypdf", "pillow"] [project.scripts] convert-paper = "arxiv_doc_builder.convert_paper:main" [build-system] requires = ["hatchling"] build-backend = "hatchling.build" [tool.hatch.build.targets.wheel] packages = ["arxiv_doc_builder"] [dependency-groups] # Pull the pdf extra into the dev env (self-reference) so plain `uv run pytest` # and `uv run pyright` resolve the PDF tier without a separate --extra flag. dev = ["pytest>=8", "pyyaml>=6", "pyright>=1.1.0", "ruff>=0.15.18", "arxiv-doc-builder[pdf]"] [tool.pytest.ini_options] testpaths = ["tests"] [tool.pyright] # Typecheck the whole package surface plus the test suite. The PDF converter # tier (the convert_pdf_*.py scripts plus the shared pdf_converter_lib.py / # pdf_image_lib.py) is now importable: the heavy third-party imports it depends # on — pdfplumber / pdf2image / pypdf / pillow — resolve via the `pdf` extra that # the dev group pulls in, so pyright no longer reports them as unresolved. The # two refactored shims import only sibling modules plus stdlib; the heavy # imports live in the libs (and the pdfplumber-tier scripts). include = ["arxiv_doc_builder", "tests"] -
SKILL.md 3.2 KB
--- name: arxiv-doc-builder description: Convert an arXiv paper to Markdown for reading or implementation reference. Use when asked to convert, fetch, or create documentation for an arXiv paper by its ID, or when a paper with a known arXiv ID needs to be read or referenced. Fetches the LaTeX source when available (plus the PDF) and converts it with pandoc; PDF-only papers get a naive single-column fallback. --- # arXiv Document Builder ## Procedure 1. Run the converter: ```bash # Using global command (recommended) convert-paper ARXIV_ID [--output-dir DIR] # Using script directly uv run --project "SKILL_DIR" "SKILL_DIR/arxiv_doc_builder/convert_paper.py" ARXIV_ID [--output-dir DIR] ``` - `SKILL_DIR`: replace with the absolute path of the directory this SKILL.md is in. It is a placeholder, not a shell variable. Keep the double quotes around it, so a path containing spaces stays one argument. The command then runs from any working directory. - `--output-dir`: Directory where `{SAFE_ID}/{SAFE_ID}.md` will be created. **Default: current working directory** (not a `papers/` subdirectory). `{SAFE_ID}`, here and below, is `ARXIV_ID` with `/` replaced by `_`, as in `hep-th/9711200` → `hep-th_9711200`. A new-style ID contains no `/`, so it is used unchanged. - Use absolute paths to control output location precisely. `convert-paper` does the metadata lookup, downloads, extraction, and directory creation itself; do not run curl, tar, or mkdir for them. 2. On success, `convert-paper` prints the Markdown file's path on its `Output:` line. The file opens with a YAML frontmatter block of provenance metadata; `references/output-format.md` documents its fields, including what `metadata_status` records. The File Organization section of `references/output-format.md` shows the layout of the paper's directory, `{SAFE_ID}/`. ## When Conversion Fails or Falls Back to PDF If you edited a file under `{SAFE_ID}/source/` and re-ran `convert-paper`, read `references/source-edits.md` first and follow it before any line below. When you made no such edit, or once that file no longer tells you to edit again or re-run, read the file that the line matching the output points to. For an output not listed here, act on what the output itself says. Whichever file you follow, change the source only as it directs. Do NOT attempt broad preprocessing (replacing documentclass, expanding `\newcommand`, removing environments, etc.) — pandoc handles revtex4/revtex4-2, custom commands, `picture` environments, and theorem environments correctly. - `convert-paper` exits with code 2 after printing `Error: Found N files with \documentclass` → `references/multiple-documentclass.md` - It prints `Pandoc conversion failed:`, and the pandoc message after it contains `unexpected (` or `unexpected [` → `references/unknown-arity-macros.md` - It prints `Pandoc conversion failed:`, and the pandoc message after it contains neither → `references/pandoc-failures.md` - It prints `Pandoc did not finish within` or `Pandoc exceeded the <N> MB memory watchdog`, or a pandoc run has not returned → `references/pandoc-runaway.md` - It prints `No LaTeX source, falling back to naive PDF conversion...` → `references/pdf-conversion.md` -
uv.lock 124.6 KB · in bundle
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.