museum-art
Source authentic, high-res PUBLIC-DOMAIN artwork from museum open-access APIs (Met, Cleveland, SMK, Rijksmuseum, NGA, Art Institute of Chicago, Getty, Smithsonian) instead of AI-generated or generic-stock imagery. The default move whenever a visual needs an aesthetic, credible im
Install
npx skills add https://github.com/huytieu/COG-second-brain/tree/main/skills/museum-art
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install huytieu-cog-second-brain@llmmart
git clone https://github.com/huytieu/COG-second-brain.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole huytieu/cog-second-brain collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
museum-art: Public-Domain Artwork for Visuals
Standing rule (adopted 2026-07-24, from Eric Li's post on museum open-access): whenever a visual needs a real image with aesthetic weight, source public-domain museum artwork first - over AI-generated imagery and over generic stock. Curated, historically significant art reads as credible and sophisticated; AI-gen reads as slop. This stacks with, and reinforces, the
no-ai-slopskill and your house image style. It does NOT replace the generativeeditorial-illustrationsskill (that owns claim-driven diagrams/figures) or your house chart style - use museum art for photographic/hero/decorative/mood imagery, generative figures for data and concept diagrams.
When to reach for this
- Blog post hero images, section breaks, mood imagery (the blog-publish image step).
- Deck/slide backgrounds and section dividers, social cards, essay figures, spec cover art.
- Any time the instinct is "generate an image" for something decorative or evocative. Stop and pull a real painting instead.
- NOT for: product screenshots, UI mockups, data charts, logos, or claim-driven explanatory diagrams (those are editorial-illustrations / real captures).
Decision: which source
- Met -> Cleveland -> SMK first. All keyless, one JSON hop, CC0/PD, broad collections. Fastest path to a hi-res image.
- Need Dutch/Flemish masters or decorative arts? Rijksmuseum (keyless, 3 hops).
- Need European antiquities/photography and keyword isn't essential? Getty (keyless, SPARQL).
- Nothing fits, or you want a cross-museum search? Wikimedia Commons API (keyless aggregator, normalized license metadata) is the best general fallback.
- "Historical illustration/engraving/old photo" rather than fine-art painting? Go straight to Internet Archive or Wikimedia Commons.
- Science/health/medical essay figure? Wellcome Collection (keyless, CC0/PD).
How to fetch in THIS environment
- Use WebFetch to hit the JSON search endpoint, then WebFetch/download the returned image URL. These APIs are server-side reachable; no browser needed.
- Exception - Art Institute of Chicago images: the JSON API (
api.artic.edu) is fine, but the image hostwww.artic.edu/iiif/...403s scripted/curl fetches via a Cloudflare bot challenge. Use the JSON metadata from AIC, but download the actual image through a real/headless browser (browser-harness) or prefer a different museum for the image bytes. - Always filter for public domain in the query AND spot-check the per-image license flag before shipping (see Licensing).
Freshness over caching (mandatory)
Fetch fresh per need. Do NOT build a reusable local pool of downloaded images to draw from. A small cached set gets reused everywhere and becomes the new "same stock photo on every post" - sameness is a form of slop, and the variety of a huge open collection is the entire point. Fetching is keyless and sub-second, so there is no cost reason to cache pixels.
- Cache recipes/metadata, not images - that is what this skill's
references/already are. - Commit an image only into the specific artifact that uses it (a post's
assets/, a deck's media) once chosen - for provenance and offline builds. That is an artifact asset, never a shared library other artifacts pull from. - Each new visual = a fresh query. Vary the search terms and the source museum so consecutive posts do not converge on the same few crowd-pleasers.
Verified keyless recipes (2026-07-24, all live-tested)
1. The Met - best all-around default
- Base:
https://collectionapi.metmuseum.org/public/collection/v1/- no key, 80 req/s, CC0 whereisPublicDomain:true. - Search:
GET /search?q=<term>&hasImages=true&isPublicDomain=true->{total, objectIDs[]} - Object:
GET /objects/{id}-> readprimaryImage(full-res JPEG, static, no IIIF hop) orprimaryImageSmall(web-large). - Example:
https://collectionapi.metmuseum.org/public/collection/v1/search?q=sunflowers&hasImages=true&isPublicDomain=true
2. Cleveland Museum of Art - only one with archival TIFF
- Base:
https://openaccess-api.clevelandart.org/api/artworks/- no key, CC0. - Search:
GET /api/artworks/?cc0=1&has_image=1&limit=10(add&q=monet,&skip=10). - Image fields on each result:
images.web.url(900px),images.print.url(3400px JPEG),images.full.url(archival TIFF). Directly downloadable fromopenaccess-cdn.clevelandart.org. - Confirm
share_license_status == "CC0"per result.
3. SMK (Denmark) - one-hop, broad European/Nordic
- Base:
https://api.smk.dk/api/v1- no key. License: Public Domain Mark 1.0 (functionally CC0). - Search:
GET /art/search?keys=*&filters=[public_domain:true]&filters=[has_image:true] - Read
image_nativedirectly from each result (no chain). - Gotcha (verified): pass
filtersas a repeated query param, one[field:value]bracket each. Concatenatingfilters=[public_domain:true][has_image:true]returns 200 but silently ignores the second condition.
4. Rijksmuseum - Dutch/Flemish masters (keyless, 3 hops)
- Base:
https://data.rijksmuseum.nl/search/collection(Linked Art, no key). GET /search/collection?type=painting&imageAvailable=true-> walk object -> VisualItem -> DigitalObject ->access_point[0].idis a ready IIIF URL.- Mixed license: mostly CC0/PD but some CC BY 4.0 - check the per-object rights block (a CC BY item requires attribution).
5. NGA (Washington) - bulk/offline, no live search
- CSV:
https://raw.githubusercontent.com/NationalGalleryOfArt/opendata/main/data/published_images.csv- filteropenaccess=1. - Image (IIIF):
https://api.nga.gov/iiif/{uuid}/full/full/0/default.jpg - Per-image
openaccess=0rows are NOT open (resolution-capped, rights-restricted). Onlyopenaccess=1is free.
6. Art Institute of Chicago - Impressionism/European (image host caveat)
- JSON:
GET https://api.artic.edu/api/v1/artworks/search?query[term][is_public_domain]=true&fields=id,title,image_id- no key, CC0. - Image: build
https://www.artic.edu/iiif/2/{image_id}/full/843,/0/default.jpg- but Cloudflare 403s scripted fetches; download via browser or use another museum for bytes.
7. Getty - European painting/photography/antiquities (SPARQL)
- Base:
https://data.getty.edu/museum/collection/- no key. No keyword search; use SPARQL with?obj crm:P138i_has_representation ?imgwhich returns a ready IIIF image URL in the same query. - Dataset is CC0; per-image rights inconsistently populated - cross-check the object's public getty.edu page (a "Download" button flags true Open Content) before commercial use.
Key-gated + aggregator sources (know these too)
- Smithsonian (needs free
api.data.govkey;DEMO_KEYworks at 30 req/hr):GET https://api.si.edu/openaccess/api/v1.0/search?q=<term> AND online_media_type:Images&api_key=KEY; CC0; checkmedia[].usage.access=="CC0"per image. Extremely broad (SAAM, NPG portraits, Freer/Sackler Asian, Cooper Hewitt design). - Wikimedia Commons (keyless, best aggregator fallback):
https://commons.wikimedia.org/w/api.php?action=query&generator=search&gsrsearch=<term>&gsrnamespace=6&prop=imageinfo&iiprop=url|extmetadata; license per file inextmetadata.LicenseShortName; raw bytes viaSpecial:FilePath/{filename}. - Wellcome Collection (keyless):
https://api.wellcomecollection.org/catalogue/v2/works?query=<term>; CC0/PD medical/scientific/historical imagery. - Internet Archive (keyless):
https://archive.org/advancedsearch.php?q=<term>&fl[]=identifier&output=json; PD book illustrations, engravings, historical photos, ephemera. - Europeana (free key):
https://api.europeana.eu/record/v2/search.json?wskey=<KEY>&query=<term>; 50M+ items across European institutions. - Harvard Art Museums (free key):
https://api.harvardartmuseums.org/object?apikey=<KEY>&q=<term>. - Yale LUX (keyless, Linked Art like Getty):
https://lux.collections.yale.edu/api/search/...; strong British art + rare books. - Paris Musees (keyless, explicit CC0):
parismuseescollections.paris.fr/opendata.paris.fr; French painting/decorative arts. - NYPL Digital Collections (free key):
https://api.repository.library.nyc/...; PD prints, maps, photos, illustrations. - DPLA (free key):
https://api.dp.la/v2/items?q=<term>&api_key=<KEY>; federated US collections.
Licensing rules (apply before shipping any image)
- True CC0 (no attribution required): Met, Cleveland, Art Institute of Chicago, Smithsonian, NGA (dataset-level; image gated by per-image
openaccess=1). - Public Domain Mark / open-but-not-CC0 (free to use, credit requested not required): SMK, and the PD subset of Rijksmuseum.
- Verify per-image, do NOT blanket-trust collection-level "open access": Getty, Rijksmuseum (CC BY items require attribution), NGA and Smithsonian per-image flags.
- Rule: when the API exposes a rights/license field, filter on it in the query AND spot-check it on the chosen image. Never assume "the collection is open access" implies "this specific image is."
- Good practice everywhere: a one-line credit ("Digital image courtesy of [Museum]") costs nothing and covers mixed-license collections. When masking/compositing per house style, keep the credit.
Relationship to other skills
- Stacks with no-ai-slop and the house image style: real museum art is the anti-slop default for evocative imagery.
- Complements editorial-illustrations (generative claim-driven diagrams) and dataviz (charts) - those own explanatory graphics; this owns photographic/artwork/mood imagery.
- Feeds blog-publish, social-media-kit, weekly-ai-slide, deck and essay work at their image-sourcing step.
References (full per-museum recipes, verified)
references/_synthesis.md (cheat-sheet) + one file per museum (met.md, cleveland.md, smk.md, rijksmuseum.md, nga.md, artic.md, getty.md, smithsonian.md) with live-tested example URLs, field maps, and gotchas.
Files (cog-second-brain)
-
references
-
artic.md 6.2 KB
# Art Institute of Chicago — Open-Access Image API Recipe Verified: 2026-07-24 ## Base URL `https://api.artic.edu/api/v1` Official docs: https://api.artic.edu/docs/ (confirmed live and current). ## API Key **None required.** Fully open, anonymous access. - Rate limit: 60 requests/minute per IP (throttled if exceeded, no auth error). - Docs explicitly ask scrapers to self-throttle to ~1 req/sec and avoid parallel hammering — courtesy limit, not enforced by a key. ## Image License **CC0 1.0 (Creative Commons Zero)** for artwork data and images flagged `is_public_domain: true`. - Confirmed via live API response `info.license_text`: > "All other data in this response is licensed under a Creative Commons Zero (CC0) 1.0 designation and the Terms and Conditions of artic.edu." (The `description` field specifically is CC-BY 4.0 — everything else, including the image, is CC0.) - Suggested (not legally required for CC0, but good practice) attribution: "Digital image courtesy of the Art Institute of Chicago." ## Search Recipe ### Step 1 — Search for public-domain artworks that have an image ``` GET https://api.artic.edu/api/v1/artworks/search?query[term][is_public_domain]=true&fields=id,title,image_id,artist_display&limit=20 ``` - `query[term][is_public_domain]=true` filters to CC0/public-domain works. - `fields=...,image_id` is required — without explicitly requesting `image_id` it won't be in the payload. - Note: `is_public_domain=true` does not strictly guarantee `image_id` is non-null for every record (a few PD works lack digitized images) — check `image_id` is present/non-empty in each result before building an image URL. Safer alternative used by many integrations: also filter `query[exists][field]=image_id` or just skip results where `image_id` is null/empty. ### Step 2 — Build the full-resolution image URL from a result's `image_id` ``` https://www.artic.edu/iiif/2/{image_id}/full/843,/0/default.jpg ``` - `843,` = width 843px, height auto — this is the museum's own site default and the most-likely-cached size. - For actual full/max resolution, replace the size segment with `full` (i.e. `.../full/full/0/default.jpg`) — larger, uncached, slower. - The IIIF base (`https://www.artic.edu/iiif/2`) is also returned dynamically in every API response under `config.iiif_url` — prefer reading it from there over hardcoding, in case they ever move image hosting. ## Example URLs (real, from live query) **exampleApiUrl** (search): ``` https://api.artic.edu/api/v1/artworks/search?query[term][is_public_domain]=true&fields=id,title,image_id,artist_display&limit=3 ``` Verified response (2026-07-24) included, among others: - id 28560, "The Bedroom", Vincent van Gogh, image_id `6644829f-f292-c5c4-a73c-0356a6fdbf0d` - id 21023, "Buddha Shakyamuni Seated in Meditation (Dhyanamudra)", image_id `0675f9a9-1a7b-c90a-3bb6-7f7be2afb678` - id 20684, "Paris Street; Rainy Day", Gustave Caillebotte, image_id `f8fd76e9-c396-5678-36ed-6a348c904d27` - Total public-domain-with-fields matches in collection: 61,568 (paginated, 20,523 pages at limit=3) - `config.iiif_url` in response: `https://www.artic.edu/iiif/2` **exampleImageUrl** (from van Gogh's "The Bedroom", id 28560): ``` https://www.artic.edu/iiif/2/6644829f-f292-c5c4-a73c-0356a6fdbf0d/full/843,/0/default.jpg ``` Also confirmed the single-artwork endpoint works: ``` GET https://api.artic.edu/api/v1/artworks/28560?fields=id,title,image_id ``` → returned `{"data":{"id":28560,"title":"The Bedroom","image_id":"6644829f-f292-c5c4-a73c-0356a6fdbf0d"}, "info":{"license_text":"...CC0 1.0..."}}` ## Verification Notes - **JSON API: verified working.** Both the search endpoint and the single-artwork endpoint were fetched live and returned real, current data (61,568 public-domain artworks, correct image_id, correct CC0 license text in `info`). - **IIIF image URL: correct per official docs, but NOT independently confirmed as a working binary download by this research pass.** Direct `curl`/WebFetch requests to `www.artic.edu/iiif/2/.../default.jpg` were blocked with `HTTP 403` + `cf-mitigated: challenge` — this is Cloudflare's bot-management challenge on the `artic.edu` web/image domain, triggered regardless of User-Agent (tried default curl UA, desktop Chrome UA, iPhone Safari UA, and no UA — all 403'd identically, which points to TLS/JA3 fingerprinting rather than header-based blocking). This is a known behavior of Cloudflare-protected static asset domains and does not indicate the URL pattern itself is wrong — it is the officially documented pattern, and it is what the museum's own website uses to render images in a real browser. - I attempted to verify via a real headed browser (browser-harness/CDP) to bypass the bot check, but that tool requires one-time manual "Allow remote debugging" approval in the user's Chrome, which wasn't available in this pass. - **Practical implication for whoever builds against this recipe:** plain `curl`/`requests`/serverless-function fetches of the IIIF image may hit the same Cloudflare challenge. A real browser, a headless browser with a full JS-capable engine, or a fetch routed through something that passes Cloudflare's bot checks (residential proxy, or a client with a legitimate browser TLS fingerprint) is likely needed for automated bulk image downloading. The JSON metadata API (`api.artic.edu`) has no such protection and worked cleanly every time. ## Gotchas Summary - No API key, but self-throttle to ~1 req/sec for bulk work; hard limit 60 req/min anonymous. - Must explicitly request `image_id` via `fields=` — not returned by default. - Not all `is_public_domain:true` records have a non-null `image_id`; check before building the URL. - IIIF size param: `843,` = site-default/cached; `full` = max resolution but uncached/slower; can also request specific `{width},{height}` or `{width},` / `,{height}`. - `www.artic.edu` (the IIIF image host) sits behind Cloudflare bot management — scripted/curl-style requests to the image URLs get a 403 challenge page even though the URL pattern is correct and it works in normal browser use. Budget for this if building an automated downloader. - Attribution "Digital image courtesy of the Art Institute of Chicago" is good practice even though CC0 doesn't legally require it. -
cleveland.md 5.5 KB
# Cleveland Museum of Art — Open Access Image Recipe (VERIFIED) ## TL;DR - **Base URL:** `https://openaccess-api.clevelandart.org/api/artworks/` - **Key required:** No — fully open, no key/token. - **License:** CC0 (public domain), designated by CMA. No attribution required for CC0 works. - **Verified 2026-07-24:** live API call returned real JSON with CC0 records; the returned image URL resolves to a real JPEG (396.3KB, confirmed via fetch). ## API Basics - Docs: https://openaccess-api.clevelandart.org/ (interactive docs, last updated 2025-07-11) - GitHub mirror (full dataset dump, JSON/CSV, updated weekly): https://github.com/ClevelandMuseumArt/openaccess - Program overview / license page: https://www.clevelandart.org/open-access - Endpoints: - `GET /api/artworks/` — search/list artworks - `GET /api/artworks/{id}` — single artwork - `GET /api/creators/`, `/api/creators/{id}` - `GET /api/exhibitions/`, `/api/exhibitions/{id}` ## Key Query Params (for CC0 + image) | Param | Meaning | |---|---| | `cc0=1` | only works licensed CC0 (public domain, no restrictions) | | `has_image=1` | only artworks that have a web image asset | | `q=<term>` | free-text search | | `limit=<n>` | page size | | `skip=<n>` | pagination offset | | `copyrighted` (opposite of cc0) | filters copyrighted works instead — do NOT use if you want open images | Response field `share_license_status` will read `"CC0"`, `"Copyrighted"`, or `"Other"` — filter/inspect this if double-checking after the fact. ## Image URL Fields Each artwork's `images` object contains up to 3 renditions, each with `url`, `filename`, `filesize`, `width`, `height`: - `images.web` — JPEG, 900px longest side, 300dpi (good default "full-res-enough" download) - `images.print` — JPEG, 3400px longest side, 300dpi (high-res) - `images.full` — TIFF, variable dimensions/dpi (archival max-res) `alternate_images` may hold extra views in the same 3 renditions. Image CDN host: `https://openaccess-cdn.clevelandart.org/...` — direct hotlinkable JPEG/TIFF, no auth. ## Step-by-Step Recipe **(a) Search for public-domain artworks with images:** ``` GET https://openaccess-api.clevelandart.org/api/artworks/?cc0=1&has_image=1&limit=10 ``` Optional: add `&q=monet` (or any term) to narrow by keyword; add `&skip=10` to paginate. **(b) Get a full-res image URL from a result:** From each returned artwork object, read: ``` result.images.print.url # high-res JPEG (3400px) result.images.full.url # archival TIFF (max res), if present result.images.web.url # 900px JPEG, smaller/fast ``` No further auth or signing needed — the URL is directly downloadable. ## Concrete Verified Example **exampleApiUrl** (fetched live, returned real data): ``` https://openaccess-api.clevelandart.org/api/artworks/?cc0=1&has_image=1&limit=3 ``` Sample of what it returned: ```json [ { "id": 94979, "title": "Nathaniel Hurd", "share_license_status": "CC0", "images.web.url": "https://openaccess-cdn.clevelandart.org/1915.534/1915.534_web.jpg" }, { "id": 92937, "title": "Stag at Sharkey's", "share_license_status": "CC0", "images.web.url": "https://openaccess-cdn.clevelandart.org/1922.1133/1922.1133_web.jpg" } ] ``` **exampleImageUrl** (verified resolves to a real JPEG, 396.3KB, image/jpeg): ``` https://openaccess-cdn.clevelandart.org/1915.534/1915.534_web.jpg ``` (This is the "web" 900px rendition for accession 1915.534, "Nathaniel Hurd," CC0. Swap `_web` for `_print` in the same path pattern for the 3400px rendition where available — but always trust the JSON's `images.print.url` field over guessing the filename pattern, since not every record has a print/full tier.) ## License & Attribution - CMA designates open-access content as **CC0** — no copyright, no restriction, no attribution required. - Non-CC0 records exist in the same API (`share_license_status: "Copyrighted"` or `"Other"`) — always filter with `cc0=1` (or check the field) before treating an image as free-use. - CMA suggests (optional, not required for CC0) citing: Artist, Title, Date, Medium, Dimensions, Institution, Credit Line, Accession Number, URL. ## Gotchas - **No rate limit documented** — be a good citizen anyway (add delays for bulk scraping); no official published number. - **Not every record has all 3 image tiers.** Always read the actual `images.*.url` fields from the JSON rather than assuming a naming convention holds for every accession — some records only have `web`, not `print`/`full`. - **No IIIF image API** — CMA does NOT expose IIIF (no `/iiif/.../full/full/0/default.jpg` sizing syntax). Sizing is fixed to the 3 pre-rendered tiers (web/print/full), not parametric. - **`cc0` vs `copyrighted` params are opposite filters** — don't confuse them; use `cc0=1` for open images. - **Full dataset dump available on GitHub** (`ClevelandMuseumArt/openaccess`, JSON/CSV, updated weekly) if you want to avoid live API pagination for bulk work. - Image CDN (`openaccess-cdn.clevelandart.org`) is a separate host from the API host (`openaccess-api.clevelandart.org`) — don't assume same-origin for CORS purposes if building a browser app. ## Sources - https://openaccess-api.clevelandart.org/ (API docs) - https://www.clevelandart.org/open-access (license/program page) - https://github.com/ClevelandMuseumArt/openaccess (dataset mirror) - Live verification: `GET https://openaccess-api.clevelandart.org/api/artworks/?cc0=1&has_image=1&limit=3` fetched 2026-07-24, returned real JSON. - Live verification: `https://openaccess-cdn.clevelandart.org/1915.534/1915.534_web.jpg` fetched 2026-07-24, confirmed real JPEG (image/jpeg, 396.3KB). -
getty.md 9.9 KB
# The Getty — Open-Access Image Recipe (VERIFIED) ## TL;DR - **Base URL (Linked Art / JSON-LD API):** `https://data.getty.edu/museum/collection/` - **API key:** NOT required. All endpoints below returned `HTTP 200` with a plain unauthenticated `curl`. - **License:** Dataset/collection metadata is **CC0 1.0** (confirmed live in every record's `subject_to` "License for Collection Metadata" block, and in the official docs text). Images come from the **Getty Open Content Program** — most are CC0/public domain, but **not all**; each image reference should be checked (see Gotchas). - **Images:** served via **IIIF Image API** at `https://media.getty.edu/iiif/image/<image-id>/...` — full resolution, no key, `Access-Control-Allow-Origin: *`. - **Search:** there is **no keyword/full-text REST search endpoint** (the docs explicitly say so — see Gotchas). Discovery is via the public **SPARQL endpoint** `https://data.getty.edu/museum/collection/sparql`, which I ran live and got real results back. --- ## 1. What the API actually is Official docs: https://data.getty.edu/museum/collection/docs/ (a Nuxt SPA; scraped its raw payload to get the real text since WebFetch truncates it — Bash `curl` worked fine). - Model: **Linked.Art** (a CIDOC-CRM profile) + **JSON-LD**. - Entity types: `object`, `place`, `document`, `group`, `person`, `exhibition`, `activity`. - Record URL pattern: `https://data.getty.edu/museum/collection/<ENTITY_TYPE>/<ENTITY_ID>` — returns JSON-LD directly, no `Accept` header needed. - Change tracking: ActivityStreams feed at `https://data.getty.edu/museum/collection/activity-stream`. - Graph queries: public **SPARQL** endpoint + a browser UI at `https://data.getty.edu/museum/collection/sparql-ui`. - Images: Getty-wide **IIIF** API at `https://media.getty.edu` (Image API 2.1.1 + Presentation API 2.1.1/3). Docs quote (verbatim, from the rendered page): *"With some exceptions, the data available from this API is made available under [CC0]. Check the Usage Guidelines section... for more details."* and *"We currently don't provide a way to get a list of all of the objects or other entity types in the dataset... We also don't provide a way to download all the data in the dataset."* — i.e. **no bulk list, no keyword search REST endpoint**, by design, confirmed straight from their own docs. ## 2. Key requirement None. Every call in this recipe (`object` record fetch, SPARQL query, IIIF image fetch) succeeded with plain unauthenticated HTTP GET. No signup, no token, no rate-limit header observed. ## 3. License — precisely - **Dataset/metadata:** CC0 1.0 Universal, unconditionally, per Getty's own statement and confirmed live in every object record I pulled (`subject_to` → `Right` → `classified_as` → `http://creativecommons.org/publicdomain/zero/1.0/`, display name `"Public Domain"`, description `"No Copyright"`). - **Images (Exception #1 in the docs):** *"Many of the linked images are part of Getty's Open Content program and can also be used without permission under CC0 — but not all of the images are."* The docs say each image reference should carry its own rights block (`VisualItem` → `subject_to` → `classified_as` with the CC0 URI, or something else if restricted). In practice, on the two live records I sampled, that per-image rights sub-block was not populated (older TMS-migrated records) — so **don't assume every image is CC0 purely from the API response**; cross-check on the object's public page at `getty.edu` (Open Content items show a "Download" button) if you need certainty for a specific artwork, or prefer objects you already know are Open Content (e.g. van Gogh's *Irises*, used below). - **Written descriptions/biographies (Exception #2):** mixed — some CC BY 4.0, some third-party copyright. Same per-block check pattern (`referred_to_by[].subject_to[].classified_as[].id`). - Attribution is **not required** but requested/appreciated (no fixed credit-line string was given beyond "provided by the J. Paul Getty Museum"). - Official program description (from `getty.edu/projects/open-content-program`, verified fetch): *"Initiative granting free access to images of public domain artworks in Getty's collections."* ## 4. Search recipe (concrete, step-by-step) There is no `?q=keyword` REST search. Use the SPARQL endpoint to discover objects that have an image, then pull each object's full record for metadata + rights + all image links. **Step A — find N objects that have a representation image (SPARQL, GET, JSON by default):** ``` GET https://data.getty.edu/museum/collection/sparql?query=<url-encoded SPARQL> ``` SPARQL body used (real, tested): ```sparql PREFIX crm: <http://www.cidoc-crm.org/cidoc-crm/> SELECT ?obj ?label ?img WHERE { ?obj a crm:E22_Human-Made_Object . ?obj rdfs:label ?label . ?obj crm:P138i_has_representation ?img . } LIMIT 5 ``` This returns the object's URI, its label, and a ready-to-use IIIF image URL — no follow-up call needed for a quick image grab. **Step B — get the full record (metadata + rights + every image/IIIF manifest link) for one object:** ``` GET https://data.getty.edu/museum/collection/object/<uuid-from-step-A> ``` Look in the JSON for: - `representation[].id` → direct JPEG (deprecated field, still live, smaller res) - `subject_of[]` where `_label` = "IIIF Manifest URL" → full IIIF Presentation manifest (has every canvas/image + real per-canvas dims) - `subject_to[]` → the rights/license block (check `classified_as[].id` for the CC0 URI) **Step C — full-resolution image URL (IIIF Image API), from any `image-id` you have:** ``` https://media.getty.edu/iiif/image/<image-id>/full/full/0/default.jpg # full native resolution https://media.getty.edu/iiif/image/<image-id>/full/!600,/0/default.jpg # thumbnail, max 600px, aspect preserved https://media.getty.edu/iiif/image/<image-id>/<x,y,w,h>/<w,h>/0/default.jpg # cropped region (used on Getty's own site) ``` **Step D — filter for CC0/public domain with confidence:** either (a) rely on the per-image `subject_to`/`classified_as` block when populated, or (b) pick artworks you can confirm are in the Open Content Program via the human-facing collection page (`getty.edu/art/collection/object/...`), which flags Open Content items with a visible "Download" affordance. ## 5. Verified example URLs **exampleApiUrl** (SPARQL — real results, confirmed via curl, HTTP 200, default response is already `application/sparql-results+json` even with no Accept header): ``` https://data.getty.edu/museum/collection/sparql?query=PREFIX%20crm%3A%20%3Chttp%3A//www.cidoc-crm.org/cidoc-crm/%3E%0ASELECT%20%3Fobj%20%3Flabel%20%3Fimg%20WHERE%20%7B%0A%20%20%3Fobj%20a%20crm%3AE22_Human-Made_Object%20.%0A%20%20%3Fobj%20rdfs%3Alabel%20%3Flabel%20.%0A%20%20%3Fobj%20crm%3AP138i_has_representation%20%3Fimg%20.%0A%7D%20LIMIT%205 ``` Live sample of what it returns (first row): ```json {"obj":"https://data.getty.edu/museum/collection/object/84eb7a1d-f806-4da9-a0ed-77d6b355df7e", "label":"West Front, Looking North (84.XB.950.7.28)", "img":"https://media.getty.edu/iiif/image/f45355f5-8ae6-4813-bf7a-dc94997f76f0/full/full/0/default.jpg"} ``` Also a plain record fetch, no query params, well-known Open Content artwork (van Gogh's *Irises*, 90.PA.20): ``` https://data.getty.edu/museum/collection/object/c88b3df0-de91-4f5b-a9ef-7b2b9a6d8abb ``` **exampleImageUrl** (full resolution — verified with `curl -I`: `HTTP/2 200`, `content-type: image/jpeg`, `content-disposition: inline; filename="8c255d80-7382-46db-9fa8-892c0d37247e_9021x7122.jpg"`, `access-control-allow-origin: *`): ``` https://media.getty.edu/iiif/image/8c255d80-7382-46db-9fa8-892c0d37247e/full/full/0/default.jpg ``` (This is the Irises main image, 9021×7122px native.) ## 6. Gotchas - **No full-text/faceted search REST endpoint, no bulk listing/dump.** Explicitly stated as a current limitation in Getty's own docs ("it's on our roadmap"). SPARQL is the only queryable discovery mechanism from the API itself; the human search UI at `getty.edu/art/collection/search/` is the practical alternative for browsing by keyword and then feeding object IDs back into the JSON API. - **Per-image rights blocks are inconsistently populated.** The docs describe a `VisualItem.subject_to.classified_as` CC0 marker per image, but live sampled records didn't always carry it — don't blanket-assume every image is CC0 just because the record fetch succeeded; verify against the public object page or stick to artworks you know are Open Content. - **The `representation` field is marked deprecated** in favor of the `shows` → IIIF Manifest route for full-size images going forward (Getty's own deprecation note: future `representation` images "will be smaller than that currently offered"). The `full/full/0/default.jpg` IIIF Image API URL is the durable way to get max resolution regardless. - **IIIF sizing syntax:** `/full/full/...` = native size; `/full/!W,/...` = fit within width W preserving aspect (the `!` matters); `/full/W,/...` = force width W; region can be `x,y,w,h` pixels instead of `full` to crop. - **The `docs/` page is a client-rendered SPA** — plain `WebFetch`/simple scrapers get truncated boilerplate; had to pull the Nuxt `_payload.json` directly to get the real doc text. If automating doc reads again, target that payload endpoint, not the rendered HTML. - **No observed rate limiting** in this session, but Getty gives no published quota — be a reasonable citizen (this is a small non-profit-run API, not a commercial CDN). - Attribution not contractually required (CC0) but Getty explicitly asks to be told how you used the data (`MuseumCollections@getty.edu`) and offers a courtesy credit line: "J. Paul Getty Museum". ## Sources - https://data.getty.edu/museum/collection/docs/ (API docs, scraped via curl on `_payload.json`) - https://www.getty.edu/projects/open-content-program/ (Open Content Program description) - https://www.getty.edu/projects/open-data-apis/ (open data/APIs overview) - Live verified endpoints: `data.getty.edu/museum/collection/object/*`, `data.getty.edu/museum/collection/sparql`, `media.getty.edu/iiif/image/*` -
met.md 6.2 KB
# The Metropolitan Museum of Art — Open Access Image Recipe Status: **VERIFIED** (both the search endpoint and the resulting image URL were fetched live and returned real data / HTTP 200). ## Base URL `https://collectionapi.metmuseum.org/public/collection/v1/` Docs: https://metmuseum.github.io/ (official, GitHub Pages) Initiative overview: https://www.metmuseum.org/hubs/open-access Repo: https://github.com/metmuseum/openaccess ## API key **None required.** The docs state explicitly: "At this time, we do not require API users to register or obtain an API key to use the service." Rate limit: **80 requests/second** (documented; be polite and add a small delay/backoff for bulk jobs anyway). Contact for questions: `openaccess@metmuseum.org`. ## License **CC0 (Creative Commons Zero)** for objects flagged `isPublicDomain: true` — The Met has waived all copyright and related/neighboring rights on this subset of the dataset (both the metadata and the associated images). No attribution is legally required, though crediting "The Metropolitan Museum of Art" is good practice / house style. Caveat: not every object in the collection is CC0 — only ones where `isPublicDomain` is `true`. Objects where it's `false` still have metadata returned by the API but the image (if any) is NOT open-licensed for reuse — always check the flag per-object, don't assume every API result is free to use. ## Endpoints used in this recipe | Endpoint | Purpose | |---|---| | `GET /search?...` | Search for object IDs matching filters (query term, `hasImages`, `isPublicDomain`, department, date range, etc.) | | `GET /objects/{objectID}` | Full record for one object: title, artist, `isPublicDomain`, `primaryImage`, `primaryImageSmall`, rights fields | | `GET /departments` | List of department IDs/names, for scoping search | Confirmed search params (from live docs): `q`, `isHighlight`, `title`, `tags`, `departmentId`, `isOnView`, `artistOrCulture`, `medium`, `hasImages`, `geoLocation`, `dateBegin` + `dateEnd` (must be given together). **`isPublicDomain` is also a live, working param** even though it isn't prominently listed in every doc rendering — confirmed by direct test below. ## Step-by-step recipe 1. **Search** for public-domain artworks with images: ``` GET https://collectionapi.metmuseum.org/public/collection/v1/search?q=<term>&hasImages=true&isPublicDomain=true ``` Response: `{"total": N, "objectIDs": [id1, id2, ...]}`. If `objectIDs` is null, no matches — try a broader `q` or drop a filter. 2. **Fetch one object's full record:** ``` GET https://collectionapi.metmuseum.org/public/collection/v1/objects/{objectID} ``` Check `isPublicDomain === true` before reuse (belt-and-suspenders even though the search already filtered on it — the object may have been re-flagged since indexing). 3. **Grab the image URL directly from the object JSON** — no extra IIIF/image-service call needed: - `primaryImage` — full original resolution JPEG (can be very large, several MB, up to ~8MB+ observed). - `primaryImageSmall` — "web-large" size, smaller JPEG, good default for web use. - `additionalImages` — array of extra full-res image URLs if the object has more than one photographed view. No IIIF Image API / sizing-parameter syntax is involved — Met just serves static JPEGs at fixed pre-rendered sizes (`original`, `web-large` in the URL path), not a dynamic IIIF resizer. ## Example URLs (live-tested) **Search:** ``` https://collectionapi.metmuseum.org/public/collection/v1/search?q=sunflowers&hasImages=true&isPublicDomain=true ``` Verified live: returned `"total": 40` and a real `objectIDs` array (e.g. 544320, 310453, 200668, 437261, 824771, 36225, ... 436535 among later results is a different van Gogh but 436535 below was pulled independently for the full example). **Object record:** ``` https://collectionapi.metmuseum.org/public/collection/v1/objects/436535 ``` Verified live via curl — returns real JSON: - `objectID`: 436535 - `title`: "Wheat Field with Cypresses" - `artistDisplayName`: Vincent van Gogh - `department`: European Paintings - `isPublicDomain`: `true` - `primaryImage`: `https://images.metmuseum.org/CRDImages/ep/original/DP-42549-001.jpg` - `primaryImageSmall`: `https://images.metmuseum.org/CRDImages/ep/web-large/DP-42549-001.jpg` **Full-res image URL (verified):** ``` https://images.metmuseum.org/CRDImages/ep/original/DP-42549-001.jpg ``` `curl -sI` on this URL returned `HTTP/2 200`, `content-type: image/jpeg`, `content-length: 8291194` (~8.3MB) — confirmed real, downloadable, full-resolution JPEG. ## Gotchas - **Rate limit 80 req/sec** per the docs — fine for interactive/scripted use, but throttle bulk crawls (e.g. iterate all `/objects`) to be a good citizen; the server is fronted by Imperva/Incapsula (visible in response headers) which may rate-limit/challenge aggressive traffic beyond the documented cap. - **No CORS problem** — `access-control-allow-origin: *` is set on the image CDN (`images.metmuseum.org`), so these URLs are directly usable from browser JS (e.g. `<img>` tags, canvas/fetch) without a proxy. - **`objects` bulk listing endpoint** (`/objects?metadataDate=...&departmentIds=...`) returns *all* object IDs for the filter, not just public-domain/has-image ones — you still need to check `isPublicDomain` + `primaryImage` (non-empty) per object, or better, use `/search?hasImages=true&isPublicDomain=true` up front to pre-filter. - **`primaryImage` can be empty string** even for public-domain objects that haven't been photographed — always check `primaryImage !== ""` (or non-null) before using. - **No IIIF sizing params** — unlike some museum APIs (e.g. Smithsonian, some Europeana sources), the Met does not expose a dynamic image resize API. You get exactly two fixed sizes baked into the URL (`original`, `web-large`) plus optional `additionalImages`. If you need a specific pixel size, resize client-side after download. - **Attribution not legally required (CC0)** but recommended: "Image courtesy of The Metropolitan Museum of Art" / link back to the object's page (`https://www.metmuseum.org/art/collection/search/{objectID}`). - **`rightsAndReproduction` field is often blank** even on legitimate public-domain records — don't treat an empty rights field as a red flag; `isPublicDomain: true` is the authoritative signal. -
nga.md 8.1 KB
# National Gallery of Art (Washington) — Open Access Image Recipe Status: VERIFIED (live IIIF fetch succeeded, real image bytes returned). ## Base URL(s) - **Data/metadata**: CSV dumps in the GitHub repo `github.com/NationalGalleryOfArt/opendata` (data lives under `/data/*.csv`, fetch raw via `raw.githubusercontent.com/NationalGalleryOfArt/opendata/main/data/<file>.csv`). No search API — this is a bulk CSV export, updated frequently (daily per the repo docs). - **Image server (IIIF Image API 2.0, level1)**: `https://api.nga.gov/iiif/{uuid}/{region}/{size}/{rotation}/{quality}.{format}` - `info.json` at `https://api.nga.gov/iiif/{uuid}/info.json` confirms native width/height and supported sizes/qualities. There is no public REST/search API beyond the CSV files + IIIF image server — no `api.nga.gov` object-search endpoint documented. Treat this as "CSV catalog + IIIF images," not a queryable API. ## Key requirement **None.** No API key, no auth, no rate-limit header observed. CSVs are public GitHub raw files; the IIIF image server responds with `access-control-allow-origin: *` and no auth challenge. ## License **CC0-1.0** (Creative Commons Zero / public domain dedication) for the dataset itself — repo `LICENSE` + README state NGA "waives any copyright or related rights that it might have in this dataset." Attribution is *requested* (not required) for research use citing "National Gallery of Art Open Data Program." Per-image rights are **not uniformly CC0** — see gotcha below on the `openaccess` flag. Only rows where `published_images.openaccess = 1` are the institution's actual open-access (full-resolution, reuse-cleared) images; `openaccess = 0` rows are still shown via IIIF but resolution-capped (see `maxpixels`), meaning rights-restricted/third-party-copyright works. ## Relevant CSV files (from `data/` dir) - `objects.csv` — one row per artwork: `objectid, uuid, title, attribution (artist), displaydate, medium, dimensions, classification, creditline, departmentabbr, wikidataid, ...`. No explicit "is public domain" column at the object level. - `published_images.csv` — one row per digitized image: `uuid, iiifurl, iiifthumburl, viewtype, sequence, width, height, maxpixels, openaccess, depictstmsobjectid, assistivetext`. This is the file that actually gates open access: - `iiifurl` = the IIIF base identifier URL (append IIIF path params to get an image). - `iiifthumburl` = pre-built 200×200 thumbnail convenience URL. - `openaccess` = `1` → full resolution downloadable via IIIF `full/full` or `full/max`; `0` → IIIF still serves an image but the server enforces a `maxpixels` ceiling (e.g. 900px on the long edge) — these are NOT open access. - `depictstmsobjectid` = join key back to `objects.csv.objectid` for title/artist metadata. - `assistivetext` = an auto-generated alt-text description of the image (handy bonus field). ## Search recipe ### (a) Find public-domain artworks that have an image No live search API — do it against the CSVs (download or stream them): 1. Fetch `https://raw.githubusercontent.com/NationalGalleryOfArt/opendata/main/data/published_images.csv`. 2. Filter rows where `openaccess == "1"` (string "1") AND `viewtype == "primary"` (primary image, not alternate crops/details) → these are full-res, reuse-cleared images. 3. Take the `depictstmsobjectid` from each surviving row and join against `objects.csv.objectid` to pull `title`, `attribution` (artist), `displaydate`, `medium`, `classification`. 4. `uuid` (or equivalently `iiifurl`'s trailing path segment) is the identifier to build image-request URLs. No pagination/rate-limit concerns since it's a flat file — just stream/filter client-side (Python `csv`/`pandas`, or `curl` + `awk`/`grep` for quick checks). Both CSVs are large (130k+ objects, more image rows); use streaming reads, not full in-memory loads if resource-constrained. ### (b) Get a full-resolution downloadable image URL from a result Given a `uuid` (e.g. from `published_images.csv`, row with `openaccess=1`): ``` GET https://api.nga.gov/iiif/{uuid}/full/full/0/default.jpg ``` or equivalently ``` GET https://api.nga.gov/iiif/{uuid}/full/max/0/default.jpg ``` Both returned the object's native full resolution in testing (3365×4332 px, 2.57 MB JPEG for the example below). Use `info.json` first if you want to confirm native dimensions or pick a specific IIIF `size` token (e.g. `!1600,1600` to cap the long edge) instead of `full`. ## Example URLs (live-tested 2026-07-24) - **exampleApiUrl** (metadata, first bytes of the open-access image catalog): `https://raw.githubusercontent.com/NationalGalleryOfArt/opendata/main/data/published_images.csv` First data row: `uuid=00007f61-4922-417b-8f27-893ea328206c, iiifurl=https://api.nga.gov/iiif/00007f61-4922-417b-8f27-893ea328206c, openaccess=1, depictstmsobjectid=17387` (join to `objects.csv` objectid 17387 for title/artist). - **info.json check**: `https://api.nga.gov/iiif/00007f61-4922-417b-8f27-893ea328206c/info.json` → returned valid IIIF Image API 2.0 descriptor, native size 3365×4332, level1 profile, jpg format, supports `sizeAboveFull`. - **exampleImageUrl** (full-resolution, downloadable): `https://api.nga.gov/iiif/00007f61-4922-417b-8f27-893ea328206c/full/full/0/default.jpg` → `HTTP/2 200`, `content-type: image/jpeg`, `content-length: 2569142` (2.57 MB), served via Cloudflare + IIPImage, `access-control-allow-origin: *`. ## Verification performed - Fetched `published_images.csv` raw (byte-range) — confirmed real header + rows, confirmed `openaccess`/`maxpixels` semantics by comparing an `openaccess=1` row (no maxpixels cap) against an `openaccess=0` row (maxpixels=900). - Fetched `objects.csv` raw (byte-range) — confirmed schema and that title/artist metadata lives here, joined via `objectid`. - `curl` GET `info.json` for a real uuid — valid IIIF descriptor returned. - `curl -I` (HEAD) on `full/full/0/default.jpg` and `full/max/0/default.jpg` — both HTTP 200, image/jpeg, multi-MB content-length, i.e. genuinely full resolution, not a redirect or error page. - `curl -I` on the `iiifthumburl` pattern (`full/!200,200/0/default.jpg`) — HTTP 200, small JPEG (9.7 KB), confirming the thumbnail convenience URL also works. **verified = true.** ## Gotchas - **`openaccess` flag is per-image, not per-object.** Always filter on it — do not assume every row in `published_images.csv` is reuse-cleared. Rows with `openaccess=0` are capped by IIIF server-side (`maxpixels`, e.g. 900px long edge) — these are rights-restricted (e.g. copyrighted contemporary works, loans) and should be excluded from a "CC0 image" pipeline. - **IIIF profile is level1** — supports `sizeAboveFull` per the `info.json` profile, which is why `full/full` and `full/max` both return native resolution rather than erroring; not all IIIF servers allow this (level0 servers only serve pre-baked sizes). - **No live object/artwork search API.** Anyone wanting "search by artist/title/date" must filter the CSV client-side (or load into SQLite/DuckDB — the repo ships `sql_tables/` schema helpers for exactly this). Don't assume a `?q=` REST endpoint exists. - **CSV files are large** (130k+ objects across multiple linked tables — `objects.csv`, `published_images.csv`, `constituents.csv`, `objects_constituents.csv`, etc. per the repo's `data/` dir) — use streaming parses or DuckDB/SQLite rather than loading everything into memory naively. - **Rate limits**: none encountered on either GitHub raw or `api.nga.gov` in this test; Cloudflare fronts `api.nga.gov` and sets `__cf_bm` cookies but did not block repeated HEAD requests. Be a reasonably polite bulk client anyway (the data updates frequently — no need to re-download the whole CSV more than daily). - **Attribution**: not legally required (CC0) but NGA requests citing "National Gallery of Art Open Data Program" for datasets built on this data; for images, credit lines are in `objects.csv.creditline` per object (nice to surface even though not mandatory). - **`assistivetext`** in `published_images.csv` is a free, pre-generated alt-text/description per image — useful if you need accessible captions without running your own vision model. -
rijksmuseum.md 7.7 KB
# Rijksmuseum — Open-Access Image Recipe (VERIFIED 2026-07-24) ## TL;DR - **Starting hint was stale.** `rijksmuseum.nl/api/en/collection` (the old apikey-based "Collection API") is **deprecated**. The current, live, documented API is the **Linked Art Search API** at `data.rijksmuseum.nl/search/collection`. - **No API key needed** for the current API. It's fully open. - Images come via a separate **IIIF** image service (`iiif.micr.io`), reached by walking the linked-art graph from a search result. - License: **CC0 / Public Domain** for the vast majority of digitized objects (confirmed on the test object). Attribution is requested but not legally required for CC0/PD items. ## Base URL - Search: `https://data.rijksmuseum.nl/search/collection` - Object resolver (linked data / JSON-LD): `https://id.rijksmuseum.nl/{objectId}` - Image (IIIF, via Micrio): `https://iiif.micr.io/{imageId}/{region}/{size}/{rotation}/{quality}.{format}` ## API Key **Not required.** The new Search API (`data.rijksmuseum.nl`) is public, no registration, no key, no auth header. (The legacy `rijksmuseum.nl/api/en/collection` API *did* require a free key via account registration + profile settings — but that endpoint is marked DEPRECATED in the current docs and should not be used for new integrations.) ## License - **CC0 / Public Domain** for most digitized collection objects — confirmed directly on the test object's `VisualItem` record: `subject_to` → `classified_as` → `https://creativecommons.org/publicdomain/mark/1.0/` ("Public Domain") and a nested `subject_of` rights block citing `https://creativecommons.org/publicdomain/zero/1.0/` (CC0). - Some items are CC BY 4.0 (attribution required) or fully copyrighted/restricted — always read the object's own rights block rather than assuming. - Museum's policy page (`data.rijksmuseum.nl/policy/`): attribution is *requested* ("kindly ask you to credit the Rijksmuseum") but not legally mandatory for CC0/PD works. ## Search Recipe (step by step) ### (a) Search for public-domain artworks that have images 1. Call `GET https://data.rijksmuseum.nl/search/collection?type=painting&imageAvailable=true` - `imageAvailable=true` filters to objects with a digital reproduction. - Other useful params: `creator`, `creationDate` (wildcards `*`/`?`), `description`, `material`, `technique`, `title`, `type`, `objectNumber`, `memberOfSetId`, `aboutActor`. Repeat a param to OR multiple values (e.g. two `material=`). - There is **no direct license/CC0 filter param** — the API doesn't expose rights as a search facet. In practice, filter/verify license per-object (see step 4 below) since almost everything with `imageAvailable=true` in the general collection is Public Domain/CC0. - Response is a Linked-Art `OrderedCollectionPage`: `orderedItems: [{id: "https://id.rijksmuseum.nl/{objectId}", type: "HumanMadeObject"}, ...]`, plus `partOf.totalItems` and a `next.id` URL (contains an opaque `pageToken`) for pagination — just follow `next.id` verbatim for the next page. Page size is capped at 100. 2. Pick an object id from `orderedItems`, e.g. `https://id.rijksmuseum.nl/200105887`. ### (b) Get a full-resolution downloadable image URL from a result 3. `GET https://id.rijksmuseum.nl/{objectId}` (Accept: application/json) → JSON-LD `HumanMadeObject` record. Find the `shows` array → `{id: "https://id.rijksmuseum.nl/{visualItemId}", type: "VisualItem"}`. 4. `GET https://id.rijksmuseum.nl/{visualItemId}` → `VisualItem` record. Check: - `subject_to[].classified_as[].id` for the rights statement (look for `creativecommons.org/publicdomain/...`). - `digitally_shown_by` → `{id: "https://id.rijksmuseum.nl/{digitalObjectId}", type: "DigitalObject"}`. 5. `GET https://id.rijksmuseum.nl/{digitalObjectId}` → `DigitalObject` record. Its `access_point[0].id` is the **ready-to-use full-resolution IIIF image URL**, already in the form `https://iiif.micr.io/{imageId}/full/max/0/default.jpg` — no further construction needed, just use it as-is. ### IIIF sizing (if you want other resolutions/crops) Template: `https://iiif.micr.io/{imageId}/{region}/{size}/{rotation}/{quality}.{format}` - `region`: `full` (whole image) or `x,y,w,h` pixel box - `size`: `max` (native full res), `!2000,2000` (fit within box), `800,` (width 800, auto height) - `rotation`: `0` normally - `quality`: `default` (or `gray`, `bitonal`) - `format`: `jpg`, `png`, `webp`, etc. Docs: `data.rijksmuseum.nl/docs/iiif/image` (IIIF Image API), `data.rijksmuseum.nl/docs/iiif/presentation` (manifests), `data.rijksmuseum.nl/docs/iiif/` (overview). ## Concrete Example (fully verified live, 2026-07-24) **exampleApiUrl** (search): ``` https://data.rijksmuseum.nl/search/collection?type=painting&imageAvailable=true ``` Verified via `curl` → HTTP 200, valid JSON, `totalItems: 4916`, first result `https://id.rijksmuseum.nl/200100988`. **Object chain used for the image example:** - Object: `https://id.rijksmuseum.nl/200105887` → "Cat at Play" (Katjesspel), Henriëtte Ronner-Knip, c.1860-1878, object number SK-A-3089 - VisualItem: `https://id.rijksmuseum.nl/202105887` → rights = Public Domain / CC0 - DigitalObject: `https://id.rijksmuseum.nl/5001087555671055286110` → `access_point[0].id` **exampleImageUrl**: ``` https://iiif.micr.io/YAxov/full/max/0/default.jpg ``` Verified via `curl -I` and full download: - HTTP/2 200, `content-type: image/jpeg`, `content-length: 1,463,989 bytes` - Actual pixel dimensions: **3720×2696**, baseline JPEG - `access-control-allow-origin: *` (CORS-open, safe to hotlink/fetch client-side) - Served via Cloudflare, `cache-control: public, max-age=31536000` (1yr cache — fine to cache aggressively) ## Gotchas - **Two generations of API coexist in search results/docs.** Don't follow the old `rijksmuseum.nl/api/en/collection` (`key=...&imgonly=True` style) — it's the deprecated Collection API. Use `data.rijksmuseum.nl/search/collection` instead. - **No single-call shortcut for the image.** Unlike some museum APIs (e.g. a flat `webImage.url` field), Rijksmuseum's linked-art model requires **3 sequential GETs** (object → VisualItem → DigitalObject) to reach the actual image URL. Budget for that in any pipeline (or cache the chain). - **No license filter in search params.** `imageAvailable=true` gets you images, not necessarily license — verify CC0/PD per object via the VisualItem's `subject_to` block if you need to be strict (though in practice public-collection paintings/prints are overwhelmingly Public Domain/CC0). - **Library/archive records excluded** from the Search API (museum objects only). - **Pagination**: don't hand-build `pageToken` — always follow the exact `next.id` URL returned in the response. - **Rate limits**: not documented/published for the new API; no auth means no per-key throttling was observed in this test, but be a good citizen (no aggressive parallel hammering). - **Attribution**: not legally required for CC0/PD, but the museum requests a credit line ("Rijksmuseum") — cheap to add and avoids any ambiguity for CC BY items mixed into broader queries. - **IIIF host is `iiif.micr.io`** (third-party Micrio infrastructure), not `data.rijksmuseum.nl` itself — don't assume same-origin/rate-limit policy as the main API. ## Sources - https://data.rijksmuseum.nl/docs/ (API overview) - https://data.rijksmuseum.nl/docs/search (Search API reference) - https://data.rijksmuseum.nl/docs/api/collection (old Collection API — marked DEPRECATED) - https://data.rijksmuseum.nl/docs/iiif/ , /docs/iiif/image , /docs/iiif/presentation (IIIF docs) - https://data.rijksmuseum.nl/policy/ (licensing/rights policy) - Live verified via `curl`: `data.rijksmuseum.nl/search/collection`, `id.rijksmuseum.nl/200105887`, `id.rijksmuseum.nl/202105887`, `id.rijksmuseum.nl/5001087555671055286110`, `iiif.micr.io/YAxov/full/max/0/default.jpg` -
smithsonian.md 7.4 KB
# Smithsonian Open Access API — Verified Recipe Status: **VERIFIED** (live queries executed 2026-07-24, all succeeded with real data and a resolvable full-res image). ## Base URL ``` https://api.si.edu/openaccess/api/v1.0/ ``` Key endpoints: - `search` — full-text/field search across ~7.5M+ records (Solr-backed) - `content/{id}` — fetch one record by its `id` (from a search result) - `stats` — collection unit counts - `metadata/v2.0/terms/{category}` — controlled-vocabulary term lists All requests go through **api.data.gov** as the API gateway (hostname is `api.si.edu` but auth/quota is api.data.gov's). ## API Key **Required: yes**, via the query param `api_key`. - Free signup: https://api.data.gov/signup/ (name + email, key emailed immediately — standard api.data.gov self-serve flow, no approval wait, no cost). - For quick testing without signing up, the shared `DEMO_KEY` works (used below) but is rate-limited much harder. - Rate limits: `DEMO_KEY` = 30 requests/hour/IP; a registered personal key = 1,000 requests/hour. (Per api.data.gov standard tiers; Smithsonian doesn't publish a separate limit.) ## License **CC0 1.0 Universal (public domain dedication)** for everything tagged Open Access. No attribution legally required, though crediting "Smithsonian Institution" is good practice. Two places the CC0 flag shows up in the JSON, both worth checking: - `content.descriptiveNonRepeating.metadata_usage.access` = `"CC0"` (record-level) - `content.descriptiveNonRepeating.online_media.media[].usage.access` = `"CC0"` (per-image-level — check this one, since a record can be CC0 but an individual attached media item can carry different rights) Not every one of the Smithsonian's ~157M total records is Open Access — only records with `metadata_usage.access: "CC0"` are released for unrestricted reuse. Filtering on this field (or the `online_media_type` field, see below) is what separates "any record" from "downloadable public-domain asset." ## Search Recipe ### (a) Search for public-domain artworks that have images ``` GET https://api.si.edu/openaccess/api/v1.0/search ?q=<TERMS> AND online_media_type:Images AND unit_code:<UNIT> &rows=10 &start=0 &api_key=<YOUR_KEY> ``` - `online_media_type:Images` — restricts to records that have at least one attached image-type media object. (This is the field that matters; do **not** rely on adding `cc0:CC0` alone — that clause was accepted by the query parser without error but did not reliably filter, see Gotchas.) - `unit_code:<UNIT>` — scope to one museum, e.g. `SAAM` (Smithsonian American Art Museum), `NPG` (National Portrait Gallery), `FSG` (Freer|Sackler), `CHNDM` (Cooper Hewitt), `NMNHBIRDS`, etc. Omit for cross-collection search. - Free-text `q=` terms combine with `AND`/`OR` Solr syntax, e.g. `q=sunflower AND online_media_type:Images`. - **After getting results, always check `content.descriptiveNonRepeating.online_media.media[].usage.access == "CC0"`** on each hit before treating its image as free-to-use — some non-Open-Access records still surface in a broad search. ### (b) Get a full-resolution downloadable image URL from a result For each hit, walk: `content.descriptiveNonRepeating.online_media.media[]` — an array (a record can have multiple images). Each media object has: - `media.usage.access` — CC0 check (per above) - `media.content` — an IDS delivery-service URL, e.g. `https://ids.si.edu/ids/deliveryService?id=<idsId>` — **calling this with no size param returns the full-resolution original** (verified: 1.8MB JPEG). - `media.resources[]` — explicit named download links when the record has them pre-generated: `"High-resolution TIFF"`, `"High-resolution JPEG"` (with `width`/`height` in pixels), `"Screen Image"`, `"Thumbnail Image"`. Not every record has the high-res TIFF/JPEG resources array populated — some only expose `Screen Image`/`Thumbnail Image`, in which case use the `deliveryService` content URL directly for full res. **IIIF-style resizing**: append `&max=<pixels>` to the `deliveryService` URL to cap the longest edge, e.g. `&max=2000` (verified: dropped a 1.8MB image to 828KB at max=2000). Omit `max` entirely for the original full-size file. ## Example URLs (both verified live, 2026-07-24) **exampleApiUrl** (search — Smithsonian American Art Museum CC0 images): ``` https://api.si.edu/openaccess/api/v1.0/search?q=cc0:CC0%20AND%20online_media_type:Images%20AND%20unit_code:SAAM&api_key=DEMO_KEY&rows=5 ``` Verified response: HTTP 200, `rowCount: 12999`, returned real SAAM artwork records (e.g. "A Chiefe Herowan," object id `saam_1985.66.403_410`, record link https://americanart.si.edu/collections/search/artwork/?id=18695) each carrying `metadata_usage.access: "CC0"` and an `online_media.media[]` array. **exampleImageUrl** (full-resolution, from a National Museum of Natural History Birds specimen record returned by the broader query `q=cc0:CC0 AND online_media_type:Images`): ``` https://ids.si.edu/ids/download?id=NMNH-vol.090_449776-449800.jpg ``` Verified: `curl -sL` → HTTP 200, `image/jpeg`, 23,659,074 bytes, resolves (307 redirect) to `https://smithsonian-open-access.s3-us-west-2.amazonaws.com/media/nmnh/NMNH-vol.090_449776-449800.jpg`. `usage.access: "CC0"` on the media object. Alternate example (SAAM artwork, screen-res since no TIFF resource present, still CC0): ``` https://ids.si.edu/ids/deliveryService?id=SAAM-1985.66.403410_1 ``` Verified: HTTP 200, `image/jpeg`, 1,833,356 bytes. ## Gotchas - **`cc0:CC0` as a query clause is not a reliable filter.** It doesn't error, but adding it to a query did not change result counts predictably in testing — always independently verify `metadata_usage.access` / `media.usage.access` == `"CC0"` in the returned JSON rather than trusting the query string to have filtered correctly. - **`online_media_type:Images` also isn't a guaranteed hard filter on every unit.** Several paintings-category test queries returned zero records with a populated `online_media` object despite the filter term, while bird/NMNH and SAAM units reliably returned populated media. Best practice: request extra rows and skip any result whose `descriptiveNonRepeating` lacks an `online_media` key. - **Rate limits are api.data.gov's, not Smithsonian's**: DEMO_KEY = 30 req/hr/IP (easy to exhaust in a scripting loop — get a real key for anything beyond a handful of test calls). - **IDS delivery URLs sometimes 307-redirect to S3** (`smithsonian-open-access.s3-us-west-2.amazonaws.com`) — follow redirects (`curl -L`) or your HTTP client's default redirect-follow. - **Attribution not legally required** under CC0, but Smithsonian's own guidance asks for a credit line where practical (e.g. "Smithsonian American Art Museum"). - **Two rights fields to reconcile**: record-level `metadata_usage.access` and per-media `media.usage.access` can theoretically diverge (a CC0 record could contain a rights-restricted third-party image) — always check the media-level flag before using a specific image. - Full docs referenced (not independently fetchable due to JS-rendered pages, but corroborated via GitHub/Postman/search): https://www.si.edu/openaccess/devfaq, https://edan.si.edu/openaccess/docs/, https://github.com/Smithsonian/OpenAccess ## Sources - https://www.si.edu/openaccess/faq - https://edan.si.edu/openaccess/docs/ - https://github.com/Smithsonian/OpenAccess - https://github.com/Smithsonian/smithsonian-openaccess (Python client) - https://api.data.gov/signup/ - Live API responses captured via curl, 2026-07-24 (this session) -
smk.md 7.5 KB
# SMK (National Gallery of Denmark) — Open Access Image Recipe Status: **VERIFIED working** (2026-07-24, live curl + WebFetch checks against the real API). ## Base URL ``` https://api.smk.dk/api/v1 ``` Official docs (Swagger UI, embedded OpenAPI 3.0.3 spec — page itself is JS-rendered but the spec JSON is inlined in `swagger-ui-init.js`): - Human docs: https://api.smk.dk/api/v1/docs/ - Article: https://www.smk.dk/en/article/smk-api/ - Contact: smkapi@smk.dk ## API key **None needed.** Confirmed empirically — 8+ unauthenticated GET requests in a row all returned HTTP 200 with no `Authorization` header, no key, no `X-Api-Key`. The API is explicitly described as "free to use." No signup, no throttling encountered in this test burst. ## License Individual objects carry a `rights` field. For public-domain works it is: ``` "rights": "https://creativecommons.org/publicdomain/mark/1.0/" ``` i.e. **Public Domain Mark 1.0** (SMK calls it "Public Domain" on their license page, https://www.smk.dk/en/license/public-domain/ — not literally the CC0 waiver, but functionally equivalent: no copyright restrictions, reuse/modify/redistribute freely, attribution not mandatory but good practice — credit "SMK / Statens Museum for Kunst"). Filter for it with the `public_domain:true` facet filter (below) rather than parsing `rights` per item. ## The gotcha the starting hint got wrong The hinted query shape `filters=[public_domain:true][has_image:true]` (both brackets concatenated into ONE `filters=` value) is **accepted by the server (200 OK) but silently wrong** — it does not AND the two conditions. Verified by comparing `found` counts: | Query | `found` | |---|---| | `filters=[public_domain:true]` only | 150,301 | | `filters=[has_image:true]` only | 54,393 | | `filters=[public_domain:true][has_image:true]` (single concatenated value) | 150,301 (wrong — ignored the second bracket) | | `filters=[public_domain:true]&filters=[has_image:true]` (**repeated param, one bracket each**) | **39,479** (correct intersection) | **Rule: pass `filters` as a repeated query parameter, one `[field:value]` bracket per occurrence, not concatenated.** ## Recipe ### (a) Search for public-domain artworks that have an image ``` GET https://api.smk.dk/api/v1/art/search ?keys=* &filters=[public_domain:true] &filters=[has_image:true] &offset=0 &rows=10 ``` URL-encoded (paste-able): ``` https://api.smk.dk/api/v1/art/search?keys=*&filters=%5Bpublic_domain:true%5D&filters=%5Bhas_image:true%5D&offset=0&rows=10 ``` Other useful params (from the OpenAPI spec, `/art/search` GET): - `keys` (required) — search keywords; `*` = match all. - `rows` — page size, max 2000, default 10. - `offset` — pagination start. - `fields` — restrict returned fields (array param). - `sort` — sort field, default relevance. - Other facet filters follow the same `[field:value]` bracket syntax, e.g. `[has_3d_file:true]`, `[on_display:true]`, `[collection:...]`. ### (b) Get a full-resolution downloadable image from a result item Each item in `items[]` already carries everything needed — no second API call required: - `image_native` — direct downloadable full-res JPEG URL (this is the one to use for "download the image"). - `image_thumbnail` — pre-sized ~1024px-wide JPEG via the IIIF thumb server. - `image_iiif_id` / `image_iiif_info` — the raw IIIF Image API base + `info.json`, for requesting any custom size/region/rotation via standard IIIF syntax `{iiif_id}/{region}/{size}/{rotation}/{quality}.{format}`. - `image_width` / `image_height` / `image_size` — native pixel dimensions and byte size. Example from a verified live item (object KKS5261, "Augustus og den tiburtinske sibylle"): ```json { "id": "1170000001_object", "object_number": "KKS5261", "public_domain": true, "rights": "https://creativecommons.org/publicdomain/mark/1.0/", "image_width": 4992, "image_height": 6287, "image_size": 32476791, "image_thumbnail": "https://iip-thumb.smk.dk/iiif/jp2/qz20sx771_kks5261.tif.jp2/full/!1024,/0/default.jpg", "image_native": "https://api.smk.dk/api/v1/download/W3siaW1nX3VybCI6Imh0dHBzOi8vaWlwLnNtay5kay9paWlmL2pwMi9xejIwc3g3NzFfa2tzNTI2MS50aWYuanAyL2Z1bGwvZnVsbC8wL25hdGl2ZS5qcGciLCJwdWJsaWNfZG9tYWluIjp0cnVlLCJvYmplY3RfbnVtYmVyIjoiS0tTNTI2MSIsIm51bSI6IiJ9XQ==/KKS5261.jpg", "image_iiif_id": "https://iip.smk.dk/iiif/jp2/qz20sx771_kks5261.tif.jp2", "image_iiif_info": "https://iip.smk.dk/iiif/jp2/qz20sx771_kks5261.tif.jp2/info.json" } ``` To fetch a specific single object later by its `object_number`, use: ``` GET https://api.smk.dk/api/v1/art?object_number=KKS5261 ``` (the `object_url` field on every item is pre-built this way). ## Example URLs (both verified live, 2026-07-24) **exampleApiUrl** (search query — returns JSON, confirmed 200 with 39,479 total matches): ``` https://api.smk.dk/api/v1/art/search?keys=*&filters=%5Bpublic_domain:true%5D&filters=%5Bhas_image:true%5D&offset=0&rows=2 ``` **exampleImageUrl** (full-resolution downloadable JPEG, confirmed HTTP 200, `Content-Type: image/jpeg`, `Content-Length: 24043853` bytes, 4992x6287 px): ``` https://api.smk.dk/api/v1/download/W3siaW1nX3VybCI6Imh0dHBzOi8vaWlwLnNtay5kay9paWlmL2pwMi9xejIwc3g3NzFfa2tzNTI2MS50aWYuanAyL2Z1bGwvZnVsbC8wL25hdGl2ZS5qcGciLCJwdWJsaWNfZG9tYWluIjp0cnVlLCJvYmplY3RfbnVtYmVyIjoiS0tTNTI2MSIsIm51bSI6IiJ9XQ==/KKS5261.jpg ``` Alternative (IIIF-native, resize on the fly, also confirmed 200): ``` https://iip.smk.dk/iiif/jp2/qz20sx771_kks5261.tif.jp2/full/1024,/0/default.jpg ``` ## Verification log - `curl -sI` on `image_native` URL → `HTTP/1.1 200 OK`, `Content-Type: image/jpeg`, `Content-Length: 24043853`, `Access-Control-Allow-Origin: *`. - `curl -s` on `art/search` example URL → HTTP 200, valid JSON, `found: 39479`, `items[0].public_domain == true`, `items[0].has_image` implied by presence of `image_native`. - `curl -sI` on IIIF `info.json` → HTTP 200, valid IIIF Image API 2.0 manifest with `sizes` array. - Burst of 5 sequential unauthenticated requests → all 200, no throttling/key errors observed. ## Gotchas 1. **`filters` must be repeated, not concatenated** — see table above. This is the single biggest trap; the naive single-string form silently returns the wrong (larger, un-intersected) result set instead of erroring. 2. **`keys` is a required param** even for "give me everything" — use `keys=*`. 3. `image_hires` field exists in the schema but was `None` on the sampled item; `image_native` is the reliable full-res download link, not `image_hires`. 4. `image_native` URLs are single-use-looking base64-ish opaque tokens embedding the source IIIF path + object metadata — they are stable (not time-limited signed URLs) but don't try to hand-construct them; always take them verbatim from the API response. 5. IIIF server (`iip.smk.dk`) supports standard region/size/rotation/quality params if you want sizes other than native — `sizes` in `info.json` lists SMK's precomputed steps (156px up to 2496px wide) but arbitrary `w,` / `w,h` sizes also work via `full/1024,/0/default.jpg` syntax. 6. No published rate limit was hit in testing; be a good citizen (the API is free, maintained by a small team — contact smkapi@smk.dk for anything at bulk-harvest scale). 7. `rights` is per-item — not every item with `public_domain:true` necessarily has an identical `rights` URL, but in the sample it was the CC Public Domain Mark 1.0 link. 8. `object_number` (e.g. `KKS5261`) is the human-facing ID; `id` (e.g. `1170000001_object`) is internal. Use `object_number` for the `/art?object_number=` single-item lookup. -
_synthesis.md 9.2 KB
# Public-Domain Art Sourcing Cheat-Sheet Consolidated from 8 verified museum-API recipes (2026-07-24). For an agent sourcing hi-res public-domain artwork for blog heroes, decks, social cards, essay figures. ## Quick-pick table (ranked: fastest/most reliable first) | Rank | Museum | Keyless? | License | How to get a hi-res PD image (one line) | Best for | |---|---|---|---|---|---| | 1 | **Met (Metropolitan Museum)** | Yes, no key | CC0 (per-object flag) | `GET /search?q=X&hasImages=true&isPublicDomain=true` → `/objects/{id}` → read `primaryImage` (static JPEG, no IIIF hop) | Broadest single source — Western painting, Asian art, Egyptian antiquities, American art | | 2 | **Cleveland Museum of Art** | Yes, no key | CC0 | `GET /api/artworks/?cc0=1&has_image=1` → read `images.print.url` (3400px JPEG) or `images.full.url` (archival TIFF) — no hop | Broad European/American painting + Asian art; archival-res TIFFs available | | 3 | **SMK (National Gallery of Denmark)** | Yes, no key | Public Domain Mark (functionally CC0) | `GET /art/search?keys=*&filters=[public_domain:true]&filters=[has_image:true]` (repeat `filters=`, don't concatenate!) → read `image_native` directly | Danish/Nordic + European painting, decorative arts | | 4 | **Rijksmuseum** | Yes, no key | CC0/PD (mostly), check per-object | `GET /search/collection?type=painting&imageAvailable=true` → walk object → VisualItem → DigitalObject → `access_point[0].id` is the ready IIIF URL (3 sequential GETs) | Dutch/Flemish masters, prints, decorative arts | | 5 | **NGA (National Gallery of Art, DC)** | Yes, no key | CC0 (dataset); per-image `openaccess` flag gates rights | Download `published_images.csv` from GitHub, filter `openaccess=1`, then `GET https://api.nga.gov/iiif/{uuid}/full/full/0/default.jpg` | American + European painting/sculpture; best for bulk/offline querying (no live search API) | | 6 | **Art Institute of Chicago** | Yes, no key | CC0 | `GET /artworks/search?query[term][is_public_domain]=true&fields=...,image_id` → build `https://www.artic.edu/iiif/2/{image_id}/full/full/0/default.jpg` — **but** the image host (`www.artic.edu`) 403s scripted/curl fetches (Cloudflare bot challenge); JSON metadata API is fine, image download needs a real/headless browser | Impressionism, European + American painting, Buddhist/Asian art | | 7 | **Getty** | Yes, no key | CC0 (dataset); images "mostly" CC0 but per-image rights block inconsistently populated — verify before trusting | No keyword search exists — use SPARQL (`?obj crm:P138i_has_representation ?img`) which returns a ready-to-use IIIF image URL directly in the same query | European painting, photography, antiquities, decorative arts | | 8 | **Smithsonian** | **No** — needs `api_key` (free instant signup, or shared `DEMO_KEY` at 30 req/hr) | CC0 | `GET /search?q=X AND online_media_type:Images&api_key=KEY` → `media[].content` (IDS deliveryService URL); check `media[].usage.access=="CC0"` per-image | Extremely broad: American art (SAAM), portraiture (NPG), Asian art (Freer\|Sackler), design (Cooper Hewitt), natural history/specimens | ## The 2-3 best keyless APIs to reach for FIRST 1. **The Met Collection API** — single JSON call, no IIIF/linked-data hop, `primaryImage` field is a directly-downloadable full-res JPEG, CC0 is a simple boolean flag, 80 req/sec, huge and diverse collection. Best all-around default. 2. **Cleveland Museum of Art Open Access API** — same one-hop simplicity as the Met, plus it's the only one of the 8 offering an archival-resolution TIFF tier (`images.full`) alongside a 3400px print JPEG, no rate limit trouble observed. 3. **SMK (Denmark) Art API** — also one-hop (`image_native` ready in the search result, no chain), broad European painting/decorative holdings. Caveat: the `filters=` param MUST be repeated (one bracket each), not concatenated into one string, or the AND silently fails. Honorable mention: if the ask is specifically Dutch/Flemish masters or decorative arts, Rijksmuseum is worth the extra 2 hops. If it's European antiquities/photography and a keyword isn't essential, Getty's SPARQL trick returns an image URL in one query. ## Attribution / licensing notes - **True CC0 (no legal attribution required, safe to treat as public domain outright):** Met, Cleveland, Art Institute of Chicago, NGA (dataset-level; images gated by per-image `openaccess` flag), Smithsonian. - **"Public Domain Mark" / open-access-but-not-technically-CC0 (functionally free to use, museum requests but does not require a credit line):** Rijksmuseum (mixed — some items are CC BY and DO require attribution, always check the object's own rights block), SMK (Public Domain Mark 1.0 — functionally equivalent to CC0 for reuse, credit "SMK" requested). - **Mixed/needs per-image verification — do not blanket-trust the collection-level CC0 claim:** Getty (dataset is CC0 but per-image `VisualItem.subject_to` rights block is inconsistently populated on older records — cross-check the object's public page, which flags true Open Content items with a "Download" button, before using an image commercially), NGA (per-image `openaccess=0` rows are NOT open — resolution-capped and rights-restricted), Smithsonian (record-level `metadata_usage.access` and per-media `media.usage.access` can theoretically diverge — check the media-level flag, not just the record flag). - **Good practice everywhere even when not legally required:** a short credit line ("Digital image courtesy of [Museum]") costs nothing and avoids ambiguity, especially for mixed-license collections (Rijksmuseum, Getty) where a CC BY item could slip through a broad query. - **Practical rule for the agent:** when a museum's API exposes a rights/license field, filter AND spot-check it per image before shipping — never assume "the collection is open access" implies "this specific image is." The 3 collections where this bites hardest are Getty, Rijksmuseum, and NGA/Smithsonian's per-image flags. ## Gaps — major keyless/near-keyless PD art sources the 8 museums miss None of the 8 notes covered these; an agent building a general-purpose "find me a public-domain artwork" tool should know about them: - **Wikimedia Commons API** — `https://commons.wikimedia.org/w/api.php?action=query&generator=search&gsrsearch=<term>&gsrnamespace=6&prop=imageinfo&iiprop=url|extmetadata` (fully keyless). Aggregates PD/CC-licensed images from dozens of museums already normalized into one schema; `extmetadata.LicenseShortName` gives the license per file. Best single fallback when a specific museum's own API comes up empty — search `Category:Paintings_by_...` or a direct `File:` page via `Special:FilePath/{filename}` for the raw image bytes. - **Europeana API** — `https://api.europeana.eu/record/v2/search.json?wskey=<KEY>&query=<term>` — aggregates 50M+ items from thousands of European cultural institutions in one schema; **requires a free API key** (instant self-serve signup, not fully keyless like the others). - **Harvard Art Museums API** — `https://api.harvardartmuseums.org/object?apikey=<KEY>&q=<term>` — **requires a free key** (instant signup). Strong for teaching-collection-style Western art and object photography. - **Yale (LUX / Yale Center for British Art)** — `https://lux.collections.yale.edu/api/search/...` (Linked Art model, similar shape to Getty's) — keyless, aggregates Yale University Art Gallery + Yale Center for British Art + Beinecke; strong British art and rare books. - **Paris Musées (Paris city museums collections)** — open-data portal at `https://www.parismuseescollections.paris.fr/en/collections` with a CC0 bulk dataset also mirrored on Paris's open-data platform (`opendata.paris.fr`) — keyless, French painting/decorative arts, explicit CC0. - **NYPL Digital Collections API** — `https://api.repository.library.nyc/...` — **requires a free key** (instant signup via `api.repository.library.nyc/register`). Huge trove of digitized public-domain prints, maps, photographs, illustrations (not "museum paintings" but excellent for editorial/essay figures). - **Wellcome Collection API** — `https://api.wellcomecollection.org/catalogue/v2/works?query=<term>` — keyless, no key needed. CC0/PD medical, historical, and scientific imagery — a good niche source for essay figures on health/science topics that the 8 art museums won't have. - **Internet Archive** — `https://archive.org/advancedsearch.php?q=<term>&fl[]=identifier&output=json` — keyless. Not a museum but a deep well of PD book illustrations, historical photographs, and scanned ephemera; useful when the need is "old illustration/engraving" rather than "fine-art painting." - **DPLA (Digital Public Library of America)** — `https://api.dp.la/v2/items?q=<term>&api_key=<KEY>` — **requires a free key**. Aggregates US libraries/archives/museums (including several of the 8 above) into one federated search — useful as a cross-collection fallback once a key is obtained. **Practical recommendation for the agent:** try Met → Cleveland → SMK first (keyless, one-hop). If nothing fits the brief, fall back to Wikimedia Commons (keyless aggregator across everything). If the visual need is more "historical illustration/engraving" than "museum painting," go straight to Internet Archive or Wikimedia Commons instead of the fine-art APIs.
-
-
SKILL.md 10.2 KB
--- name: museum-art description: Source authentic, high-res PUBLIC-DOMAIN artwork from museum open-access APIs (Met, Cleveland, SMK, Rijksmuseum, NGA, Art Institute of Chicago, Getty, Smithsonian) instead of AI-generated or generic-stock imagery. The default move whenever a visual needs an aesthetic, credible image (blog heroes, decks, social cards, essay/spec figures). Verified keyless recipes + licensing rules inside. --- # museum-art: Public-Domain Artwork for Visuals > Standing rule (adopted 2026-07-24, from Eric Li's post on museum open-access): **whenever a visual needs a real image with aesthetic weight, source public-domain museum artwork first** - over AI-generated imagery and over generic stock. Curated, historically significant art reads as credible and sophisticated; AI-gen reads as slop. This stacks with, and reinforces, the `no-ai-slop` skill and your house image style. It does NOT replace the generative `editorial-illustrations` skill (that owns claim-driven diagrams/figures) or your house chart style - use museum art for photographic/hero/decorative/mood imagery, generative figures for data and concept diagrams. ## When to reach for this - Blog post hero images, section breaks, mood imagery (the blog-publish image step). - Deck/slide backgrounds and section dividers, social cards, essay figures, spec cover art. - Any time the instinct is "generate an image" for something decorative or evocative. Stop and pull a real painting instead. - NOT for: product screenshots, UI mockups, data charts, logos, or claim-driven explanatory diagrams (those are editorial-illustrations / real captures). ## Decision: which source 1. **Met -> Cleveland -> SMK first.** All keyless, one JSON hop, CC0/PD, broad collections. Fastest path to a hi-res image. 2. **Need Dutch/Flemish masters or decorative arts?** Rijksmuseum (keyless, 3 hops). 3. **Need European antiquities/photography and keyword isn't essential?** Getty (keyless, SPARQL). 4. **Nothing fits, or you want a cross-museum search?** Wikimedia Commons API (keyless aggregator, normalized license metadata) is the best general fallback. 5. **"Historical illustration/engraving/old photo" rather than fine-art painting?** Go straight to Internet Archive or Wikimedia Commons. 6. **Science/health/medical essay figure?** Wellcome Collection (keyless, CC0/PD). ## How to fetch in THIS environment - Use **WebFetch** to hit the JSON search endpoint, then WebFetch/download the returned image URL. These APIs are server-side reachable; no browser needed. - **Exception - Art Institute of Chicago images:** the JSON API (`api.artic.edu`) is fine, but the image host `www.artic.edu/iiif/...` 403s scripted/curl fetches via a Cloudflare bot challenge. Use the JSON metadata from AIC, but download the actual image through a real/headless browser (browser-harness) or prefer a different museum for the image bytes. - Always **filter for public domain in the query AND spot-check the per-image license flag** before shipping (see Licensing). ## Freshness over caching (mandatory) **Fetch fresh per need. Do NOT build a reusable local pool of downloaded images to draw from.** A small cached set gets reused everywhere and becomes the new "same stock photo on every post" - sameness is a form of slop, and the variety of a huge open collection is the entire point. Fetching is keyless and sub-second, so there is no cost reason to cache pixels. - **Cache recipes/metadata, not images** - that is what this skill's `references/` already are. - **Commit an image only into the specific artifact that uses it** (a post's `assets/`, a deck's media) once chosen - for provenance and offline builds. That is an artifact asset, never a shared library other artifacts pull from. - Each new visual = a fresh query. Vary the search terms and the source museum so consecutive posts do not converge on the same few crowd-pleasers. ## Verified keyless recipes (2026-07-24, all live-tested) ### 1. The Met - best all-around default - Base: `https://collectionapi.metmuseum.org/public/collection/v1/` - no key, 80 req/s, CC0 where `isPublicDomain:true`. - Search: `GET /search?q=<term>&hasImages=true&isPublicDomain=true` -> `{total, objectIDs[]}` - Object: `GET /objects/{id}` -> read `primaryImage` (full-res JPEG, static, no IIIF hop) or `primaryImageSmall` (web-large). - Example: `https://collectionapi.metmuseum.org/public/collection/v1/search?q=sunflowers&hasImages=true&isPublicDomain=true` ### 2. Cleveland Museum of Art - only one with archival TIFF - Base: `https://openaccess-api.clevelandart.org/api/artworks/` - no key, CC0. - Search: `GET /api/artworks/?cc0=1&has_image=1&limit=10` (add `&q=monet`, `&skip=10`). - Image fields on each result: `images.web.url` (900px), `images.print.url` (3400px JPEG), `images.full.url` (archival TIFF). Directly downloadable from `openaccess-cdn.clevelandart.org`. - Confirm `share_license_status == "CC0"` per result. ### 3. SMK (Denmark) - one-hop, broad European/Nordic - Base: `https://api.smk.dk/api/v1` - no key. License: Public Domain Mark 1.0 (functionally CC0). - Search: `GET /art/search?keys=*&filters=[public_domain:true]&filters=[has_image:true]` - Read `image_native` directly from each result (no chain). - **Gotcha (verified):** pass `filters` as a **repeated** query param, one `[field:value]` bracket each. Concatenating `filters=[public_domain:true][has_image:true]` returns 200 but silently ignores the second condition. ### 4. Rijksmuseum - Dutch/Flemish masters (keyless, 3 hops) - Base: `https://data.rijksmuseum.nl/search/collection` (Linked Art, no key). - `GET /search/collection?type=painting&imageAvailable=true` -> walk object -> VisualItem -> DigitalObject -> `access_point[0].id` is a ready IIIF URL. - Mixed license: mostly CC0/PD but some CC BY 4.0 - **check the per-object rights block** (a CC BY item requires attribution). ### 5. NGA (Washington) - bulk/offline, no live search - CSV: `https://raw.githubusercontent.com/NationalGalleryOfArt/opendata/main/data/published_images.csv` - filter `openaccess=1`. - Image (IIIF): `https://api.nga.gov/iiif/{uuid}/full/full/0/default.jpg` - **Per-image `openaccess=0` rows are NOT open** (resolution-capped, rights-restricted). Only `openaccess=1` is free. ### 6. Art Institute of Chicago - Impressionism/European (image host caveat) - JSON: `GET https://api.artic.edu/api/v1/artworks/search?query[term][is_public_domain]=true&fields=id,title,image_id` - no key, CC0. - Image: build `https://www.artic.edu/iiif/2/{image_id}/full/843,/0/default.jpg` - **but Cloudflare 403s scripted fetches**; download via browser or use another museum for bytes. ### 7. Getty - European painting/photography/antiquities (SPARQL) - Base: `https://data.getty.edu/museum/collection/` - no key. No keyword search; use SPARQL with `?obj crm:P138i_has_representation ?img` which returns a ready IIIF image URL in the same query. - Dataset is CC0; **per-image rights inconsistently populated** - cross-check the object's public getty.edu page (a "Download" button flags true Open Content) before commercial use. ## Key-gated + aggregator sources (know these too) - **Smithsonian** (needs free `api.data.gov` key; `DEMO_KEY` works at 30 req/hr): `GET https://api.si.edu/openaccess/api/v1.0/search?q=<term> AND online_media_type:Images&api_key=KEY`; CC0; check `media[].usage.access=="CC0"` per image. Extremely broad (SAAM, NPG portraits, Freer/Sackler Asian, Cooper Hewitt design). - **Wikimedia Commons** (keyless, best aggregator fallback): `https://commons.wikimedia.org/w/api.php?action=query&generator=search&gsrsearch=<term>&gsrnamespace=6&prop=imageinfo&iiprop=url|extmetadata`; license per file in `extmetadata.LicenseShortName`; raw bytes via `Special:FilePath/{filename}`. - **Wellcome Collection** (keyless): `https://api.wellcomecollection.org/catalogue/v2/works?query=<term>`; CC0/PD medical/scientific/historical imagery. - **Internet Archive** (keyless): `https://archive.org/advancedsearch.php?q=<term>&fl[]=identifier&output=json`; PD book illustrations, engravings, historical photos, ephemera. - **Europeana** (free key): `https://api.europeana.eu/record/v2/search.json?wskey=<KEY>&query=<term>`; 50M+ items across European institutions. - **Harvard Art Museums** (free key): `https://api.harvardartmuseums.org/object?apikey=<KEY>&q=<term>`. - **Yale LUX** (keyless, Linked Art like Getty): `https://lux.collections.yale.edu/api/search/...`; strong British art + rare books. - **Paris Musees** (keyless, explicit CC0): `parismuseescollections.paris.fr` / `opendata.paris.fr`; French painting/decorative arts. - **NYPL Digital Collections** (free key): `https://api.repository.library.nyc/...`; PD prints, maps, photos, illustrations. - **DPLA** (free key): `https://api.dp.la/v2/items?q=<term>&api_key=<KEY>`; federated US collections. ## Licensing rules (apply before shipping any image) - **True CC0 (no attribution required):** Met, Cleveland, Art Institute of Chicago, Smithsonian, NGA (dataset-level; image gated by per-image `openaccess=1`). - **Public Domain Mark / open-but-not-CC0 (free to use, credit requested not required):** SMK, and the PD subset of Rijksmuseum. - **Verify per-image, do NOT blanket-trust collection-level "open access":** Getty, Rijksmuseum (CC BY items require attribution), NGA and Smithsonian per-image flags. - **Rule:** when the API exposes a rights/license field, filter on it in the query AND spot-check it on the chosen image. Never assume "the collection is open access" implies "this specific image is." - **Good practice everywhere:** a one-line credit ("Digital image courtesy of [Museum]") costs nothing and covers mixed-license collections. When masking/compositing per house style, keep the credit. ## Relationship to other skills - Stacks with **no-ai-slop** and the house image style: real museum art is the anti-slop default for evocative imagery. - Complements **editorial-illustrations** (generative claim-driven diagrams) and **dataviz** (charts) - those own explanatory graphics; this owns photographic/artwork/mood imagery. - Feeds **blog-publish**, **social-media-kit**, **weekly-ai-slide**, deck and essay work at their image-sourcing step. ## References (full per-museum recipes, verified) `references/_synthesis.md` (cheat-sheet) + one file per museum (`met.md`, `cleveland.md`, `smk.md`, `rijksmuseum.md`, `nga.md`, `artic.md`, `getty.md`, `smithsonian.md`) with live-tested example URLs, field maps, and gotchas.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.