Claude Skill

api-connector-builder

Use when writing a client for someone else's REST or GraphQL API: auth flow choice and token refresh, pagination to exhaustion, retry-with-jitter on transient failures only, rate-limit-aware throttling. NOT inbound callbacks (that is `webhooks`), NOT chaining services (that is `a

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download ericrisco-rsc-harness-skills_api-connector-builder-953fef5.zip · 16 KB
Part of ericrisco/rsc-harness — 46 skills

Install

skills CLI npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/api-connector-builder
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
Git git clone https://github.com/ericrisco/rsc-harness.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

API connector builder — auth, pagination, retries, rate limits

You are writing a client for someone else's HTTP API. You do not own the contract; you obey it. The deliverable is one typed connector module per vendor that authenticates, walks the whole result set, retries only transient failures with backoff, and stays under the rate limit without getting the key banned. One connector = one vendor: mixed clients tangle two auth schemes and two rate-limit budgets into something no one can reason about.

Four pillars, every time: auth, pagination, retries, rate limits. If your connector skips any one of them it works in the demo and breaks in production — on page 2, on token expiry, on the first 429, or on a flaky network.

Step 0 — read the contract

Read the vendor docs before writing a line — invented endpoints and guessed field names 404 in prod. Extract these first; each one changes what you write, so missing one means a rewrite.

Find in their docs Why it changes your code
Auth scheme + token TTL Picks the flow below; TTL decides if you need refresh
Base URL and version Wrong version = silent 404s or deprecated field shapes
Rate-limit header names You cannot throttle to a budget you cannot read
Pagination style Cursor vs offset vs Link vs Relay = different loop
Error envelope shape Where the real error code lives (body, not always status)
Idempotency support Decides whether POST is safe to retry (key header?)

If a fact is not in the docs, probe one real call and read the response headers and body — do not assume.

Auth — pick the flow, then store the secret right

OAuth 2.1 is the current baseline. Three deltas you must honor: PKCE is mandatory for every authorization-code client, the Implicit and Resource-Owner-Password grants are removed, and bearer tokens may not travel in query strings (header only) — query strings end up in logs and referrers. (oauth.net/2.1, accessed 2026-06-02.)

Flow Pick it when Secret lives Refresh strategy
API key in header Simple server-to-server, vendor issues key env / secret store none; rotate manually
Bearer static token Personal access token, long-lived env / secret store none; treat as a key
OAuth2 Client Credentials Machine-to-machine, no end user client_id + secret re-mint on 401; cache until expiry
OAuth2 Auth-Code + PKCE Acting on behalf of a user refresh token rotate on every refresh (see below)
Device flow CLI / TV / input-constrained device refresh token poll then rotate as auth-code

Refresh-token rule (public clients): refresh tokens must be sender- constrained or one-time-use — rotated on every refresh, with the old one invalidated. Keep access tokens short-lived. Never log a raw token; log a partial or hash if you must correlate. (OAuth 2.0 Security BCP / RFC 9700, accessed 2026-06-02.) Full per-flow walkthroughs, token storage, and a DPoP note are in references/auth-flows.md.

Retries + backoff

Retry only transient failures, and only when the operation is safe to repeat.

Status / error Retry? Note
429 Too Many Requests yes honor Retry-After (see rate limits)
502 / 503 / 504 yes server-side transient
Connection reset / timeout / DNS yes network transient
400 / 401 / 403 / 404 / 409 / 422 no replay returns the same error — fix the call
2xx n/a success

Idempotency gate. GET/PUT/DELETE are idempotent by semantics and safe to retry. POST is not — only retry it if you send a stable idempotency key so the server dedupes the duplicate. (AWS Builders' Library, accessed 2026-06-02.)

Backoff with jitter. Use delay = min(cap, base * 2^attempt) + random_jitter. The jitter is not optional: without it, every client that failed at the same instant retries at the same instant — a synchronized thundering herd that DDoSes the recovering server. Full or decorrelated jitter is preferred. Always set stop conditions: max attempts 3–5, a total deadline, and a per-attempt timeout. (AWS Builders' Library, accessed 2026-06-02.)

# Bad: retries everything, flat sleep, no jitter, no cap, no deadline.
for _ in range(10):
    r = httpx.get(url)
    if r.status_code == 200:
        return r.json()
    time.sleep(1)  # 4xx will never recover; herds synchronize on flat 1s
# Good: transient-only, exponential backoff WITH jitter, bounded.
import random, time, httpx

RETRYABLE = {429, 502, 503, 504}

def get(url, *, attempts=5, base=0.5, cap=20.0, deadline=60.0):
    started = time.monotonic()
    for attempt in range(attempts):
        try:
            r = httpx.get(url, timeout=10.0)  # per-attempt timeout
        except (httpx.ConnectError, httpx.ReadTimeout):
            pass  # network transient -> fall through to backoff
        else:
            if r.status_code == 200:
                return r.json()
            if r.status_code not in RETRYABLE:
                r.raise_for_status()  # 4xx: do not retry, surface it
        if time.monotonic() - started > deadline:
            raise TimeoutError("retry deadline exceeded")
        delay = min(cap, base * 2 ** attempt) + random.uniform(0, base)
        time.sleep(delay)
    raise RuntimeError("max attempts exhausted")

In Python prefer tenacity 9.1.4 (@retry with wait_exponential_jitter, stop_after_attempt, retry_if_exception_type) over a hand loop; in Node/TS use undici (the engine behind global fetch() since Node 18) with its retry interceptor. (PyPI/tenacity, nodejs/undici, accessed 2026-06-02.)

Rate limits

Do not guess the wait — the server tells you. Precedence:

  1. Retry-After present (on a 429 or 503) → wait exactly that. It is either a number of seconds or an HTTP-date; handle both.
  2. No Retry-After → compute the wait from X-RateLimit-Reset (a Unix epoch or a delta-seconds, per vendor docs).
  3. Proactively throttle on X-RateLimit-Remaining — slow down before you hit zero rather than absorbing a wall of 429s. (iotools.cloud rate-limiting guidance, accessed 2026-06-02.)
# Bad: ignore the headers, hammer, eat 429s, get the key throttled or banned.
while more:
    resp = client.get(next_url)   # no remaining check, no Retry-After
    process(resp)
# Good: header-aware. Honor Retry-After, else reset; brake near the limit.
def wait_for_rate_limit(resp):
    ra = resp.headers.get("Retry-After")
    if ra is not None:
        return float(ra) if ra.isdigit() else _seconds_until_httpdate(ra)
    remaining = int(resp.headers.get("X-RateLimit-Remaining", "1"))
    if remaining <= 1:
        reset = float(resp.headers.get("X-RateLimit-Reset", "0"))
        return max(0.0, reset - time.time())  # or reset directly if delta-seconds
    return 0.0

For sustained pulls, gate every request through a token-bucket sized to the documented budget (e.g. 60 req/min → refill 1 token/sec, capacity 60) so bursts smooth out instead of slamming the wall.

Pagination — loop to exhaustion

There is no single best strategy; the vendor chose one and you follow it. The universal rule: loop until the API says there is no next page — never a fixed page count. A hardcoded for page in range(10) silently drops everything after page 10.

Style Signal of "next" Trade-off
Offset / page number ?offset= / ?page= until empty page simple; drifts on inserts, slow at depth
Cursor / keyset opaque next_cursor in body stable under writes, fixed cost — prefer it
Link header (REST) Link: <...>; rel="next" parse the header, follow until no next
Relay (GraphQL) pageInfo.hasNextPage + endCursor pass endCursor as after next query

(graphql.org/learn/pagination + pagination pattern guides, accessed 2026-06-02.)

Expose results as a generator / async-iterator so callers stream instead of buffering everything:

def iter_records(client):
    cursor = None
    while True:
        page = client.get("/items", params={"cursor": cursor, "limit": 100})
        body = page.json()
        yield from body["data"]
        cursor = body.get("next_cursor")
        if not cursor:           # exhaustion signal, not a counter
            return

Full code for every style — offset, keyset, Link-header parsing, GraphQL Relay connection walking, in Python and TS, plus dedup-on-overlap notes — is in references/pagination.md.

Putting it together

A minimal connector wires all four pillars plus env config and structured logging.

# connector.py — Python: httpx + tenacity
import os, logging, httpx
from tenacity import retry, stop_after_attempt, wait_exponential_jitter, retry_if_exception_type

log = logging.getLogger("connector")
TOKEN = os.environ["VENDOR_API_TOKEN"]          # from env, never hardcoded

class Transient(Exception): ...

@retry(stop=stop_after_attempt(5),
       wait=wait_exponential_jitter(initial=0.5, max=20),
       retry=retry_if_exception_type(Transient))
def _request(client, method, path, **kw):
    r = client.request(method, path, timeout=10.0, **kw)   # per-request timeout
    log.info("req id=%s %s status=%s", r.headers.get("x-request-id"), path, r.status_code)
    if r.status_code in (429, 502, 503, 504):
        raise Transient(r.status_code)
    r.raise_for_status()
    return r

def client():
    return httpx.Client(base_url="https://api.vendor.com/v2",
                        headers={"Authorization": f"Bearer {TOKEN}"})  # header, not query
// connector.ts — Node/TS: global fetch (undici) + bounded retry
const TOKEN = process.env.VENDOR_API_TOKEN!;          // from env, never hardcoded
const RETRYABLE = new Set([429, 502, 503, 504]);

export async function request(path: string, init: RequestInit = {}, attempt = 0): Promise<Response> {
  const res = await fetch(`https://api.vendor.com/v2${path}`, {
    ...init,
    headers: { Authorization: `Bearer ${TOKEN}`, ...init.headers },
    signal: AbortSignal.timeout(10_000),              // per-request timeout
  });
  console.info(JSON.stringify({ id: res.headers.get("x-request-id"), path, status: res.status, attempt }));
  if (RETRYABLE.has(res.status) && attempt < 4) {
    const ra = Number(res.headers.get("retry-after"));
    const wait = Number.isFinite(ra) && ra > 0 ? ra * 1000 : Math.min(20_000, 500 * 2 ** attempt) + Math.random() * 500;
    await new Promise((r) => setTimeout(r, wait));
    return request(path, init, attempt + 1);
  }
  return res;                                          // caller checks res.ok / paginates
}

Anti-patterns

Anti-pattern Consequence Fix
Retry 4xx (401/403/404/422) Burns attempts; same error every time Retry only 429 + 5xx + network errors
for page in range(N) fixed loop Silently drops records past page N Loop on cursor / Link / hasNextPage
Flat sleep(1) between retries Synchronized herd hammers recovering API Exponential backoff with jitter + cap
Log the token / Authorization header Leaked credential in shipped logs Log request id + status + attempt only
Hardcode the API key in source Leaks via git; cannot rotate cleanly Read from env / secret store
No request timeout One hung socket stalls the whole run Set a per-request timeout always
Ignore Retry-After Keep 429-ing; key gets throttled/banned Honor Retry-After, else X-RateLimit-Reset
Re-POST on retry with no idempotency Duplicate charges / records Send an idempotency key, or do not retry POST
Token in query string Token ends up in logs / referrers Bearer in the Authorization header
Implicit / password OAuth grant Removed in OAuth 2.1; insecure Auth-Code + PKCE, or Client Credentials

Verify

Run scripts/verify.sh over the connector you write: it greps for hardcoded secrets, asserts a retry mechanism and a pagination loop and a request timeout exist, and flags localStorage token storage or plaintext token logging. It is a structure linter (read-only), not a behavior test — exit 0 on a clean target.

Files (rsc-harness)
  • evals
    • cases.yaml 4 KB
      skill: api-connector-builder
      
      should_trigger:
        - prompt: "Write a connector for the Linear GraphQL API that pulls every issue, not just the first page."
          why: "Core task: a typed client against a third-party GraphQL API with Relay cursor connections, walked to exhaustion. The skill's flagship deliverable."
        - prompt: "This vendor API keeps returning 429. Add backoff and respect their rate limit so we stop getting throttled."
          why: "Rate-limit + retry symptom. Honor Retry-After / X-RateLimit-* with jittered backoff is squarely this skill."
        - prompt: "We only ever get the first 100 results back from this API. Paginate through everything."
          why: "Pagination-to-exhaustion. The fixed-first-page bug is the canonical thing this skill fixes."
        - prompt: "Our access token expires halfway through a long run and the whole job dies. Add refresh."
          why: "Non-obvious phrasing — never says 'OAuth' or 'connector', but mid-run token expiry is the auth/refresh concern the skill owns (refresh rotation)."
        - prompt: "Conector para una API de terceros que nos banea la key cuando vamos demasiado rápido."
          why: "Spanish trigger combining rate-limit + auth (banned key from hammering). Vendor-agnostic outbound client = this skill."
        - prompt: "Wrap this vendor REST API as a typed SDK with retries and per-request timeouts."
          why: "Typed-client framing. 'Wrap a REST API as an SDK' with resilience is the connector-module deliverable."
        - prompt: "Build me a client for the Shopify Admin API that reads the cursor-based pages and handles the leaky-bucket rate limit."
          why: "Vendor with no dedicated skill, cursor pagination + bucket rate limit — the generic connector case this skill covers."
      
      should_not_trigger:
        - prompt: "Verify the incoming webhook signature from Stripe on our /hooks endpoint and reject replays."
          route_to: webhooks
          why: "Inbound callback — receiving and verifying events the vendor POSTs to us, the inverse of calling their API."
        - prompt: "Design the versioning, resource model, and error envelope for OUR new public API."
          route_to: api-design
          why: "Authoring a contract others will consume, not consuming someone else's. Opposite side of the boundary."
        - prompt: "Chain Gmail to Sheets to Slack into one automated if-this-then-that workflow."
          route_to: automation-flows
          why: "Multi-service orchestration across many tools, not building one connector for one vendor."
        - prompt: "Scrape the product prices off this storefront — it has no public API at all."
          route_to: data-scraper
          why: "HTML/DOM scraping with no documented API; this skill requires a real API contract to obey."
        - prompt: "Pull the structured fields out of this LLM's free-text answer into JSON."
          route_to: structured-extraction
          why: "Extracting structure from a model response, not making HTTP calls to a third-party API."
      
      capability:
        - scenario: "Build a Python connector for a paginated REST API (cursor-based) that requires a Bearer token and rate-limits at 60 req/min with X-RateLimit-* headers. It must fetch all records."
          must_include:
            - "Reads the Bearer token from an environment variable / secret store — never hardcoded in source."
            - "Sets a per-request timeout on every call (no unbounded socket wait)."
            - "Retries only transient statuses (429, 502, 503, 504) and network errors; never retries other 4xx."
            - "Uses exponential backoff WITH jitter and explicit stop conditions (max attempts 3–5 and/or a total deadline)."
            - "Honors Retry-After when present, otherwise computes the wait from X-RateLimit-Reset; throttles proactively on X-RateLimit-Remaining."
            - "Paginates by following the cursor until exhausted (no fixed page-count loop); ideally exposes a generator/iterator."
            - "Logs request id / status / attempt but never the raw token or Authorization header."
            - "Sends the token in the Authorization header, not a query string."
            - "Returns or streams the FULL record set across all pages, not just the first page."
      
    • README.md 687 B
      # Evals — api-connector-builder
      
      These cases are a hand-run rubric, not an automated test harness; no scoring
      script ships. To run them: paste a `should_trigger` prompt into a fresh agent
      session and confirm `api-connector-builder` is the skill that fires (and not a
      sibling). Paste a `should_not_trigger` prompt and confirm the agent routes to the
      named sibling instead. For the `capability` scenario, have the agent build the
      connector, then grade the result against the `must_include` list — every bullet
      should be satisfied by the code it produces. A connector that hardcodes the token,
      retries all 4xx, or stops after one page fails the rubric regardless of how clean
      it looks.
      
  • references
    • auth-flows.md 5.3 KB
      # Auth flows — per-flow walkthroughs
      
      OAuth 2.1 baseline (oauth.net/2.1, accessed 2026-06-02):
      
      - **PKCE is mandatory** for every authorization-code client, public and
        confidential alike.
      - **Implicit grant and Resource-Owner-Password grant are removed.** If a vendor's
        docs still show them, do not use them — pick Auth-Code + PKCE or Client
        Credentials instead.
      - **Bearer tokens must not travel in query strings.** Header only. Query strings
        leak into access logs, browser history, and `Referer` headers.
      
      ## API key in header
      
      Simplest scheme. The vendor issues a static key; you send it on every request.
      Store it in env, never source. There is no refresh — rotate manually when it
      leaks or on a schedule.
      
      ```python
      import os, httpx
      KEY = os.environ["VENDOR_API_KEY"]
      client = httpx.Client(base_url="https://api.vendor.com/v1",
                            headers={"X-API-Key": KEY})  # header name per vendor docs
      ```
      
      Some vendors want `Authorization: Bearer <key>` instead of a custom header — read
      the docs; do not guess the header name.
      
      ## Bearer static token (personal access token)
      
      Treat exactly like an API key: long-lived, env-stored, no refresh. The only
      difference is the standard `Authorization: Bearer <token>` header.
      
      ```typescript
      const TOKEN = process.env.VENDOR_PAT!;
      const headers = { Authorization: `Bearer ${TOKEN}` };
      ```
      
      ## OAuth2 Client Credentials (machine-to-machine)
      
      No end user is involved — your service authenticates as itself. You hold a
      `client_id` + `client_secret`, exchange them for a short-lived access token at the
      token endpoint, cache it until it expires, and re-mint on expiry or a 401.
      
      ```python
      import os, time, httpx
      
      _cache = {"token": None, "exp": 0.0}
      
      def access_token():
          if _cache["token"] and time.time() < _cache["exp"] - 30:  # 30s safety margin
              return _cache["token"]
          r = httpx.post("https://auth.vendor.com/oauth/token", data={
              "grant_type": "client_credentials",
              "client_id": os.environ["VENDOR_CLIENT_ID"],
              "client_secret": os.environ["VENDOR_CLIENT_SECRET"],
              "scope": "read:items",
          }, timeout=10.0)
          r.raise_for_status()
          body = r.json()
          _cache.update(token=body["access_token"], exp=time.time() + body["expires_in"])
          return _cache["token"]
      ```
      
      ## OAuth2 Authorization Code + PKCE (acting for a user)
      
      The user-facing flow. PKCE prevents an intercepted authorization code from being
      redeemed by an attacker.
      
      1. Generate a `code_verifier` (43–128 chars, random) and its
         `code_challenge = base64url(sha256(verifier))`.
      2. Redirect the user to the authorize endpoint with
         `response_type=code`, `code_challenge`, `code_challenge_method=S256`,
         `redirect_uri`, `scope`, and a random `state`.
      3. On callback, verify `state`, then POST `grant_type=authorization_code` with
         the `code` **and the original `code_verifier`** to the token endpoint.
      4. Store the resulting refresh token securely; use the access token for calls.
      
      ```python
      import base64, hashlib, os
      verifier = base64.urlsafe_b64encode(os.urandom(64)).rstrip(b"=").decode()
      challenge = base64.urlsafe_b64encode(
          hashlib.sha256(verifier.encode()).digest()).rstrip(b"=").decode()
      # -> send `challenge` (S256) on /authorize, keep `verifier` for the token exchange
      ```
      
      ### Refresh-token rotation (RFC 9700 / OAuth 2.0 Security BCP)
      
      Refresh tokens for public clients must be **sender-constrained or one-time-use**:
      every refresh returns a *new* refresh token and invalidates the old one. If the
      old token is ever replayed, the server detects the reuse and revokes the whole
      chain (a sign of theft). Persist the latest refresh token atomically — a crash
      between "got new token" and "saved it" locks you out.
      
      ```python
      def refresh(refresh_token: str) -> dict:
          r = httpx.post("https://auth.vendor.com/oauth/token", data={
              "grant_type": "refresh_token",
              "refresh_token": refresh_token,
              "client_id": os.environ["VENDOR_CLIENT_ID"],
          }, timeout=10.0)
          r.raise_for_status()
          body = r.json()
          save_refresh_token(body["refresh_token"])  # ROTATE: store the new one, drop old
          return body
      ```
      
      Keep access tokens short-lived (minutes). Never log a raw token — log a partial
      (`tok…3f9c`) or a hash if you must correlate across services.
      
      ## Device flow (input-constrained devices)
      
      For CLIs, TVs, and devices with no browser. Request a device + user code, show the
      user a URL and code to enter on another device, then **poll** the token endpoint
      until they approve (respect the `interval`; back off on `slow_down`). Once
      granted, you receive access + refresh tokens and rotate them exactly like the
      auth-code flow above.
      
      ## DPoP (sender-constraining, brief)
      
      DPoP (Demonstrating Proof-of-Possession) binds a token to a client-held key: each
      request carries a signed `DPoP` header proving you hold the private key, so a
      stolen bearer token alone is useless. If a vendor offers DPoP-bound access tokens,
      prefer them for high-value scopes — it is the practical way to satisfy the
      "sender-constrained" requirement without mTLS infrastructure.
      
      ## Storage rules (all flows)
      
      - Secrets in env vars or a secret manager — never in source, never in
        `localStorage` (XSS-readable) for anything long-lived.
      - One credential set per vendor per environment; never share a prod key into dev.
      - Rotate on any suspected leak; the flows above make rotation cheap by design.
      
    • pagination.md 4.1 KB
      # Pagination — full code per style
      
      The universal rule: **loop until the API signals no next page.** Never a fixed
      page count. Expose results as a generator / async-iterator so callers stream
      instead of buffering the whole set in memory.
      
      (graphql.org/learn/pagination + REST pagination pattern guides, accessed
      2026-06-02.)
      
      ## Offset / page number
      
      Simple, but it **drifts** when rows are inserted or deleted mid-walk (you skip or
      double-read) and gets **slow at large offsets** (the server scans and discards).
      Prefer cursor/keyset when the API offers it.
      
      ```python
      def iter_offset(client, path, page_size=100):
          offset = 0
          while True:
              body = client.get(path, params={"offset": offset, "limit": page_size}).json()
              rows = body["data"]
              if not rows:                  # empty page = exhausted
                  return
              yield from rows
              offset += len(rows)
      ```
      
      ```typescript
      export async function* iterOffset(get: (q: Record<string, number>) => Promise<any>, pageSize = 100) {
        let offset = 0;
        for (;;) {
          const body = await get({ offset, limit: pageSize });
          const rows = body.data as unknown[];
          if (rows.length === 0) return;     // exhausted
          yield* rows;
          offset += rows.length;
        }
      }
      ```
      
      ## Cursor / keyset
      
      The server returns an opaque cursor for the next page; you pass it back. **Stable
      under concurrent writes and fixed cost regardless of depth** — the strategy to
      prefer.
      
      ```python
      def iter_cursor(client, path, page_size=100):
          cursor = None
          while True:
              body = client.get(path, params={"cursor": cursor, "limit": page_size}).json()
              yield from body["data"]
              cursor = body.get("next_cursor")
              if not cursor:                # cursor exhausted
                  return
      ```
      
      ## Link header (REST)
      
      The next page lives in the `Link` response header with `rel="next"`. Follow it
      until the header has no `next`.
      
      ```python
      import re
      
      def iter_link(client, url):
          while url:
              resp = client.get(url)
              yield from resp.json()
              link = resp.headers.get("Link", "")
              m = re.search(r'<([^>]+)>;\s*rel="next"', link)
              url = m.group(1) if m else None   # no rel="next" = exhausted
      ```
      
      ```typescript
      function nextLink(header: string | null): string | null {
        if (!header) return null;
        const m = header.match(/<([^>]+)>;\s*rel="next"/);
        return m ? m[1] : null;
      }
      
      export async function* iterLink(start: string, get: (u: string) => Promise<Response>) {
        let url: string | null = start;
        while (url) {
          const res = await get(url);
          yield* (await res.json()) as unknown[];
          url = nextLink(res.headers.get("Link"));
        }
      }
      ```
      
      ## GraphQL — Relay Cursor Connections
      
      The Relay spec returns `edges` (each with a `node` and `cursor`) and a `pageInfo`
      with `endCursor` and `hasNextPage`. Pass `endCursor` as the `after` argument of
      the next query and stop when `hasNextPage` is false.
      
      ```python
      QUERY = """
      query($after: String) {
        issues(first: 100, after: $after) {
          edges { node { id title } }
          pageInfo { endCursor hasNextPage }
        }
      }
      """
      
      def iter_relay(client, url):
          after = None
          while True:
              body = client.post(url, json={"query": QUERY, "variables": {"after": after}}).json()
              conn = body["data"]["issues"]
              for edge in conn["edges"]:
                  yield edge["node"]
              info = conn["pageInfo"]
              if not info["hasNextPage"]:   # exhausted
                  return
              after = info["endCursor"]
      ```
      
      ## Dedup on overlap
      
      With offset pagination under concurrent inserts, the same row can appear on two
      pages. If exactness matters, track seen ids and skip duplicates:
      
      ```python
      def dedup(records, key="id"):
          seen = set()
          for r in records:
              k = r[key]
              if k in seen:
                  continue
              seen.add(k)
              yield r
      ```
      
      Cursor/keyset pagination does not have this problem for stable sort keys — another
      reason to prefer it.
      
      ## Combine with retries + rate limits
      
      Each `client.get(...)` above should go through the retry + rate-limit wrapper from
      `SKILL.md` (transient-only retry with jittered backoff, `Retry-After` honored).
      Pagination is the loop; the per-request call inside it is where resilience lives.
      
  • scripts
    • verify.sh 7.4 KB
      #!/usr/bin/env bash
      #
      # verify.sh — connector-module structure/banlist linter for `api-connector-builder`.
      #
      # WHAT IT DOES (read-only; never edits a file)
      #   Lints the connector source the skill produces against the four-pillar shape:
      #   secrets, retries, pagination, timeouts. It checks the artifact's STRUCTURE,
      #   not its runtime behavior — appropriate for a connector emitted into an
      #   arbitrary target repo. It never false-fails: a repo with no connector-like
      #   source is simply skipped.
      #
      #   A connector file = a .py / .ts / .js / .mjs / .tsx file that looks like an
      #   HTTP client (mentions a base URL, fetch/httpx/requests/undici/axios, or an
      #   Authorization/Bearer header). Other files are ignored.
      #
      #   Checks per connector file:
      #     1. FAIL  hardcoded secret  — an inline API key / bearer literal / long
      #              token-looking string assigned in source (not read from env).
      #     2. FAIL  no retry wired    — no tenacity import, no backoff loop, and no
      #              undici/axios retry interceptor anywhere in the file.
      #     3. FAIL  no pagination loop — uses pagination-style params (cursor/offset/
      #              page/after/Link) but has no loop (while/for/async generator) to
      #              walk them: a lone "get first page" that ignores the next cursor.
      #     4. FAIL  no request timeout — issues HTTP calls but sets no timeout /
      #              AbortSignal.timeout on them.
      #     5. FAIL  token in localStorage  — long-lived token stored in localStorage.
      #     6. WARN  token logging      — a log/console call that emits the token,
      #              Authorization header, or Bearer value.
      #
      # HOW TO RUN (inside YOUR project, not the skills repo)
      #   ./verify.sh                 # scan ./ for connector source
      #   ./verify.sh --path src      # scan a subdirectory
      #   ./verify.sh --strict        # treat any warning as a failure (exit 1)
      #
      # EXIT CODES
      #   0  clean, or warnings only without --strict (also: nothing to check)
      #   1  a structural failure, or --strict with a warning
      #   2  bad usage
      #
      # Runs on stock macOS bash 3.2 — no mapfile, no associative arrays.
      
      set -euo pipefail
      
      if [ -t 1 ]; then
        RED=$'\033[31m'; GREEN=$'\033[32m'; YELLOW=$'\033[33m'; NC=$'\033[0m'
      else
        RED=''; GREEN=''; YELLOW=''; NC=''
      fi
      
      ok_count=0; skip_count=0; warn_count=0; fail_count=0
      ok()   { printf '%s[ ok ]%s %s\n' "$GREEN"  "$NC" "$*"; ok_count=$((ok_count + 1)); }
      skip() { printf '%s[skip]%s %s\n' "$YELLOW" "$NC" "$*"; skip_count=$((skip_count + 1)); }
      warn() { printf '%s[warn]%s %s\n' "$YELLOW" "$NC" "$*"; warn_count=$((warn_count + 1)); }
      fail() { printf '%s[fail]%s %s\n' "$RED"    "$NC" "$*"; fail_count=$((fail_count + 1)); }
      
      usage() { sed -n '2,40p' "$0" | sed 's/^# \{0,1\}//'; }
      
      SCAN_PATH="."
      STRICT=0
      while [ $# -gt 0 ]; do
        case "$1" in
          --path)    SCAN_PATH="${2:?--path needs a value}"; shift 2 ;;
          --strict)  STRICT=1; shift ;;
          -h|--help) usage; exit 0 ;;
          *) printf '%sUnknown argument: %s%s\n\n' "$RED" "$1" "$NC"; usage; exit 2 ;;
        esac
      done
      
      if [ ! -e "$SCAN_PATH" ]; then
        printf '%sPath not found: %s%s\n' "$RED" "$SCAN_PATH" "$NC"; exit 2
      fi
      
      TMPDIR_V="$(mktemp -d 2>/dev/null || printf '/tmp/apiconn-verify.%s' "$$")"
      mkdir -p "$TMPDIR_V" 2>/dev/null || true
      cleanup() { rm -rf "$TMPDIR_V" 2>/dev/null || true; }
      trap cleanup EXIT
      
      ALL="$TMPDIR_V/all"
      FILES="$TMPDIR_V/files"
      
      # Candidate source files (skip node_modules / vendored / build dirs).
      find "$SCAN_PATH" -type f \
        \( -name '*.py' -o -name '*.ts' -o -name '*.tsx' -o -name '*.js' -o -name '*.mjs' \) \
        2>/dev/null \
        | grep -Ev '/(node_modules|\.git|dist|build|\.venv|venv|__pycache__)/' \
        > "$ALL" || true
      
      # Keep only files that look like an HTTP client (a connector).
      : > "$FILES"
      while IFS= read -r f; do
        [ -z "$f" ] && continue
        if grep -Eiq 'https?://|\bfetch\(|\bhttpx\b|\brequests\.|\bundici\b|\baxios\b|Authorization|Bearer ' "$f" 2>/dev/null; then
          printf '%s\n' "$f" >> "$FILES"
        fi
      done < "$ALL"
      
      if [ ! -s "$FILES" ]; then
        skip "no connector-like source (.py/.ts/.js with an HTTP client) under: $SCAN_PATH"
        printf '\nok=%d skip=%d warn=%d fail=%d\n' "$ok_count" "$skip_count" "$warn_count" "$fail_count"
        exit 0
      fi
      
      # --- per-file checks ---------------------------------------------------------
      while IFS= read -r f; do
        [ -z "$f" ] && continue
        file_ok=1
      
        # 1. Hardcoded secret: an inline literal assigned to a key/token/secret-looking
        #    name, OR a Bearer/sk- literal embedded in source. Reading from env is fine.
        if grep -Eni \
            '(api[_-]?key|secret|token|bearer|password|client[_-]?secret)[[:space:]]*[:=][[:space:]]*["'\''][A-Za-z0-9_\-]{16,}["'\'']' \
            "$f" 2>/dev/null | grep -Eiv 'process\.env|os\.environ|getenv|import\.meta\.env|<[^>]*>|YOUR_|EXAMPLE|xxxx|\.\.\.' >/dev/null; then
          fail "hardcoded secret literal: $f"
          file_ok=0
        fi
        if grep -Eni '["'\''](sk_live_|sk_test_|ghp_|xox[baprs]-|AKIA)[A-Za-z0-9_\-]{8,}' "$f" 2>/dev/null >/dev/null; then
          fail "hardcoded provider token literal: $f"
          file_ok=0
        fi
      
        # 2. Retry mechanism wired somewhere in the file.
        if grep -Eiq 'tenacity|@retry|wait_exponential|backoff|retry_if|maxRetries|max_retries|interceptors\.|retry[_-]?interceptor|exponential' "$f" 2>/dev/null; then
          : # has a retry primitive
        elif grep -Eiq 'for[[:space:]].*attempt|while[[:space:]].*attempt|range\([[:space:]]*[0-9]|attempt[[:space:]]*[+<]' "$f" 2>/dev/null \
             && grep -Eiq 'sleep|setTimeout|delay|backoff' "$f" 2>/dev/null; then
          : # has a hand-rolled retry loop with a delay
        else
          fail "no retry mechanism wired (no tenacity/backoff/interceptor or retry loop): $f"
          file_ok=0
        fi
      
        # 3. Pagination loop, only required when the file uses pagination params.
        if grep -Eiq 'cursor|next[_-]?cursor|endCursor|hasNextPage|pageInfo|offset|\bpage\b|\bafter\b|rel="next"|Link' "$f" 2>/dev/null; then
          if grep -Eiq 'while[[:space:]]|for[[:space:]]|yield|async[[:space:]]*function\*|def[[:space:]]+iter|->[[:space:]]*Generator|AsyncIterator' "$f" 2>/dev/null; then
            : # walks the pages
          else
            fail "pagination params used but no loop to walk them (first-page-only): $f"
            file_ok=0
          fi
        fi
      
        # 4. Request timeout set on the HTTP calls.
        if grep -Eiq 'timeout|AbortSignal\.timeout|AbortController|signal[[:space:]]*[:=]|connect_timeout|read_timeout' "$f" 2>/dev/null; then
          : # a timeout is configured
        else
          fail "no request timeout set (a hung socket will stall the run): $f"
          file_ok=0
        fi
      
        # 5. Token in localStorage (long-lived secret in an XSS-readable store).
        if grep -Eiq 'localStorage\.(setItem|getItem)?[[:space:]]*\(?[^)]*\b(token|bearer|access[_-]?token|refresh[_-]?token|api[_-]?key)\b' "$f" 2>/dev/null \
           || grep -Eiq 'localStorage[^;]*\b(token|access_token|refresh_token|apiKey)\b' "$f" 2>/dev/null; then
          fail "token stored in localStorage (XSS-readable): $f"
          file_ok=0
        fi
      
        # 6. WARN: logging the token / Authorization header.
        if grep -Eni '(console\.(log|info|debug|warn|error)|logg?(er)?\.|print|logging\.)[^;\n]*\b(token|bearer|authorization)\b' "$f" 2>/dev/null \
           | grep -Eiv 'x-request-id|status|attempt|request[_-]?id|//|#' >/dev/null; then
          warn "possible token/Authorization logging: $f"
        fi
      
        if [ "$file_ok" -eq 1 ]; then
          ok "connector shape ok: $f"
        fi
      done < "$FILES"
      
      printf '\nok=%d skip=%d warn=%d fail=%d\n' "$ok_count" "$skip_count" "$warn_count" "$fail_count"
      
      if [ "$fail_count" -gt 0 ]; then exit 1; fi
      if [ "$STRICT" -eq 1 ] && [ "$warn_count" -gt 0 ]; then exit 1; fi
      exit 0
      
  • SKILL.md 14.2 KB
    ---
    name: api-connector-builder
    description: "Use when writing a client for someone else's REST or GraphQL API: auth flow choice and token refresh, pagination to exhaustion, retry-with-jitter on transient failures only, rate-limit-aware throttling. NOT inbound callbacks (that is `webhooks`), NOT chaining services (that is `automation-flows`), NOT designing your own API (that is `api-design`)."
    tags: [api, rest, graphql, oauth, pagination, retries, rate-limiting, http-client, connector]
    recommends: [webhooks, automation-flows, api-design, data-scraper, secure-coding, error-handling, structured-extraction]
    origin: risco
    ---
    
    # API connector builder — auth, pagination, retries, rate limits
    
    You are writing a client for **someone else's** HTTP API. You do not own the
    contract; you obey it. The deliverable is one typed connector module per vendor
    that authenticates, walks the whole result set, retries only transient failures
    with backoff, and stays under the rate limit without getting the key banned.
    One connector = one vendor: mixed clients tangle two auth schemes and two
    rate-limit budgets into something no one can reason about.
    
    Four pillars, every time: **auth, pagination, retries, rate limits.** If your
    connector skips any one of them it works in the demo and breaks in production —
    on page 2, on token expiry, on the first 429, or on a flaky network.
    
    ## Step 0 — read the contract
    
    Read the vendor docs before writing a line — invented endpoints and guessed
    field names 404 in prod. Extract these first; each one changes what you write,
    so missing one means a rewrite.
    
    | Find in their docs        | Why it changes your code                                  |
    | ------------------------- | --------------------------------------------------------- |
    | Auth scheme + token TTL   | Picks the flow below; TTL decides if you need refresh     |
    | Base URL **and version**  | Wrong version = silent 404s or deprecated field shapes    |
    | Rate-limit header names   | You cannot throttle to a budget you cannot read           |
    | Pagination style          | Cursor vs offset vs Link vs Relay = different loop         |
    | Error envelope shape      | Where the real error code lives (body, not always status) |
    | Idempotency support       | Decides whether POST is safe to retry (key header?)       |
    
    If a fact is not in the docs, probe one real call and read the response headers
    and body — do not assume.
    
    ## Auth — pick the flow, then store the secret right
    
    OAuth 2.1 is the current baseline. Three deltas you must honor: **PKCE is
    mandatory for every authorization-code client**, the **Implicit** and
    **Resource-Owner-Password** grants are **removed**, and **bearer tokens may not
    travel in query strings** (header only) — query strings end up in logs and
    referrers. (oauth.net/2.1, accessed 2026-06-02.)
    
    | Flow                          | Pick it when                                | Secret lives        | Refresh strategy                       |
    | ----------------------------- | ------------------------------------------- | ------------------- | -------------------------------------- |
    | **API key in header**         | Simple server-to-server, vendor issues key  | env / secret store  | none; rotate manually                  |
    | **Bearer static token**       | Personal access token, long-lived           | env / secret store  | none; treat as a key                   |
    | **OAuth2 Client Credentials** | Machine-to-machine, no end user             | client_id + secret  | re-mint on 401; cache until expiry     |
    | **OAuth2 Auth-Code + PKCE**   | Acting on behalf of a user                  | refresh token       | rotate on every refresh (see below)    |
    | **Device flow**               | CLI / TV / input-constrained device         | refresh token       | poll then rotate as auth-code          |
    
    **Refresh-token rule (public clients):** refresh tokens must be **sender-
    constrained or one-time-use** — rotated on every refresh, with the old one
    invalidated. Keep access tokens short-lived. Never log a raw token; log a partial
    or hash if you must correlate. (OAuth 2.0 Security BCP / RFC 9700, accessed
    2026-06-02.) Full per-flow walkthroughs, token storage, and a DPoP note are in
    `references/auth-flows.md`.
    
    ## Retries + backoff
    
    Retry **only transient failures**, and only when the operation is safe to repeat.
    
    | Status / error                       | Retry? | Note                                          |
    | ------------------------------------ | ------ | --------------------------------------------- |
    | 429 Too Many Requests                | yes    | honor `Retry-After` (see rate limits)         |
    | 502 / 503 / 504                      | yes    | server-side transient                         |
    | Connection reset / timeout / DNS     | yes    | network transient                             |
    | 400 / 401 / 403 / 404 / 409 / 422    | **no** | replay returns the same error — fix the call  |
    | 2xx                                  | n/a    | success                                       |
    
    **Idempotency gate.** GET/PUT/DELETE are idempotent by semantics and safe to
    retry. **POST is not** — only retry it if you send a stable **idempotency key**
    so the server dedupes the duplicate. (AWS Builders' Library, accessed 2026-06-02.)
    
    **Backoff with jitter.** Use `delay = min(cap, base * 2^attempt) + random_jitter`.
    The jitter is not optional: without it, every client that failed at the same
    instant retries at the same instant — a synchronized thundering herd that DDoSes
    the recovering server. Full or decorrelated jitter is preferred. Always set stop
    conditions: **max attempts 3–5, a total deadline, and a per-attempt timeout.**
    (AWS Builders' Library, accessed 2026-06-02.)
    
    ```python
    # Bad: retries everything, flat sleep, no jitter, no cap, no deadline.
    for _ in range(10):
        r = httpx.get(url)
        if r.status_code == 200:
            return r.json()
        time.sleep(1)  # 4xx will never recover; herds synchronize on flat 1s
    ```
    
    ```python
    # Good: transient-only, exponential backoff WITH jitter, bounded.
    import random, time, httpx
    
    RETRYABLE = {429, 502, 503, 504}
    
    def get(url, *, attempts=5, base=0.5, cap=20.0, deadline=60.0):
        started = time.monotonic()
        for attempt in range(attempts):
            try:
                r = httpx.get(url, timeout=10.0)  # per-attempt timeout
            except (httpx.ConnectError, httpx.ReadTimeout):
                pass  # network transient -> fall through to backoff
            else:
                if r.status_code == 200:
                    return r.json()
                if r.status_code not in RETRYABLE:
                    r.raise_for_status()  # 4xx: do not retry, surface it
            if time.monotonic() - started > deadline:
                raise TimeoutError("retry deadline exceeded")
            delay = min(cap, base * 2 ** attempt) + random.uniform(0, base)
            time.sleep(delay)
        raise RuntimeError("max attempts exhausted")
    ```
    
    In Python prefer `tenacity` 9.1.4 (`@retry` with `wait_exponential_jitter`,
    `stop_after_attempt`, `retry_if_exception_type`) over a hand loop; in Node/TS use
    `undici` (the engine behind global `fetch()` since Node 18) with its retry
    interceptor. (PyPI/tenacity, nodejs/undici, accessed 2026-06-02.)
    
    ## Rate limits
    
    Do not guess the wait — the server tells you. **Precedence:**
    
    1. **`Retry-After` present** (on a 429 or 503) → wait exactly that. It is either
       a number of seconds or an HTTP-date; handle both.
    2. **No `Retry-After`** → compute the wait from `X-RateLimit-Reset` (a Unix epoch
       or a delta-seconds, per vendor docs).
    3. **Proactively throttle** on `X-RateLimit-Remaining` — slow down *before* you
       hit zero rather than absorbing a wall of 429s. (iotools.cloud rate-limiting
       guidance, accessed 2026-06-02.)
    
    ```python
    # Bad: ignore the headers, hammer, eat 429s, get the key throttled or banned.
    while more:
        resp = client.get(next_url)   # no remaining check, no Retry-After
        process(resp)
    ```
    
    ```python
    # Good: header-aware. Honor Retry-After, else reset; brake near the limit.
    def wait_for_rate_limit(resp):
        ra = resp.headers.get("Retry-After")
        if ra is not None:
            return float(ra) if ra.isdigit() else _seconds_until_httpdate(ra)
        remaining = int(resp.headers.get("X-RateLimit-Remaining", "1"))
        if remaining <= 1:
            reset = float(resp.headers.get("X-RateLimit-Reset", "0"))
            return max(0.0, reset - time.time())  # or reset directly if delta-seconds
        return 0.0
    ```
    
    For sustained pulls, gate every request through a **token-bucket** sized to the
    documented budget (e.g. 60 req/min → refill 1 token/sec, capacity 60) so bursts
    smooth out instead of slamming the wall.
    
    ## Pagination — loop to exhaustion
    
    There is no single best strategy; the vendor chose one and you follow it. The
    universal rule: **loop until the API says there is no next page — never a fixed
    page count.** A hardcoded `for page in range(10)` silently drops everything after
    page 10.
    
    | Style                      | Signal of "next"                          | Trade-off                                   |
    | -------------------------- | ----------------------------------------- | ------------------------------------------- |
    | **Offset / page number**   | `?offset=` / `?page=` until empty page    | simple; drifts on inserts, slow at depth    |
    | **Cursor / keyset**        | opaque `next_cursor` in body              | stable under writes, fixed cost — prefer it |
    | **Link header (REST)**     | `Link: <...>; rel="next"`                 | parse the header, follow until no `next`    |
    | **Relay (GraphQL)**        | `pageInfo.hasNextPage` + `endCursor`      | pass `endCursor` as `after` next query      |
    
    (graphql.org/learn/pagination + pagination pattern guides, accessed 2026-06-02.)
    
    Expose results as a generator / async-iterator so callers stream instead of
    buffering everything:
    
    ```python
    def iter_records(client):
        cursor = None
        while True:
            page = client.get("/items", params={"cursor": cursor, "limit": 100})
            body = page.json()
            yield from body["data"]
            cursor = body.get("next_cursor")
            if not cursor:           # exhaustion signal, not a counter
                return
    ```
    
    Full code for every style — offset, keyset, `Link`-header parsing, GraphQL Relay
    connection walking, in Python and TS, plus dedup-on-overlap notes — is in
    `references/pagination.md`.
    
    ## Putting it together
    
    A minimal connector wires all four pillars plus env config and structured logging.
    
    ```python
    # connector.py — Python: httpx + tenacity
    import os, logging, httpx
    from tenacity import retry, stop_after_attempt, wait_exponential_jitter, retry_if_exception_type
    
    log = logging.getLogger("connector")
    TOKEN = os.environ["VENDOR_API_TOKEN"]          # from env, never hardcoded
    
    class Transient(Exception): ...
    
    @retry(stop=stop_after_attempt(5),
           wait=wait_exponential_jitter(initial=0.5, max=20),
           retry=retry_if_exception_type(Transient))
    def _request(client, method, path, **kw):
        r = client.request(method, path, timeout=10.0, **kw)   # per-request timeout
        log.info("req id=%s %s status=%s", r.headers.get("x-request-id"), path, r.status_code)
        if r.status_code in (429, 502, 503, 504):
            raise Transient(r.status_code)
        r.raise_for_status()
        return r
    
    def client():
        return httpx.Client(base_url="https://api.vendor.com/v2",
                            headers={"Authorization": f"Bearer {TOKEN}"})  # header, not query
    ```
    
    ```typescript
    // connector.ts — Node/TS: global fetch (undici) + bounded retry
    const TOKEN = process.env.VENDOR_API_TOKEN!;          // from env, never hardcoded
    const RETRYABLE = new Set([429, 502, 503, 504]);
    
    export async function request(path: string, init: RequestInit = {}, attempt = 0): Promise<Response> {
      const res = await fetch(`https://api.vendor.com/v2${path}`, {
        ...init,
        headers: { Authorization: `Bearer ${TOKEN}`, ...init.headers },
        signal: AbortSignal.timeout(10_000),              // per-request timeout
      });
      console.info(JSON.stringify({ id: res.headers.get("x-request-id"), path, status: res.status, attempt }));
      if (RETRYABLE.has(res.status) && attempt < 4) {
        const ra = Number(res.headers.get("retry-after"));
        const wait = Number.isFinite(ra) && ra > 0 ? ra * 1000 : Math.min(20_000, 500 * 2 ** attempt) + Math.random() * 500;
        await new Promise((r) => setTimeout(r, wait));
        return request(path, init, attempt + 1);
      }
      return res;                                          // caller checks res.ok / paginates
    }
    ```
    
    ## Anti-patterns
    
    | Anti-pattern                          | Consequence                              | Fix                                          |
    | ------------------------------------- | ---------------------------------------- | -------------------------------------------- |
    | Retry 4xx (401/403/404/422)           | Burns attempts; same error every time    | Retry only 429 + 5xx + network errors        |
    | `for page in range(N)` fixed loop     | Silently drops records past page N       | Loop on cursor / Link / hasNextPage          |
    | Flat `sleep(1)` between retries        | Synchronized herd hammers recovering API | Exponential backoff **with** jitter + cap    |
    | Log the token / Authorization header  | Leaked credential in shipped logs        | Log request id + status + attempt only       |
    | Hardcode the API key in source        | Leaks via git; cannot rotate cleanly     | Read from env / secret store                 |
    | No request timeout                    | One hung socket stalls the whole run     | Set a per-request timeout always             |
    | Ignore `Retry-After`                  | Keep 429-ing; key gets throttled/banned  | Honor `Retry-After`, else `X-RateLimit-Reset`|
    | Re-POST on retry with no idempotency  | Duplicate charges / records              | Send an idempotency key, or do not retry POST|
    | Token in query string                 | Token ends up in logs / referrers        | Bearer in the `Authorization` header         |
    | Implicit / password OAuth grant       | Removed in OAuth 2.1; insecure           | Auth-Code + PKCE, or Client Credentials      |
    
    ## Verify
    
    Run `scripts/verify.sh` over the connector you write: it greps for hardcoded
    secrets, asserts a retry mechanism and a pagination loop and a request timeout
    exist, and flags `localStorage` token storage or plaintext token logging. It is a
    structure linter (read-only), not a behavior test — exit 0 on a clean target.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related