design-token-governance-review
Reviews design-token source of truth and build pipelines for hardcoded-value drift and resolved WCAG 1.4.3/1.4.11 contrast compliance across theme variants (light, dark, high-contrast), grounded in the W3C Design Tokens format and current WCAG success criteria.
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/frontend/design-token-governance-review
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Design Token Governance Review
Purpose
A token system only delivers its promise (consistent, themeable, accessible UI) if it is actually the single source of truth. This skill catches the two most common failure modes: hardcoded values that duplicate an existing token (silent drift) and token pairings that resolve to an inaccessible contrast ratio in some theme variant that nobody re-checked after the token value changed.
When to use
Use this skill when the user asks to:
- review a design-token source file or build/transform pipeline (Style Dictionary, Tokens Studio, W3C DTCG format),
- audit whether hardcoded colors/spacing values have crept back into components,
- verify WCAG 1.4.3 (text contrast) or 1.4.11 (non-text/UI-component contrast) compliance for token pairings,
- review a proposed rebrand, dark-mode, or high-contrast-mode token change for contrast regressions.
Context7 Documentation Protocol
Before making any claim about token-file format, $value/$type semantics, reference resolution, transform behavior, or Storybook a11y tooling, ground the claim in current library docs:
- Call
mcp__Context7__resolve-library-idfor the library in scope (Style Dictionary,Storybook) if not already resolved in this session. - Call
mcp__Context7__query-docswith a specific query before asserting version-specific behavior (transform config shape, DTCG$typeinheritance rules, a11y-addon detection coverage, etc.). - Verified library IDs for this skill's scope: Style Dictionary
/style-dictionary/style-dictionary(DTCG format,$value/$type,resolveReferences/getReferences,typeDtcgDelegatefor$typeinheritance); Storybook/storybookjs/storybook(@storybook/addon-a11y, built on axe-core, documented as catching up to ~57% of WCAG issues automatically — never represent an addon-a11y pass as full conformance). - If Context7 has no coverage for a claim, fall back to the W3C Design Tokens Community Group spec (
https://tr.designtokens.org/format/) or WCAG 2.1 normative text, and mark the claimdocumentation-basedrather thanverified. - Never invent a transform name, CLI flag, or token property that was not confirmed via Context7 or the W3C spec.
Lean operating rules
- Always compute the resolved contrast ratio from the actual token values in scope; a token's semantic name (
--color-text-secondary) does not guarantee its computed contrast is compliant. - Check every theme variant (light, dark, high-contrast) independently; a token pairing compliant in light mode can fail in dark mode with no code-level warning.
- Distinguish a documented, intentional design-system escape hatch (a one-off spacing value explicitly allowed by governance docs) from undocumented drift; only the latter is a finding.
- Apply WCAG 1.4.3 (4.5:1 normal text / 3:1 large text) to text-on-background token pairings and WCAG 1.4.11 (3:1) to UI-component and graphical-object contrast (borders, focus indicators, icons) — these are different success criteria with different thresholds; do not conflate them.
- In DTCG-format token files, resolve
{group.token}references and$typegroup-level inheritance before evaluating a value — a token's effective type/value can come from an ancestor group, not the token's own object. - Require that generated token build output (CSS custom properties, JS token modules) be committed with a reviewable diff against source tokens, not silently regenerated; a build-only change with no source-token diff is itself a governance finding.
- Treat a Storybook
@storybook/addon-a11ypass as partial automated evidence only (axe-core-class coverage, not full conformance); it does not substitute for manually verifying contrast math on token pairings the user is asking about. - Load the
visual-regression-storybook-reviewskill instead of this one when the question is about pixel-diff baseline approval rather than token source/contrast math. - This is a static-review skill: read and analyze token source files, build configs, and resolved output; do not run builds, mutate token files, or execute pipeline commands.
References
Load these only when needed:
- Token source and pipeline review — use when reviewing DTCG/Style Dictionary token structure, reference resolution,
$typeinheritance, transform/build config, or hunting for hardcoded-value drift against an existing token set. - Contrast compliance review — use when computing or verifying resolved contrast ratios for token pairings against WCAG 1.4.3/1.4.11 across theme variants.
Response minimum
Return, at minimum:
- inventory of hardcoded-value findings with file/line and the matching existing token,
- computed contrast ratio per token pairing reviewed, per theme variant, with WCAG pass/fail against the correct success criterion (1.4.3 vs 1.4.11),
- evidence level (computed from provided token values vs inferred, and Context7-grounded vs documentation-based),
- recommended governance mechanism (lint rule/CI check) to prevent recurrence,
- security caveat if the token pipeline pulls from an external API.
Files (vanguard-frontier-agentic)
-
references
-
contrast-compliance-review.md 8.1 KB
# Contrast Compliance Review Use this reference when computing or verifying resolved contrast ratios for token pairings (text-on-background, UI-component/border/icon/focus-indicator) against WCAG success criteria 1.4.3 and 1.4.11, across every theme variant a token set ships. > Version note: Storybook addon-a11y capabilities and axe-core detection coverage evolve. Verify tool behavior against installed version and current Context7 docs before asserting coverage. ## What people get wrong The naive story is: > "We ran the tokens through an automated a11y check once, so contrast is handled." Wrong, for at least three reasons: 1. **One theme is not all themes.** A token pairing that passes in light mode has no automatic bearing on the same semantic pairing resolved in dark mode or high-contrast mode — each theme resolves the same token *names* to different literal values, and each resolution must be checked independently. 2. **Automated tooling has documented, partial coverage.** Storybook's `@storybook/addon-a11y` is built on axe-core and is documented (verified via Context7, `/storybookjs/storybook`) as automatically catching up to approximately 57% of WCAG issues. A clean addon-a11y run is evidence of "no automatically detectable violation in the states exercised," not evidence of full WCAG conformance. 3. **1.4.3 and 1.4.11 are different success criteria with different scopes and thresholds.** Treating them as interchangeable produces both false passes (applying the lower non-text threshold to body text) and false alarms (applying the text threshold to a decorative icon that WCAG doesn't require to meet either). ## Officially grounded thresholds (WCAG 2.1) - **1.4.3 Contrast (Minimum):** normal text and images of text must have a contrast ratio of at least **4.5:1** against its background; large-scale text (per WCAG's definition of large text) needs at least **3:1**. This applies to text-on-background token pairings — body copy, labels, placeholder text if it conveys required information, etc. - **1.4.11 Non-text Contrast:** visual information required to identify UI components and states (borders on input fields, focus indicators, icons that convey meaning, graphical objects required to understand content) must have a contrast ratio of at least **3:1** against adjacent color(s). This applies to a different set of token pairings than 1.4.3 — component borders, focus rings, icon-on-background — and uses a flat 3:1 threshold regardless of size. - Both criteria are about **resolved, rendered** contrast — the actual sRGB values that end up on screen after every token reference and theme override is applied — not about the token's declared or "intended" relationship. ## Non-negotiable design rules ### 1. Compute the ratio; do not eyeball it or trust the token name Resolve both sides of the pairing (foreground token, background token) to final hex/RGB values for the theme variant under review, then compute the WCAG relative-luminance contrast ratio. A name like `--text-on-surface` does not guarantee the resolved pair clears any threshold — that is exactly the kind of drift this skill exists to catch after a rebrand or a single token-value edit. ### 2. Route each pairing to the correct success criterion Before computing anything, classify the pairing: is it text-on-background (1.4.3) or a UI-component/graphical-object pairing (1.4.11)? Do not apply 4.5:1 to a focus-indicator token pairing or 3:1 to body text — cite the criterion you're actually applying, and the correct threshold for it. ### 3. Every theme variant is a separate finding set Light, dark, and high-contrast (or any other shipped theme) must each be evaluated independently against the tokens that variant resolves to. A pass in one variant must never be reported as if it covers another. If a theme variant is missing from what was provided, say so explicitly as an evidence gap rather than assuming parity. ### 4. Automated-tool evidence is a floor, not a ceiling If the user supplies Storybook `addon-a11y` (or equivalent axe-core-based) results, treat a "no violations" result as partial coverage evidence only, and say so using the ~57% figure grounded via Context7. Do not upgrade a clean automated run to "WCAG 1.4.3/1.4.11 compliant" without independently computing the ratio for the specific pairings in scope. ### 5. State-dependent tokens need their own check Hover, focus, active, and disabled states often resolve to different token values than the resting state. A resting-state pass does not imply a focus-state pass — focus indicators are explicitly in scope for 1.4.11 and are a common place for contrast regressions to hide. ## Minimal safe review flow 1. Enumerate the token pairings in scope and classify each as text-on-background (1.4.3) or UI-component/graphical (1.4.11). 2. For each theme variant provided, resolve both sides of each pairing to final color values (following token references/aliases through to their literal value for that theme). 3. Compute the WCAG relative-luminance contrast ratio for each resolved pairing. 4. Compare against the correct threshold for the criterion (4.5:1 normal text / 3:1 large text for 1.4.3; 3:1 for 1.4.11), and record pass/fail per pairing per theme variant. 5. If any automated-tool evidence (addon-a11y/axe-core) was supplied, report it as supplementary partial-coverage evidence, not as the basis for the pass/fail verdict on the specific pairings you computed. 6. Flag any state (hover/focus/active/disabled) or theme variant that was not supplied as an explicit evidence gap, not a silent pass. ## Adversarial checklist Before reporting a contrast finding, answer these: - Did you resolve every token reference/alias to a literal color for the specific theme variant, or did you compute against an intermediate/aliased value? - Is this pairing actually text, or a UI component/icon/border — and does your cited threshold match that classification? - Did you check this pairing in every theme variant supplied, or only the default one? - Did you check the focus-visible state separately, given it is explicitly a 1.4.11 concern and often uses a different token than the resting border? - If you're relying on an automated tool result, did you disclose its documented partial-coverage limitation rather than presenting it as a conformance determination? If you cannot answer those, the contrast finding is not ready to report. ## High-risk assumptions to kill - "The design system says this pairing is accessible" — design-system documentation can be stale relative to the current resolved token values; compute, don't cite intent. - "It passed axe-core in Storybook, so it's WCAG compliant" — axe-core-class tooling is documented as catching a minority of WCAG issues; contrast math specifically is one of the more reliably automatable checks, but a "pass" scope is still limited to states/variants actually exercised in the stories run. - "Dark mode uses the same relationships as light mode, just inverted, so it's fine" — inversion does not preserve ratios; each variant's resolved values must be computed independently. - "It's just a decorative icon, ignore it" — if the icon is required to identify a UI component or convey state (not purely decorative), 1.4.11 applies; verify which case it is before dismissing it. ## Safe verification targets - The resolved (post-reference, post-theme-override) hex/RGB value for each side of each pairing under review, for each theme variant. - The computed contrast ratio and the specific criterion (1.4.3 vs 1.4.11) and threshold it was checked against. - Any automated-tool report supplied by the user (Storybook a11y-addon panel output, CI a11y job output), labeled as partial-coverage supplementary evidence. ## When to push back Push back if the user asks you to: - report a pairing as "compliant" based on the token name or design intent alone, with no resolved-value computation, - treat a single theme variant's result as representative of all shipped variants, - treat a clean automated a11y-addon run as a full WCAG conformance statement, - skip focus/hover/active state checks because "the resting state passed." Those are not shortcuts. They are exactly the kind of gap this review exists to close. -
token-source-and-pipeline-review.md 8.2 KB
# Token Source and Pipeline Review Use this reference when reviewing a design-token source file, a Style Dictionary (or equivalent) build/transform config, or hunting for hardcoded-value drift in components that should be referencing an existing token. > Version note: Style Dictionary's DTCG support and transform API are evolving. Verify exact config shape and CLI flags against the installed package version and current Context7-sourced docs before asserting behavior. ## What people get wrong The naive story is: > "We have a `tokens.json` file, so we have a single source of truth." Wrong. A token file is only a source of truth if: 1. every component actually consumes the *built* output of that file (not a copy-pasted value), 2. the build step is deterministic and its output is diffed/reviewed like any other generated artifact, 3. references between tokens (`{color.brand.500}` or DTCG `{group.token}`) resolve to what the author intended, including inherited `$type`, 4. there is exactly one place a given design decision (a specific blue, a specific spacing step) is expressed as a literal value. Most drift findings come from case 4 breaking silently: someone needed "almost that blue" for a one-off component, typed a hex literal instead of adding/aliasing a token, and nobody caught it because nothing greps for it. ## Officially grounded token shape (W3C DTCG + Style Dictionary) Per the W3C Design Tokens Community Group format and Style Dictionary's DTCG support (verified via Context7, `/style-dictionary/style-dictionary`): - Tokens are expressed with `$value` and `$type` (the `$`-prefixed form is the DTCG spec; Style Dictionary v3's unprefixed `value`/`type` is the legacy form — do not assume a codebase uses one or the other without checking the actual file). - `$type` can be declared on a group and inherited downward to every token in that group; a token's *effective* type is not always present on the token object itself. Style Dictionary's `typeDtcgDelegate` utility performs exactly this inheritance resolution — treat a token's type as unresolved until inheritance has been applied, not as whatever key literally appears on the leaf node. - Token values can reference other tokens using brace syntax (e.g. `{spacing.2}`, `{colors.black}`). `resolveReferences`/`getReferences` are the documented utilities for resolving and enumerating these; a raw, unresolved reference string is not a computed value and must not be treated as one when checking contrast or any numeric threshold. - Style Dictionary's DTCG conversion tooling explicitly notes that refactoring common `type` values up to a shared ancestor is a distinct, not-fully-automated step from the `value`→`$value` rename — a partially converted file (renamed values, un-hoisted types) is a legitimate mid-migration state, not automatically a defect, but it changes how you must resolve `$type` per token. ## Non-negotiable design rules ### 1. Resolve references and inheritance before judging a value Do not evaluate `$type` or `$value` from a token's own JSON object in isolation. Walk up the group chain for `$type` inheritance and resolve `{...}` references for `$value` before treating either as ground truth for a governance or contrast finding. ### 2. A literal value is not automatically drift A hardcoded value is a finding only when a token expressing the *same design decision* already exists. A one-off value with no equivalent token, explicitly scoped and documented (e.g. a documented design-system escape hatch), is not drift. Search the resolved token set for a value/near-value match before flagging. ### 3. Generated output must be reviewable, not silently regenerated CSS custom properties, JS/TS token modules, platform-specific outputs (iOS, Android) generated by the build step must be committed and diffed like source code. A change to generated output with no corresponding change to token source is itself a governance finding — it means either the build was hand-edited (breaking the source-of-truth property) or the source changed without the build being re-run and reviewed together. ### 4. Distinguish legacy Style Dictionary syntax from DTCG syntax explicitly Do not assume every token file in a repo uses the same syntax. Mixed syntax across files (or a partial per-file migration) changes how references resolve (`usesDtcg` option) and is itself worth flagging if undocumented. ### 5. Treat external token-source APIs as a supply-chain surface A pipeline that pulls tokens from a Figma/Tokens Studio API at build time introduces a live external dependency into every build. Credentials for that pull must be scoped, rotatable, and stored only in CI secrets — never inside the token-source config file itself (config files get committed and often copied across repos). ## Minimal safe review flow 1. Identify the token source format in use (legacy Style Dictionary `value`/`type` vs DTCG `$value`/`$type`) — do not assume; read the file. 2. Resolve `$type` inheritance and `{...}` references for every token pairing under review before doing anything else with the values. 3. Grep component source for literal values (hex colors, px/rem spacing, font-weight numbers) that match or nearly match a resolved token value. 4. For each match, confirm whether an equivalent token already exists; if yes, it is a drift finding, if no, confirm whether it is a documented escape hatch. 5. Check whether generated build output is committed and whether its last change correlates with a source-token change. 6. If the pipeline pulls from an external token-source API, confirm credential handling before treating the pipeline as internally sound. ## Adversarial checklist Before closing out a token-source review, answer these: - Does every component in scope actually import the *built* token output, or do some import raw source JSON directly (bypassing the build's guarantees)? - Is there a single token file, or multiple files that could disagree on the same design decision (e.g. two different "brand primary" definitions)? - If `$type` is declared only at a group level, did you resolve it before comparing the token's type to what a hardcoded value implies? - Does the build pipeline pin its Style Dictionary (or equivalent) version, or could an unpinned transform silently change output between builds? - If tokens are pulled from an external API at build time, what happens on API failure — does the build fail loud, or silently fall back to a stale cached copy? If you cannot answer those, the review is incomplete. ## High-risk assumptions to kill - "It's named `--color-text-secondary`, so it must be the secondary text color used everywhere" — semantic naming does not guarantee usage consistency; grep for actual consumption. - "The token file is DTCG format because it has `$value`" — a file can be partially migrated; check every token, not the first one. - "Generated CSS is just a build artifact, no need to review it" — a hand-edited generated file silently breaks the single-source-of-truth property. - "One-off pixel value, not worth a token" — repeated across enough components, undocumented one-offs are exactly what token systems exist to eliminate. ## Safe verification targets - Diff the token source file against the last reviewed/approved version. - Diff generated build output (CSS/JS) against the previous commit and correlate with the source-token diff. - Grep component source trees for literal color/spacing values matching resolved token values. - If Style Dictionary is in use, check the config file's `source`, `platforms`, and any custom transforms for anything that could alter resolved values silently (e.g. a custom color transform). ## When to push back Push back if the user asks you to: - approve a generated-output change with no corresponding source-token diff ("just trust the build"), - treat a partially DTCG-migrated file as fully migrated without checking every token, - skip the drift search because "we'll catch it in visual review" (visual review does not catch a hex literal that happens to render identically to its token equivalent today but will drift on the next rebrand), - store token-source API credentials directly in the pipeline config for convenience. Those are not shortcuts. They erase the property the token system exists to guarantee.
-
-
metadata.json 1.1 KB
{ "id": "design-token-governance-review", "name": "Design Token Governance Review", "type": "skill", "provider": "frontend", "harnesses": [ "claude-code", "cursor", "codex", "gemini", "kiro", "other" ], "summary": "Reviews design-token source of truth, build/transform pipelines, and resolved contrast ratios against WCAG 1.4.3/1.4.11 to stop hardcoded values and inaccessible token pairings from shipping across themes.", "source_type": "original", "official_docs": [ "https://www.w3.org/community/design-tokens/", "https://www.w3.org/TR/WCAG21/#contrast-minimum", "https://www.w3.org/TR/WCAG21/#non-text-contrast", "https://amzn.github.io/style-dictionary/#/" ], "security_notes": "Token-source pipelines pulling from Figma/Tokens Studio APIs must use scoped, rotatable personal access tokens stored only in CI secrets, never committed to token-source config files.", "last_verified": "2026-07-02", "path": "skills/frontend/design-token-governance-review", "author": "github: VincentChuWaiChow", "version": "0.1.0" } -
SKILL.md 5.5 KB
--- name: design-token-governance-review description: Reviews design-token source of truth and build pipelines for hardcoded-value drift and resolved WCAG 1.4.3/1.4.11 contrast compliance across theme variants (light, dark, high-contrast), grounded in the W3C Design Tokens format and current WCAG success criteria. allowed-tools: Read Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.0" updated: "2026-07-02" category: compliance --- # Design Token Governance Review ## Purpose A token system only delivers its promise (consistent, themeable, accessible UI) if it is actually the single source of truth. This skill catches the two most common failure modes: hardcoded values that duplicate an existing token (silent drift) and token pairings that resolve to an inaccessible contrast ratio in some theme variant that nobody re-checked after the token value changed. ## When to use Use this skill when the user asks to: - review a design-token source file or build/transform pipeline (Style Dictionary, Tokens Studio, W3C DTCG format), - audit whether hardcoded colors/spacing values have crept back into components, - verify WCAG 1.4.3 (text contrast) or 1.4.11 (non-text/UI-component contrast) compliance for token pairings, - review a proposed rebrand, dark-mode, or high-contrast-mode token change for contrast regressions. ## Context7 Documentation Protocol Before making any claim about token-file format, `$value`/`$type` semantics, reference resolution, transform behavior, or Storybook a11y tooling, ground the claim in current library docs: 1. Call `mcp__Context7__resolve-library-id` for the library in scope (`Style Dictionary`, `Storybook`) if not already resolved in this session. 2. Call `mcp__Context7__query-docs` with a specific query before asserting version-specific behavior (transform config shape, DTCG `$type` inheritance rules, a11y-addon detection coverage, etc.). 3. Verified library IDs for this skill's scope: Style Dictionary `/style-dictionary/style-dictionary` (DTCG format, `$value`/`$type`, `resolveReferences`/`getReferences`, `typeDtcgDelegate` for `$type` inheritance); Storybook `/storybookjs/storybook` (`@storybook/addon-a11y`, built on axe-core, documented as catching up to ~57% of WCAG issues automatically — never represent an addon-a11y pass as full conformance). 4. If Context7 has no coverage for a claim, fall back to the W3C Design Tokens Community Group spec (`https://tr.designtokens.org/format/`) or WCAG 2.1 normative text, and mark the claim `documentation-based` rather than `verified`. 5. Never invent a transform name, CLI flag, or token property that was not confirmed via Context7 or the W3C spec. ## Lean operating rules - Always compute the resolved contrast ratio from the actual token values in scope; a token's semantic name (`--color-text-secondary`) does not guarantee its computed contrast is compliant. - Check every theme variant (light, dark, high-contrast) independently; a token pairing compliant in light mode can fail in dark mode with no code-level warning. - Distinguish a documented, intentional design-system escape hatch (a one-off spacing value explicitly allowed by governance docs) from undocumented drift; only the latter is a finding. - Apply WCAG 1.4.3 (4.5:1 normal text / 3:1 large text) to text-on-background token pairings and WCAG 1.4.11 (3:1) to UI-component and graphical-object contrast (borders, focus indicators, icons) — these are different success criteria with different thresholds; do not conflate them. - In DTCG-format token files, resolve `{group.token}` references and `$type` group-level inheritance before evaluating a value — a token's effective type/value can come from an ancestor group, not the token's own object. - Require that generated token build output (CSS custom properties, JS token modules) be committed with a reviewable diff against source tokens, not silently regenerated; a build-only change with no source-token diff is itself a governance finding. - Treat a Storybook `@storybook/addon-a11y` pass as partial automated evidence only (axe-core-class coverage, not full conformance); it does not substitute for manually verifying contrast math on token pairings the user is asking about. - Load the `visual-regression-storybook-review` skill instead of this one when the question is about pixel-diff baseline approval rather than token source/contrast math. - This is a static-review skill: read and analyze token source files, build configs, and resolved output; do not run builds, mutate token files, or execute pipeline commands. ## References Load these only when needed: - [Token source and pipeline review](references/token-source-and-pipeline-review.md) — use when reviewing DTCG/Style Dictionary token structure, reference resolution, `$type` inheritance, transform/build config, or hunting for hardcoded-value drift against an existing token set. - [Contrast compliance review](references/contrast-compliance-review.md) — use when computing or verifying resolved contrast ratios for token pairings against WCAG 1.4.3/1.4.11 across theme variants. ## Response minimum Return, at minimum: - inventory of hardcoded-value findings with file/line and the matching existing token, - computed contrast ratio per token pairing reviewed, per theme variant, with WCAG pass/fail against the correct success criterion (1.4.3 vs 1.4.11), - evidence level (computed from provided token values vs inferred, and Context7-grounded vs documentation-based), - recommended governance mechanism (lint rule/CI check) to prevent recurrence, - security caveat if the token pipeline pulls from an external API.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.