{"slug":"suede-code-grader","title":"suede-code-grader","summary":"Suede Labs AI blunt A-F ship grade for a code change across correctness, security and permissions, data and state, domain truth, UX and release behavior, tests and verification, and deploy readiness, with Instant-F triggers and evidence-based grade caps on auth, payment, migratio","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-08-23T08:58:00.475569Z","repo":{"url":"https://github.com/JasonColapietro/suede-creator-skills","stars":125,"forks":12,"license":"MIT","updatedAt":"2026-09-24T09:44:24Z"},"bodyHtml":"<hr>\n<h2>name: suede-code-grader\ndescription: \"Suede Labs AI blunt A-F ship grade for a code change across correctness, security and permissions, data and state, domain truth, UX and release behavior, tests and verification, and deploy readiness, with Instant-F triggers and evidence-based grade caps on auth, payment, migration, and public-API surfaces. Use when asked to grade this, give it a letter, is this an A, how ready is this to ship, or should this merge — when the caller wants the verdict without a findings list. NOT FOR: findings, evidence, and fix briefs (use suede-code-review, or suede-code for findings plus grade); enforcing the verdict in CI (use suede-ci-gate); eval coverage for AI behavior (use suede-ai-eval).\"</h2>\n<h1>Suede Code Grader</h1>\n<h2>Gate policy — advisory, not blocking</h2>\n<p>Every claim-verification step, check, quality gate, and ship verdict in this\nskill is a <strong>recommendation to the user, not a control on the agent</strong>. This\npolicy governs every gate, check, verdict, and \"do not ship / publish /\nproceed\" line elsewhere in this skill:</p>\n<ul>\n<li>Run every check and report the results honestly. Verdicts (<code>ship</code>,\n<code>ship-with-caveats</code>, <code>hold</code>, letter grades, BLOCKED or OPEN items) are\nadvice attached to the work, not orders that change it.</li>\n<li>Never block, delay, skip, rewrite, or refuse the action the user asked for\nbecause a check failed or a gate said hold. Complete the requested action\nas asked, and deliver the gate output alongside it as a clearly labeled\nrecommendation.</li>\n<li>A failed gate changes what you report, never what you do.</li>\n<li>Single exception: if a finding is extremely risky — data loss, security or\ncredential exposure, legal or rights violations, payment mistakes, or\nirreversible public damage — pause, tell the user exactly what the risk is\nand what the options are, and let them pick. Their choice is final.</li>\n</ul>\n<p>Blunt A-F read on whether code is ready to ship. The output is a grade with evidence, not a lint score or a pile of style notes.</p>\n<h2>Source Truth</h2>\n<p>Read before grading. Do not grade from the PR description or commit message alone.</p>\n<p>Inspect:</p>\n<ul>\n<li>repo, branch, remote, dirty state, and relevant local guidance;</li>\n<li>diff, changed files, generated files, and touched routes or APIs;</li>\n<li>imports, callers, schemas, configs, env requirements, jobs, webhooks, scripts,\ntests, and docs that move with the change;</li>\n<li>build, test, lint, typecheck, browser, simulator, MCP, or live/API evidence\nthat directly exercises the changed behavior;</li>\n<li>published statements, rights/provenance claims, payment/wallet behavior, registry\nexpectations, royalty routing, and agent-commerce contracts when relevant.</li>\n</ul>\n<p>If live, test, or runtime checks are not practical, grade the source and mark those lanes as unverified.</p>\n<p><strong>Gate evidence is a command, not an impression.</strong> Run what the repo already ships and cite the command and its exit status: the typecheck, the configured linter on changed files, the test suite, and — for a release grade — the production build. For per-stack syntax (web/Node, MCP server, iOS/Swift, generic API), use the Gate Commands by Stack table in <strong>suede-code-review</strong> rather than inventing a command. Detect what exists and run only that; never introduce a tool the repo does not use, and never report a gate result you did not execute.</p>\n<h2>Instant-F Triggers</h2>\n<p>Check these before scoring any lane. Any single match is an automatic F — no other lanes matter until it is fixed. This list mirrors suede-code's canonical Step 1 list — change both together.</p>\n<p><strong>Secrets and credentials</strong> — hardcoded API key/secret/token/password in committed source; private key or certificate committed; OAuth/signing secret outside a secret manager.\n<strong>Injection</strong> — SQL built by string concatenation with user input; shell command from user input via exec/spawn/eval; template rendered with unescaped user input where XSS is reachable.\n<strong>Auth bypass</strong> — auth middleware with a path that skips it (early return, swallowed exception, always-true condition); permission check bypassable via request param; JWT accepting <code>alg: none</code> or a hardcoded secret.\n<strong>Payment and wallet</strong> — payment handler swallowing errors silently; webhook with no signature verification; amount or recipient from untrusted input without server-side validation.\n<strong>Data destruction</strong> — migration with DROP/destructive ALTER, no rollback, no tested restore; bulk delete/update with no WHERE or user-controlled WHERE; cache invalidation that clears production stores with no restore path.\n<strong>Plaintext sensitive data</strong> — password stored or logged in plaintext; PII to an unencrypted log/analytics pipeline; SSN/payment card/health data in a non-encrypted field.</p>\n<p>If any Instant-F pattern is present: stop, report it, mark the grade F, list the specific file and line, and do not grade remaining lanes. The grade cannot be raised by other lane performance.</p>\n<h2>Grade Lanes</h2>\n<p>Score each lane A-F, then give one overall grade. When grading non-Suede work, substitute \"domain truth\" for \"Suede truth\" — use whatever domain invariants apply (API contract truth, published-statement accuracy, data model truth).</p>\n<ul>\n<li><strong>Correctness:</strong> intended behavior, edge cases, error paths, async behavior,\nrouting, data flow, and regression risk.</li>\n<li><strong>Security and permissions:</strong> auth, secrets, payment, wallet, injection, path,\nSSRF, permission, and data exposure risks fail closed.</li>\n<li><strong>Data and state:</strong> schemas, migrations, caches, jobs, queues, webhooks,\nretries, idempotency, and state transitions stay consistent.</li>\n<li><strong>Suede truth:</strong> public copy, rights, provenance, registry-backed media,\nroyalty routing, licensing, agent-commerce, and product claims match the\nimplementation.</li>\n<li><strong>UX and release behavior:</strong> loading, empty, error, success, mobile/native,\nscreenshot, metadata, route, and user-visible states hold together.</li>\n<li><strong>Tests and verification:</strong> changed behavior has meaningful tests, builds,\nscreenshots, simulator runs, MCP checks, live/API readbacks, or named caveats.</li>\n<li><strong>Deploy readiness:</strong> env vars, feature flags, configs, migrations, rollback\nnotes, install paths, docs, and release sequencing are clear.</li>\n</ul>\n<h2>Grade Meaning</h2>\n<ul>\n<li><strong>A:</strong> All lanes pass. Behavior is verified at runtime. No known follow-ups. Example: new feature with unit + integration tests, live readback confirmed, env vars documented, rollback is trivial.</li>\n<li><strong>B:</strong> No blockers. One or more lanes have named, bounded follow-ups that do not affect correctness or safety in the current release. Example: happy-path tested but edge-case coverage is thin; or migration is forward-only but rollback risk is low and documented.</li>\n<li><strong>C:</strong> At least one lane has a real defect or unverified risk that could surface in production but is not immediately catastrophic. Hold until that lane is fixed and rechecked. Example: auth path not fully tested; or a data migration with no rollback plan on a low-traffic table; or a God object in a payment module that obscures correctness.</li>\n<li><strong>D:</strong> A serious defect exists that is likely to cause data loss, auth bypass, broken payments, or a user-visible production failure. Recommend not shipping until the defect is fixed and verified, and because these are extreme-risk categories, pause and put the choice to the user before any ship step. Example: missing auth check on a state-changing endpoint; migration with no tested rollback on a high-traffic table; payment flow that silently swallows errors.</li>\n<li><strong>F:</strong> Strongly recommend against shipping. The change breaks core behavior, introduces an Instant-F pattern, or verification evidence is absent for a critical surface. Example: hardcoded API key in source, SQL injection via string concatenation, auth middleware that can be bypassed, or a payment handler with zero test coverage and no live readback.</li>\n</ul>\n<h2>Grade Caps by Surface Type</h2>\n<p>Certain surfaces cannot receive A or B without specific evidence beyond passing CI.</p>\n<p><strong>Auth changes</strong> (login, session, token validation, middleware, role assignment, permission checks)</p>\n<ul>\n<li>A requires: explicit test coverage for the bypass/escalation path, not just the happy path. Named evidence (e.g., \"tested with expired token returns 401\", \"role escalation attempt returns 403\").</li>\n<li>B requires: happy-path tested plus named caveats on what is not tested.</li>\n<li>If neither condition is met: cap at C regardless of other lane performance.</li>\n</ul>\n<p><strong>Payment and wallet flows</strong> (checkout, subscription, refund, payout, wallet transfer, webhook)</p>\n<ul>\n<li>A requires: error path tested (failed charge, declined card, webhook replay), amount/recipient validated server-side, and no silent error swallowing.</li>\n<li>B requires: happy-path tested, error paths documented as follow-ups with named risk.</li>\n<li>If neither: cap at C.</li>\n</ul>\n<p><strong>Data migrations</strong> (schema changes, backfills, column drops, index changes on production tables)</p>\n<ul>\n<li>A requires: rollback plan documented, restore tested against a copy of production data (or explicitly waived with justification for low-risk/reversible migrations).</li>\n<li>B requires: rollback plan exists but restore is untested.</li>\n<li>If no rollback plan exists: cap at D.</li>\n</ul>\n<p><strong>Public-facing API changes</strong> (new endpoints, breaking changes, removed fields, changed auth)</p>\n<ul>\n<li>A requires: backward compatibility verified or explicit version bump with documented migration path.</li>\n<li>If breaking change with no migration path: cap at C minimum.</li>\n</ul>\n<p>State these caps explicitly in the output when they apply.</p>\n<h2>Technical Debt Indicators</h2>\n<p>Flag these patterns as part of the grade assessment:</p>\n<ul>\n<li><strong>Magic numbers/strings</strong>: constants with no name or explanation that appear in logic.</li>\n<li><strong>God objects/functions</strong>: a single function or class doing 5+ unrelated things.</li>\n<li><strong>Deep coupling</strong>: code that reaches across 3+ abstraction layers to access internals.</li>\n<li><strong>Missing abstraction</strong>: the same 20-line block duplicated in 3+ places.</li>\n<li><strong>Leaky abstraction</strong>: a module that requires callers to know its internal implementation details to use it correctly.</li>\n<li><strong>Implicit state</strong>: program behavior depends on hidden global or module-level state.</li>\n<li><strong>Dead code</strong>: functions, branches, or imports that can never be reached.</li>\n</ul>\n<p><strong>Grade impact depends on where the debt lives, not just what it is:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Pattern</th>\n<th>Location</th>\n<th>Grade Impact</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>God object (5+ unrelated concerns)</td>\n<td>Payment module</td>\n<td>D in Correctness</td>\n</tr>\n<tr>\n<td>God object</td>\n<td>Utility helper</td>\n<td>B in Correctness</td>\n</tr>\n<tr>\n<td>Missing abstraction (3+ duplicated blocks)</td>\n<td>Auth flow</td>\n<td>C in Security</td>\n</tr>\n<tr>\n<td>Missing abstraction</td>\n<td>UI component</td>\n<td>B in Correctness</td>\n</tr>\n<tr>\n<td>Deep coupling (3+ layer reach)</td>\n<td>Data migration</td>\n<td>C in Data and state</td>\n</tr>\n<tr>\n<td>Implicit global state</td>\n<td>API route handler</td>\n<td>C in Correctness</td>\n</tr>\n<tr>\n<td>Dead code</td>\n<td>Any</td>\n<td>Flag only; no grade impact unless it shadows live code</td>\n</tr>\n<tr>\n<td>Magic numbers in payment amounts</td>\n<td>Payment flow</td>\n<td>C in Correctness</td>\n</tr>\n<tr>\n<td>Magic numbers in UI spacing</td>\n<td>UI component</td>\n<td>No grade impact; flag as P3</td>\n</tr>\n</tbody>\n</table>\n<p>Do not block a ship on tech debt alone unless it directly obscures a P0/P1 bug. Name the debt in Required Upgrades and let the overall grade reflect it.</p>\n<h2>Red Flags — Stop</h2>\n<ul>\n<li>\"CI passed, round up\" — CI that never exercised the changed behavior raises nothing.</li>\n<li>\"The work was clearly hard\" — effort never moves a grade; evidence does.</li>\n<li>\"It's just a refactor\" — Instant-F triggers run on every grade, every time.</li>\n<li>\"Happy path works, call it an A\" — the grade caps exist because happy paths are never where the risk lives.</li>\n<li>\"The PR description is clear enough\" — grade the diff and its evidence, or mark the lane unverified.</li>\n</ul>\n<h2>Output Format</h2>\n<pre><code>Simple explanation:\nPlain-language summary of the grade and the one biggest reason.\n\nUsual breakdown:\nTarget:\nChange reviewed:\nRuntime surfaces:\n\nGrades:\nCorrectness: A-F\nSecurity and permissions: A-F\nData and state: A-F\nSuede truth: A-F\nUX and release behavior: A-F\nTests and verification: A-F\nDeploy readiness: A-F\nOverall: A-F\nGrade cap applied: [surface type] — [what evidence would lift the cap] | none\n\nWhy:\nEvidence-backed explanation of why the overall grade landed there.\n\nRequired upgrades:\n1. Highest-impact fix.\n2. Second fix.\n3. Third fix.\n\nVerification:\nChecked:\nNot checked:\nShip gate: ship | ship-with-caveats | hold\n</code></pre>\n<p>Ship gate follows the overall grade, mechanically: A → <code>ship</code>; B → <code>ship-with-caveats</code>; C, D, F → <code>hold</code>.</p>\n<p>To revise this grade: name what changed.\nTo bank a pattern: name what worked so it can be reused.\nSilence = accepted.</p>\n<h2>Boundaries</h2>\n<ul>\n<li>Do not block on style preferences unless they create real maintenance, behavior, accessibility, release, or product-risk cost.</li>\n<li>Do not invent tests, screenshots, live checks, deploy status, or evidence for published statements.</li>\n<li>Never report a C, D, or F without naming the required upgrade that would move the grade.</li>\n<li>Keep the grade independent. Do not raise a grade because the implementation was hard, because CI passed without exercising the changed behavior, or because the author explains the intent well.</li>\n</ul>\n<h2>Worked Example</h2>\n<p>One change graded end to end, showing how lanes combine into the overall letter, is\nin <code>references/worked-example.md</code>. Read it when a grade feels borderline and you\nneed to see the lane arithmetic on a real case.</p>\n<h2>Routing</h2>\n<ul>\n<li>Findings and fix briefs behind the grade → <strong>suede-code</strong> (combined) or <strong>suede-code-review</strong> (findings only, plus Accessibility/SEO lanes)</li>\n<li>Grade is C or below and the repo has no merge gate → <strong>suede-ci-gate</strong></li>\n<li>The change ships AI behavior with no eval coverage → <strong>suede-ai-eval</strong></li>\n<li>The change touches an MCP server, its catalog, or its tool/resource/prompt definitions → <strong>suede-mcp-qa</strong> for the live protocol suite before the grade counts as verified</li>\n</ul>\n","files":[{"path":"agents/openai.yaml","sizeBytes":439,"isText":true},{"path":"CARD.md","sizeBytes":4746,"isText":true},{"path":"references/worked-example.md","sizeBytes":7951,"isText":true},{"path":"SKILL.md","sizeBytes":13667,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"notes-only","suspicious":0,"notes":2,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-13T07:22:44.12861Z","sha256":"48C82B1A228BAFF64279C793DB22C72E15091596D7F2FA53A85E6E6790B06906","sizeBytes":12314},"review":null,"source":{"repositoryUrl":"https://github.com/JasonColapietro/suede-creator-skills","path":"skills/suede-code-grader","license":"MIT","commit":"05d69df3c3b6a51d497ba14b0be90cc8932216a7","subtreeSha":"DB024FCD7A0FC59B3BB055C81892DA268D65647D94AF909493BD9A94D3862BF6","lastSyncedAt":"2026-09-25T07:37:29.687342Z"},"reviewedAt":"2026-09-13T07:23:52.653484Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/JasonColapietro/suede-creator-skills/tree/main/skills/suede-code-grader"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install jasoncolapietro-suede-creator-skills@llmmart"},{"target":"git","command":"git clone https://github.com/JasonColapietro/suede-creator-skills.git"}]}