{"slug":"semgrep-rule-variant-creator","title":"semgrep-rule-variant-creator","summary":"Creates language variants of existing Semgrep rules. Use when porting a Semgrep rule to specified target languages. Takes an existing rule and target languages as input, produces independent rule+test directories for each language.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-06T18:19:41.381031Z","repo":{"url":"https://github.com/trailofbits/skills","stars":7234,"forks":616,"license":"CC-BY-SA-4.0","updatedAt":"2026-09-25T07:34:17Z"},"bodyHtml":"<hr>\n<h2>name: semgrep-rule-variant-creator\ndescription: Creates language variants of existing Semgrep rules. Use when porting a Semgrep rule to specified target languages. Takes an existing rule and target languages as input, produces independent rule+test directories for each language.\nallowed-tools: Bash Read Write Edit Glob Grep WebFetch Workflow</h2>\n<h1>Semgrep Rule Variant Creator</h1>\n<p>Port an existing Semgrep rule to other languages, one independent test-driven cycle per\nlanguage.</p>\n<p>For a new rule rather than a port, use <code>semgrep-rule-creator</code> — it takes a bug pattern\ndescription where this skill takes a finished rule. That skill is also the reference for\nrule-writing fundamentals: taint mode versus pattern matching, why tests come first, and\nhow to narrow a rule once it passes. Porting applies those same judgments in a new\nlanguage, so start there when the rule structure itself is the open question.</p>\n<h2>Run it as a workflow</h2>\n<p>Porting is the same four phases repeated per language, so the orchestration ships as a\ndynamic workflow rather than as instructions to re-follow each run:</p>\n<pre><code>/semgrep-rule-variant-creator:port-rule-to-languages\n</code></pre>\n<p>Pass the three required arguments, and <code>outputDir</code> unless the working directory is where you\nwant the variants. One language per entry: <code>\"Go and Java\"</code> ports a single language named after\nthe phrase, and the script rejects it.</p>\n<p><code>referencesDir</code> has to be a resolved absolute path. Resolve it here, because no workflow script\ncan expand a variable. Try in order, first hit wins — the <code>-d</code> is the point, since a bare <code>ls</code>\nprints the names of the files inside the directory rather than the directory itself and leaves\nnothing to copy:</p>\n<ol>\n<li><strong>Claude Code</strong> — <code>ls -d -- \"${CLAUDE_PLUGIN_ROOT}/skills/semgrep-rule-variant-creator/references\"</code></li>\n<li><strong>Codex</strong> — the same command with <code>${CODEX_PLUGIN_ROOT}</code>, if that variable is set instead</li>\n<li><strong>Neither set</strong> — <code>find ~/.claude ~/.codex . -type d -path '*/semgrep-rule-variant-creator/skills/*/references' -print -quit 2&gt;/dev/null</code></li>\n</ol>\n<p>Then confirm the directory that printed holds both reference files, with <code>ls -1 -- \"&lt;that path&gt;\"</code>.</p>\n<p>Pass the path exactly as printed. If all three come back empty, stop and say so rather than\nassembling a path by hand: the script rejects a relative path and an unexpanded token, but a\nhand-built absolute path that happens not to exist clears every guard it has, and the run then\nreports every language as passed having read no guidance at all.</p>\n<pre><code>{\n  \"rulePath\": \"&lt;path to the rule being ported&gt;\",\n  \"languages\": [\"Go\", \"Java\"],\n  \"referencesDir\": \"&lt;the absolute path the ls above printed&gt;\",\n  \"outputDir\": \"&lt;where the variant directories should land&gt;\"\n}\n</code></pre>\n<p>A workflow script cannot expand <code>{baseDir}</code> or <code>${CLAUDE_PLUGIN_ROOT}</code>, and has no filesystem\naccess to notice that it did not; an installed plugin does not sit in the user's project\neither, so <code>referencesDir</code> is the only route by which the references below reach the phase\nagents. The script rejects a run that omits it and one that passes a token instead of a path,\nrather than porting without them, since a port made without this guidance still reports every\nlanguage as passed. <code>outputDir</code> is the one optional argument, defaulting to the working\ndirectory, which is rarely what you want inside a repository.</p>\n<p>It reads the rule once, then runs each language through the full cycle independently, and\nreports which languages passed, which failed validation, which it judged not applicable, which\nSemgrep cannot analyze at all, and which it stopped on — a language key it does not recognize,\ntwo entries resolving to one directory, or a refuter that never reported back. A stop names\nwhat to change and will happen again on a re-run, which is what separates it from an agent that\ndied. The rule travels as a path, not as text: every phase\nreads the file, because an agent asked to repeat a rule back verbatim does not — one\nHTML-escaped <code>&lt;</code> and <code>&gt;</code> and broke the <code>&lt;... ...&gt;</code> operator for every phase downstream.</p>\n<p>If a run is interrupted while the session is still alive — you stopped it, or an agent hit a\nterminal error — relaunch it with\n<code>Workflow({scriptPath: \"…\", resumeFromRunId: \"&lt;runId&gt;\", args: {…}})</code>, passing the same\narguments again. Arguments are not saved with a run, so a resume that omits them fails the\npre-flight check above before replaying anything; with them, languages that finished replay\nfrom cache and only the unfinished ones re-run.</p>\n<p>Resume is same-session only, which rules it out for the interruption a long port is most\nlikely to hit: a session limit ends the session, and runs are stored under that session's own\ndirectory, so the next session cannot reach them. A run id it cannot resolve is not an error\neither — the workflow starts from scratch under that id and re-runs every language at full\ncost, with nothing saying so. Check the id is still there before counting on a resume:</p>\n<pre><code>ls -d ~/.claude/projects/*/*/subagents/workflows/*/\n</code></pre>\n<p>That is where runs land today rather than a documented interface, so an empty result may mean the\nlayout moved rather than that the run is gone. The safe reading is the same either way: if you\ncannot confirm the id, or the session ended, re-invoke with the same arguments and point\n<code>outputDir</code> somewhere fresh. The script never deletes a directory, so a language that flipped to\n<code>NOT_APPLICABLE</code> on the second run leaves the first run's variant behind.</p>\n<p>The script is <code>workflows/port-rule-to-languages.js</code> at the plugin root. It pins a\nreasoning effort per phase — cheap to read the rule, highest for translation and for the\nfix-until-green loop — and encodes the phase order, so a rule cannot be written before\nthe tests that specify it. It also keeps the two decisions that have no oracle out of any\nsingle agent's hands: a <code>NOT_APPLICABLE</code> verdict goes to an independent refuter before the\nlanguage is dropped, and failed validation is retried up to three times rather than\ntrusting one agent to iterate until green.</p>\n<p>Run the phases by hand when you are porting to a single language and want to stay in the\nloop, or when a port is already half-finished and you only need one phase. The workflow is\nthe only delegation a port needs: one agent to read the rule, and four per language when\nthe port goes green first try — a refuted verdict adds one, and so does each validation\nretry. Nothing else here is large or independent enough to be worth its own agent, so\nrunning a phase by hand means doing it yourself rather than handing it to a subagent.</p>\n<h2>The four phases</h2>\n<p>Each language runs all four before its variant is finished. A language that fails\nvalidation is unfinished; a language judged not applicable produces no directory.</p>\n<p><strong>1. Applicability analysis</strong> — decide whether the pattern belongs in the target language\nat all: does the vulnerability class exist there, does an equivalent construct exist for\neach source, sink, and sanitizer, and would the ported rule detect real risk rather than\na surface syntax match. Verdict is <code>APPLICABLE</code>, <code>APPLICABLE_WITH_ADAPTATION</code>, or\n<code>NOT_APPLICABLE</code>. <code>NOT_APPLICABLE</code> is the one verdict nothing downstream can contradict —\nit produces no tests, no rule, and no directory — so it earns a second opinion before you\nact on it. Answered separately, by running Semgrep: can Semgrep read this language at all?\nPerl has no frontend and Elixir's parser is Pro-only, and in both cases the bug class is\npresent while the rule is ungradeable — a different finding from <code>NOT_APPLICABLE</code>, which\nclaims the bug class is absent. See\n<a href=\"%7BbaseDir%7D/references/applicability-analysis.md\">applicability-analysis.md</a>\nfor worked examples of each verdict.</p>\n<p><strong>2. Test creation</strong> — write the test file first, in idiomatic target-language code. At\nleast two <code>ruleid:</code> cases and two <code>ok:</code> cases, each annotation on the line immediately\nabove the code it grades. Include the safe form that is the language's own idiom for\ndoing the thing correctly, since that is the false positive a port most often invents.</p>\n<p><strong>3. Rule translation</strong> — dump the AST for the target language and translate against what\nit shows, because pattern shape follows AST shape rather than source resemblance. Keep the\noriginal's detection intent and mode; change the id to <code>&lt;original-id&gt;-&lt;language&gt;</code>, the\n<code>languages</code> key, and add <code>original-rule</code> and <code>ported-from</code> metadata. See\n<a href=\"%7BbaseDir%7D/references/language-syntax-guide.md\">language-syntax-guide.md</a>.</p>\n<p><strong>4. Validation</strong> — <code>semgrep --test</code> is the acceptance criterion, and it must report that\nall tests passed. Missed lines mean the pattern is narrower than the vulnerability;\nincorrect lines mean it is broader. The test file is the specification, so fix the rule to\nsatisfy it. Stopping while tests still fail leaves the language unfinished, not done. See\n<a href=\"%7BbaseDir%7D/references/workflow.md\">workflow.md</a> for reading a test failure and for\ntroubleshooting when a pattern will not match or taint will not propagate.</p>\n<p>The acceptance criterion is one specific Semgrep: the version recorded when the rule was\nread. Switching binaries to get a green is the failure this guards against — an agent that\ncould not pass its Elixir tests installed the last OSS build shipping the Elixir parser and\nreported its genuine \"All tests passed\" for a port that is red here. Two other greens mean\nnothing: a rule Semgrep <em>skipped</em> still ends its run in \"All tests passed\", and so does a\ntest file whose extension Semgrep does not associate with the rule's language, since it\ngraded zero tests either way.</p>\n<h2>Output</h2>\n<p>One directory per applicable language, holding the ported rule and its test file:</p>\n<pre><code>python-command-injection-go/\n├── python-command-injection-go.yaml\n└── python-command-injection-go.go\n</code></pre>\n<p><code>All tests passed</code> means the rule and its test file agree with each other; it is not\nevidence that the vulnerability class is exploitable in the target language, since the same\ncycle wrote both, so treat a finished variant as a candidate for review rather than a\nvalidated rule.</p>\n<h2>Scope and reporting</h2>\n<p>Port the rule you were handed to the languages you were asked for. Do not repair the\noriginal, widen it to catch a nearby bug class, or add a language nobody named; if the\noriginal looks wrong or an obvious target is missing, say so in one sentence and carry on\nwith the port as asked. Every language you were given gets finished — a port is done when\nits tests pass, not when its files exist.</p>\n<p>Keep prose short and spend it on the result. Before the first tool call, say in one\nsentence what you are about to do. While a port runs, speak up when a verdict changes,\nwhen the target needs a pattern shape the original does not have, or when the tests will\nnot go green — not on every <code>semgrep --test</code> iteration. Then lead with the outcome: the\nfirst sentence says which languages passed, which failed validation, which were not\napplicable, and which Semgrep cannot analyze, with the detail after it. Correct an earlier statement when the error changes\nthe rule, the verdict, or what to do next, then keep going; a slip that changes nothing\nneeds no note.</p>\n<p>Rule and test files are the size of the problem. A test file earns its length from\ndistinct constructs and distinct safe forms rather than from restatements of the same\ncase, and neither file needs comments repeating what the code already says.</p>\n<h2>Rationalizations to Reject</h2>\n<table>\n<thead>\n<tr>\n<th>Rationalization</th>\n<th>Why It Fails</th>\n<th>Correct Approach</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>\"Pattern structure is identical\"</td>\n<td>Different ASTs across languages</td>\n<td>Always dump AST for target language</td>\n</tr>\n<tr>\n<td>\"Same vulnerability, same detection\"</td>\n<td>Data flow differs between languages</td>\n<td>Analyze target language idioms</td>\n</tr>\n<tr>\n<td>\"Rule doesn't need tests since original worked\"</td>\n<td>Language edge cases differ</td>\n<td>Write NEW test cases for target</td>\n</tr>\n<tr>\n<td>\"Skip applicability - it obviously applies\"</td>\n<td>Some patterns are language-specific</td>\n<td>Complete applicability analysis first</td>\n</tr>\n<tr>\n<td>\"I'll create all variants then test\"</td>\n<td>Errors compound, hard to debug</td>\n<td>Finish each language before the next</td>\n</tr>\n<tr>\n<td>\"Library equivalent is close enough\"</td>\n<td>Surface similarity hides differences</td>\n<td>Verify API semantics match</td>\n</tr>\n<tr>\n<td>\"Just translate the syntax 1:1\"</td>\n<td>Languages have different idioms</td>\n<td>Research target language patterns</td>\n</tr>\n<tr>\n<td>\"Most tests pass\"</td>\n<td>A partial rule reports partial truth</td>\n<td><code>All tests passed</code>, or the port is unfinished</td>\n</tr>\n<tr>\n<td>\"An older semgrep still parses this language\"</td>\n<td>A green nobody can reproduce on the semgrep the rule must run under</td>\n<td>Report the failing output and say the parser is Pro-only</td>\n</tr>\n<tr>\n<td>\"The class exists there, so the rule ports\"</td>\n<td>Semgrep has no Perl frontend and Elixir's is Pro-only; taint no-ops silently</td>\n<td>Confirm semgrep can read the language before porting</td>\n</tr>\n<tr>\n<td>\"Semgrep said all tests passed\"</td>\n<td>It says that over zero graded tests, for a rule it skipped or a file it never matched</td>\n<td>Check the rule ran and the test file's extension matches</td>\n</tr>\n</tbody>\n</table>\n<h2>Quick Reference</h2>\n<table>\n<thead>\n<tr>\n<th>Task</th>\n<th>Command</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Run tests</td>\n<td><code>semgrep --test --config rule.yaml test-file</code></td>\n</tr>\n<tr>\n<td>Validate YAML</td>\n<td><code>semgrep --validate --config rule.yaml</code></td>\n</tr>\n<tr>\n<td>Dump AST</td>\n<td><code>semgrep --dump-ast -l &lt;lang&gt; &lt;file&gt;</code></td>\n</tr>\n<tr>\n<td>Debug taint flow</td>\n<td><code>semgrep --dataflow-traces -f rule.yaml file</code></td>\n</tr>\n</tbody>\n</table>\n<h2>Documentation</h2>\n<ul>\n<li><a href=\"https://semgrep.dev/docs/writing-rules/pattern-syntax\">Pattern Syntax</a> — metavariables and matching</li>\n<li><a href=\"https://semgrep.dev/docs/writing-rules/pattern-examples\">Pattern Examples</a> — per-language references, the most useful page when translating</li>\n<li><a href=\"https://semgrep.dev/docs/writing-rules/testing-rules\">Testing Rules</a> — annotation semantics</li>\n<li><a href=\"https://appsec.guide/docs/static-analysis/semgrep/advanced/\">Trail of Bits Testing Handbook</a> — advanced taint patterns</li>\n</ul>\n","files":[{"path":"agents/openai.yaml","sizeBytes":254,"isText":true},{"path":"assets/trail-of-bits-mark.svg","sizeBytes":3084,"isText":false},{"path":"references/applicability-analysis.md","sizeBytes":10259,"isText":true},{"path":"references/language-syntax-guide.md","sizeBytes":7509,"isText":true},{"path":"references/workflow.md","sizeBytes":3970,"isText":true},{"path":"SKILL.md","sizeBytes":13749,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-17T15:59:53.242029Z","sha256":"78617BAFA47A1947536746F4D38E69C285CDD6D329CB7361D88EECACEBFEE481","sizeBytes":17199},"review":null,"source":{"repositoryUrl":"https://github.com/trailofbits/skills","path":"plugins/semgrep-rule-variant-creator/skills/semgrep-rule-variant-creator","license":"CC-BY-SA-4.0","commit":"0cc1c73a5e96749ab32d7ea5e14892fafa6972ae","subtreeSha":"6D985EDA84071794584B8D991978A2E6A43521961403684BFAA10C1A19129A0D","lastSyncedAt":"2026-09-25T07:36:46.789003Z"},"reviewedAt":"2026-09-17T16:01:14.735827Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/trailofbits/skills/tree/main/plugins/semgrep-rule-variant-creator/skills/semgrep-rule-variant-creator"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install trailofbits-skills@llmmart"},{"target":"git","command":"git clone https://github.com/trailofbits/skills.git"}]}