{"slug":"vector-forge","title":"vector-forge","summary":"Mutation-driven test vector generation. Finds implementations of a cryptographic algorithm or protocol, runs mutation testing to identify escaped mutants, then generates new test vectors that deliberately exercise the uncovered code paths. Compares before/after mutation kill rate","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-11T17:26:25.613801Z","repo":{"url":"https://github.com/trailofbits/skills","stars":7234,"forks":616,"license":"CC-BY-SA-4.0","updatedAt":"2026-09-25T07:34:17Z"},"bodyHtml":"<hr>\n<h2>name: vector-forge\ndescription: \"Mutation-driven test vector generation. Finds implementations of a cryptographic algorithm or protocol, runs mutation testing to identify escaped mutants, then generates new test vectors that deliberately exercise the uncovered code paths. Compares before/after mutation kill rates to prove vector effectiveness. Use when generating cryptographic test vectors, measuring Wycheproof coverage gaps, finding escaped mutants via mutation testing, creating cross-implementation test suites, or improving test vector coverage for crypto primitives.\"</h2>\n<h1>Vector Forge</h1>\n<p>Uses mutation testing to systematically identify gaps in test vector\ncoverage, then generates new test vectors that close those gaps.\nMeasures effectiveness by comparing mutation kill rates before and after.</p>\n<h2>When to Use</h2>\n<ul>\n<li>Generating test vectors for cryptographic algorithms or protocols</li>\n<li>Evaluating how well existing test vectors cover an implementation</li>\n<li>Finding implementation code paths that no test vector exercises</li>\n<li>Creating Wycheproof-style cross-implementation test vectors</li>\n<li>Measuring the concrete coverage value of a test vector suite</li>\n</ul>\n<h2>When NOT to Use</h2>\n<ul>\n<li>No implementations exist yet (need code to mutate)</li>\n<li>Single trivial implementation with no edge cases</li>\n<li>Testing application logic rather than algorithm implementations</li>\n<li>The algorithm has no public test vectors to compare against</li>\n</ul>\n<h2>Prerequisites</h2>\n<ul>\n<li><strong>trailmark</strong> installed — if <code>uv run trailmark</code> fails, run:\n<pre><code>uv tool install trailmark\n</code></pre>\n</li>\n</ul>\n<h1>Python snippets: uv run --with trailmark python -   (a tool env is not importable)</h1>\n<pre><code>- At least one implementation of the target algorithm in a\nlanguage with mutation testing support\n- A test harness that consumes test vectors and exercises\nthe implementation\n- A mutation testing framework for the target language\n\n---\n\n## Rationalizations to Reject\n\n| Rationalization | Why It's Wrong | Required Action |\n|-----------------|----------------|-----------------|\n| \"We have enough test vectors\" | Mutation testing proves otherwise | Run the baseline first |\n| \"The implementation's own tests are sufficient\" | Own tests often share blind spots with the impl | Cross-impl vectors catch different bugs |\n| \"FFI crates can be mutation tested at the binding layer\" | Mutations to wrappers don't affect the underlying impl | Mutate the actual implementation language |\n| \"Timeouts mean the mutation was caught\" | Timeouts are ambiguous — could be killed or alive | Resolve timeouts before drawing conclusions |\n| \"All mutants are equivalent\" | Most aren't — verify by reading the mutation | Classify each escaped mutant individually |\n| \"Checking valid vectors is enough\" | Permissive mutations survive without negative assertions | Assert rejection for every invalid vector |\n| \"Manual analysis is fine\" | Manual analysis misses what tooling catches | Install and run the tools |\n\n---\n\n## Workflow Overview\n\n</code></pre>\n<p>Phase 1: Discovery       → Find implementations to test\n↓\nPhase 2: Harness         → Write/adapt test vector harness for each impl\n↓\nPhase 3: Baseline        → Run mutation testing with existing vectors\n↓\nPhase 4: Escape Analysis → Classify escaped mutants by code path\n↓\nPhase 5: Vector Gen      → Create test vectors targeting escapes\n↓\nPhase 6: Validation      → Re-run mutation testing, compare before/after\n↓\nOutput: Coverage Report + New Test Vectors</p>\n<pre><code>\n---\n\n## Phase 1: Discovery\n\nFind implementations of the target algorithm. Look for:\n\n1. **Pure implementations** in high-level languages (Go, Rust, Python)\n   — these are the best mutation testing targets\n2. **FFI wrapper crates** — identify these early so you don't waste\n   time mutating wrapper glue code\n3. **Reference implementations** — useful for cross-verification but\n   may not be the best mutation targets\n\nFor each implementation, note:\n- Language and mutation testing framework\n- Whether it's pure code or FFI wrappers\n- Existing test suite size and coverage\n- Which API surface the test vectors will exercise\n\n### Implementation Type Classification\n\n| Type | Mutation Value | Example |\n|------|---------------|---------|\n| Pure implementation | High | zkcrypto/bls12_381 (Rust), gnark-crypto (Go) |\n| FFI bindings to C/asm | Low at binding layer | blst Rust crate |\n| C/C++ implementation | High (use Mull) | blst C library |\n| Generated code | Medium (mutations may be equivalent) | gnark-crypto generated field arithmetic |\n\n**Key insight:** If an implementation delegates to another language\nvia FFI, you must mutate the *underlying* implementation, not the\nbindings. For C/C++ underneath Rust/Go/Python, use Mull or similar.\n\n---\n\n## Phase 2: Harness\n\nFor each implementation, create a test harness that:\n\n1. Reads test vectors from JSON files (Wycheproof format recommended)\n2. Exercises the implementation's API for each vector\n3. Asserts **both acceptance and rejection**:\n   - Valid vectors: deserialization succeeds, output matches expected\n   - Invalid vectors: deserialization fails or verification rejects\n4. Adds **roundtrip assertions** for valid deserialization vectors:\n   `serialize(deserialize(bytes)) == bytes`\n5. Reports pass/fail per vector with test IDs\n\n**Critical:** A harness that only checks valid vectors will miss all\npermissive mutations (e.g., `&amp;` → `|` in validation). See\n[references/lessons-learned.md](references/lessons-learned.md) §7.\n\nThe harness must be runnable by the mutation testing framework.\nFor most frameworks this means:\n- **Go:** A `_test.go` file in the same package as the implementation\n- **Rust:** An integration test in `tests/` or inline `#[test]` functions\n- **Python:** A pytest test file\n- **C/C++:** A test binary linked against the implementation\n\n### Harness Placement\n\nThe harness must live *inside the implementation's package* so the\nmutation framework can see it. This usually means:\n\n```bash\n# Go: add test file to the package being mutated\ncp wycheproof_test.go /path/to/impl/package/\n\n# Rust: add integration test\ncp wycheproof.rs /path/to/crate/tests/\n\n# Python: add test to the test directory\ncp test_wycheproof.py /path/to/package/tests/\n</code></pre>\n<h3>Handling Existing Vectors</h3>\n<p>If the implementation already has test vectors:</p>\n<ol>\n<li>Run mutation testing with ONLY the existing vectors (baseline)</li>\n<li>Run mutation testing with ONLY your new vectors</li>\n<li>Run mutation testing with BOTH combined</li>\n<li>The delta between (1) and (3) shows the new vectors' value</li>\n</ol>\n<hr>\n<h2>Phase 3: Baseline</h2>\n<p>Run mutation testing with existing test vectors only.</p>\n<h3>Framework Selection</h3>\n<p>See <a href=\"references/mutation-frameworks.md\">references/mutation-frameworks.md</a>\nfor language-specific setup.</p>\n<table>\n<thead>\n<tr>\n<th>Language</th>\n<th>Framework</th>\n<th>Command</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Go</td>\n<td>gremlins</td>\n<td><code>gremlins unleash ./path/to/package</code></td>\n</tr>\n<tr>\n<td>Rust</td>\n<td>cargo-mutants</td>\n<td><code>cargo mutants -j N --timeout T</code></td>\n</tr>\n<tr>\n<td>Python</td>\n<td>mutmut</td>\n<td><code>mutmut run --paths-to-mutate src/</code></td>\n</tr>\n<tr>\n<td>C/C++</td>\n<td>Mull</td>\n<td><code>mull-runner -test-framework=GoogleTest binary</code></td>\n</tr>\n</tbody>\n</table>\n<h3>Parallelism</h3>\n<p>Always use parallel execution for large codebases:</p>\n<ul>\n<li><code>cargo mutants -j 8</code> (Rust, 8 parallel workers)</li>\n<li><code>gremlins unleash --timeout-coefficient 3</code> (Go, increase timeouts)</li>\n<li><code>mutmut run --runner \"pytest -x -q\"</code> (Python, fail-fast)</li>\n</ul>\n<h3>Recording Baseline Results</h3>\n<p>Capture these metrics per implementation:</p>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Total mutants</td>\n<td>Number of mutations generated</td>\n</tr>\n<tr>\n<td>Killed</td>\n<td>Mutants caught by tests</td>\n</tr>\n<tr>\n<td>Survived/Lived</td>\n<td>Mutants NOT caught (these are the targets)</td>\n</tr>\n<tr>\n<td>Not covered</td>\n<td>Code paths no test reaches at all</td>\n</tr>\n<tr>\n<td>Timed out</td>\n<td>Ambiguous — resolve before comparing</td>\n</tr>\n<tr>\n<td>Efficacy %</td>\n<td>Killed / (Killed + Survived)</td>\n</tr>\n<tr>\n<td>Coverage %</td>\n<td>(Total - Not covered) / Total</td>\n</tr>\n</tbody>\n</table>\n<p>Save the full mutation log for Phase 4 analysis.</p>\n<hr>\n<h2>Phase 4: Escape Analysis (Graph-Informed Triage)</h2>\n<p>Classify each escaped (survived + not covered) mutant using the\nTrailmark call graph for reachability and blast radius analysis.</p>\n<p><strong>This phase MUST use the genotoxic skill's triage methodology.</strong>\nThe call graph transforms mutation results from a flat list of\nsurvived mutants into an actionable, prioritized set of vector\ntargets.</p>\n<h3>Step 1: Build the Call Graph</h3>\n<p>Build a Trailmark code graph for each implementation before\ntriaging mutations:</p>\n<pre><code># Go\nuv run trailmark analyze --language go --summary {targetDir}\n\n# Rust\nuv run trailmark analyze --language rust --summary {targetDir}\n</code></pre>\n<p>The graph provides:</p>\n<ul>\n<li><strong>Caller chains</strong> — trace from public API entry points to\nmutated functions to determine reachability</li>\n<li><strong>Cyclomatic complexity</strong> — prioritize high-CC functions</li>\n<li><strong>Blast radius</strong> — functions with many callers have wider\nimpact if their mutations survive</li>\n</ul>\n<h3>Step 2: Filter to Relevant Code</h3>\n<p>Mutation frameworks test the entire package. Filter results to\nonly the files/functions that test vectors should exercise:</p>\n<pre><code># Go (gremlins)\ngrep -E \"(LIVED|NOT COVERED)\" baseline.log \\\n  | grep -E \" at (relevant|files)\" \\\n  | sort\n\n# Rust (cargo-mutants)\ncat mutants.out/missed.txt | grep \"src/relevant\"\n</code></pre>\n<h3>Step 3: Graph-Informed Classification</h3>\n<p>For each escaped mutant, map it to its containing function in the\ncall graph and apply the genotoxic triage criteria:</p>\n<table>\n<thead>\n<tr>\n<th>Graph Signal</th>\n<th>Classification</th>\n<th>Action</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No callers in graph</td>\n<td><strong>False Positive</strong></td>\n<td>Dead code, skip</td>\n</tr>\n<tr>\n<td>Only test callers</td>\n<td><strong>False Positive</strong></td>\n<td>Test infrastructure</td>\n</tr>\n<tr>\n<td>Logging/display/formatting</td>\n<td><strong>False Positive</strong></td>\n<td>Cosmetic</td>\n</tr>\n<tr>\n<td>Cross-package callers but NOT COVERED</td>\n<td><strong>Cross-Package Gap</strong></td>\n<td>See below</td>\n</tr>\n<tr>\n<td>Reachable from public API, low CC</td>\n<td><strong>Missing Vector</strong></td>\n<td>Design targeted vector</td>\n</tr>\n<tr>\n<td>Reachable from public API, high CC (&gt;10)</td>\n<td><strong>Fuzzing Target</strong></td>\n<td>Both vector + fuzz harness</td>\n</tr>\n<tr>\n<td>Validation/error-handling path</td>\n<td><strong>Negative Vector</strong></td>\n<td>Craft invalid input that triggers path</td>\n</tr>\n<tr>\n<td>Optimization path (GLV, SIMD, batch)</td>\n<td><strong>Edge-Case Vector</strong></td>\n<td>Input that triggers optimization threshold</td>\n</tr>\n<tr>\n<td><code>\\|</code>→<code>^</code> after left shift (e.g. <code>(t&lt;&lt;1) \\| carry</code>)</td>\n<td><strong>Equivalent Mutant</strong></td>\n<td>Skip — bit 0 always 0, OR=XOR</td>\n</tr>\n<tr>\n<td>ct_eq <code>&amp;</code>→<code>\\|</code> on Montgomery limbs</td>\n<td><strong>API-Unreachable</strong></td>\n<td>Needs library-internal tests, not vectors</td>\n</tr>\n<tr>\n<td>Equivalent mutation (behavior unchanged)</td>\n<td><strong>False Positive</strong></td>\n<td>Skip</td>\n</tr>\n</tbody>\n</table>\n<h3>Step 4: Identify Cross-Package Test Gaps</h3>\n<p><strong>Critical pitfall:</strong> Mutation frameworks often only run tests\nwithin the same package as the mutation. For Go (gremlins) and\nRust (cargo-mutants), this means:</p>\n<ul>\n<li>A mutation in <code>hash_to_curve/g2.go</code> only runs tests in the\n<code>hash_to_curve</code> package, NOT tests in the parent <code>bls12381</code>\npackage that imports it</li>\n<li>Functions that are fully exercised by cross-package tests\nwill appear as NOT COVERED — these are <strong>false positives</strong></li>\n<li>To confirm: check if the mutated function is called from a\ntest in a <em>different</em> package that wouldn't be run</li>\n</ul>\n<p>To resolve cross-package gaps:</p>\n<ol>\n<li>Add a thin test in the sub-package that calls through the\nsame code path as the cross-package test</li>\n<li>Or run gremlins with <code>--test-pkg ./...</code> (if supported)</li>\n<li>Or document as a framework limitation in the report</li>\n</ol>\n<h3>Step 5: Prioritize by Security Impact</h3>\n<p>Using the call graph, rank surviving mutants by impact:</p>\n<table>\n<thead>\n<tr>\n<th>Priority</th>\n<th>Criteria</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>P0 — Critical</strong></td>\n<td>Mutant weakens validation/equality/authentication</td>\n<td><code>ct_eq</code>: <code>&amp;</code> → <code>\\|</code> makes equality permissive</td>\n</tr>\n<tr>\n<td><strong>P1 — High</strong></td>\n<td>Mutant in deserialization flag parsing</td>\n<td><code>from_compressed</code>: <code>&amp;</code> → <code>\\|</code> accepts invalid flags</td>\n</tr>\n<tr>\n<td><strong>P2 — Medium</strong></td>\n<td>Mutant in field arithmetic internals</td>\n<td><code>Fp::square</code>: <code>\\|</code> → <code>^</code> corrupts computation</td>\n</tr>\n<tr>\n<td><strong>P3 — Low</strong></td>\n<td>Mutant in optimization path</td>\n<td><code>phi</code> endomorphism: only affects performance path</td>\n</tr>\n<tr>\n<td><strong>Skip</strong></td>\n<td>Formatting, display, equivalent mutation</td>\n<td><code>Debug::fmt</code> return value replacement</td>\n</tr>\n</tbody>\n</table>\n<h3>Step 6: Group by Vector Strategy</h3>\n<p>Group escaped mutants by the code path they represent and the\ntype of test vector needed:</p>\n<pre><code>Deserialization flag validation (P1):\n  - g1.rs:339,363-365,384 — from_compressed_unchecked flags\n  → Need: valid-point-wrong-flag vectors\n\nField arithmetic (P2):\n  - fp.rs:371-376,406,635-643 — subtract_p, neg, square\n  → Need: field arithmetic KATs with edge-case values\n\nOptimization thresholds (P3):\n  - g1.go:68, g2.go:75 — GLV vs windowed multiplication\n  → Need: scalar multiplication with large scalars\n\nCross-package (framework limitation):\n  - hash_to_curve/g2.go:242-278 — isogeny, sgn0\n  → Document as false positive or add sub-package test\n</code></pre>\n<p>Each group becomes a target for new test vectors in Phase 5.</p>\n<hr>\n<h2>Phase 5: Vector Generation</h2>\n<p>For each escaped code path group, design test vectors that\nforce execution through that path.</p>\n<h3>Vector Design Patterns</h3>\n<table>\n<thead>\n<tr>\n<th>Code Path Type</th>\n<th>Vector Strategy</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Point deserialization</td>\n<td>Malformed points: wrong length, invalid field elements, off-curve, wrong subgroup, identity point</td>\n</tr>\n<tr>\n<td>Signature verification</td>\n<td>Valid sig + all single-bit corruptions of sig, pk, msg</td>\n</tr>\n<tr>\n<td>Hash-to-curve</td>\n<td>Known answer tests (KATs) with edge-case inputs: empty, single byte, max length</td>\n</tr>\n<tr>\n<td>Aggregate operations</td>\n<td>1 signer, many signers, duplicate signers, mixed valid/invalid</td>\n</tr>\n<tr>\n<td>Error handling</td>\n<td>Every error path should have a vector that triggers it</td>\n</tr>\n<tr>\n<td>Arithmetic edge cases</td>\n<td>Zero, one, field modulus - 1, points at infinity</td>\n</tr>\n<tr>\n<td>Serialization flags</td>\n<td>Every valid flag combination + every invalid flag combination</td>\n</tr>\n<tr>\n<td>Roundtrip integrity</td>\n<td>For every valid deser vector, assert <code>serialize(deserialize(b)) == b</code></td>\n</tr>\n<tr>\n<td>Carry/reduction faults</td>\n<td>Reimplement at reduced limb widths, inject faults, extract distinguishing inputs</td>\n</tr>\n</tbody>\n</table>\n<h3>Single-Fault Negative Vectors</h3>\n<p>Each negative vector should have <strong>exactly one defect</strong> with\neverything else valid — this isolates which validation check is\nbeing tested. See <a href=\"references/vector-patterns.md\">references/vector-patterns.md</a>\nfor per-flag construction examples.</p>\n<h3>Fault Simulation (Limb-Width Reimplementation)</h3>\n<p>When mutation testing only applies local operator swaps, deeper\narchitectural bugs (carry propagation, reduction overflow) go\nuntested. To close this gap, reimplement the target algorithm\nat reduced limb widths (8, 16, 25, 32 bits) and deliberately\ninject faults — then generate vectors that catch them.</p>\n<p>See <a href=\"references/fault-simulation.md\">references/fault-simulation.md</a>\nfor the full methodology: limb-width selection, fault injection\ncatalog, vector extraction, and validation workflow.</p>\n<h3>Cross-Implementation Verification</h3>\n<p>Every new test vector MUST be verified against at least two\nindependent implementations before being added to the suite:</p>\n<ol>\n<li>Generate the vector using implementation A</li>\n<li>Verify with implementation B (different codebase, ideally different language)</li>\n<li>If B disagrees, investigate — one implementation has a bug</li>\n</ol>\n<h3>Vector Format</h3>\n<p>Use Wycheproof JSON format (<code>algorithm</code>, <code>testGroups[].tests[]</code>\nwith <code>tcId</code>, <code>comment</code>, <code>result</code>, <code>flags</code>). See\n<a href=\"references/vector-patterns.md\">references/vector-patterns.md</a>\nfor the full schema.</p>\n<p><strong>Wycheproof contributions:</strong> Use Wycheproof's <code>vectorgen</code> tool rather\nthan formatting vector files directly. Supply the generated changes as an\nenvelope. The <code>vectorgen</code> tool can add, update, or replace vectors while\nhandling <code>tcId</code> assignment, test counts, canonical formatting, and schema\nvalidation. Go-based generators can avoid the <code>vectorgen</code> CLI tool and instead\ncall the programmatic <code>github.com/c2sp/wycheproof/vectorgen</code> API.</p>\n<p>See <a href=\"references/lessons-learned.md\">references/lessons-learned.md</a>\n§14 and the upstream\n<a href=\"https://github.com/C2SP/wycheproof/blob/main/doc/vectorgen.md\">vectorgen guide</a>\nfor the current workflow and commands.</p>\n<hr>\n<h2>Phase 6: Validation</h2>\n<p>Re-run mutation testing with the new test vectors included.</p>\n<p><strong>Tip:</strong> Use per-file mutation testing for fast iteration during\nvector development (see <a href=\"references/lessons-learned.md\">references/lessons-learned.md</a> §12).\nOnly run full-crate tests for the final comparison.</p>\n<h3>Before/After Comparison</h3>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Baseline</th>\n<th>With New Vectors</th>\n<th>Delta</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Killed</td>\n<td>X</td>\n<td>Y</td>\n<td>Y - X</td>\n</tr>\n<tr>\n<td>Survived</td>\n<td>A</td>\n<td>B</td>\n<td>A - B (should decrease)</td>\n</tr>\n<tr>\n<td>Not Covered</td>\n<td>C</td>\n<td>D</td>\n<td>C - D (should decrease)</td>\n</tr>\n<tr>\n<td>Efficacy %</td>\n<td>E%</td>\n<td>F%</td>\n<td>F - E</td>\n</tr>\n</tbody>\n</table>\n<h3>Success Criteria</h3>\n<p>Vectors have both <strong>retroactive</strong> value (killing mutants in\nexisting code) and <strong>proactive</strong> value (catching bugs in future\nimplementations). Generate both kinds — boundary-condition vectors\nmay not improve kill rates in mature libraries but will catch bugs\nin new implementations. See\n<a href=\"references/lessons-learned.md\">references/lessons-learned.md</a> §13.</p>\n<p><strong>Retroactive (measurable):</strong> previously survived/uncovered mutants\nbecome killed, no regressions.</p>\n<p><strong>If kill rates don't change:</strong> the implementation's own tests\nlikely already cover those paths. The vectors still add\ncross-implementation verification value. Document which case\napplies.</p>\n<hr>\n<h2>Output Format</h2>\n<p>Write <code>VECTOR_FORGE_REPORT.md</code> covering: target algorithm,\nimplementations tested, baseline results, escape analysis,\nnew vectors generated, after results, before/after delta, and\nconclusions. See\n<a href=\"references/report-template.md\">references/report-template.md</a>\nfor the full template.</p>\n<hr>\n<h2>Quality Checklist</h2>\n<p>Before delivering:</p>\n<ul>\n<li><input disabled=\"disabled\" type=\"checkbox\"> At least one pure implementation mutation-tested (not just FFI wrappers)</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Baseline run completed with existing vectors</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Trailmark call graph built for each implementation</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> All escaped mutants triaged using graph-informed classification</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Cross-package false positives identified and documented</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Security-critical mutations (ct_eq, validation, auth) prioritized as P0/P1</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Fault simulation and mutation-derived vectors cross-verified against 2+ implementations</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> After run completed with new vectors included</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Before/after delta computed and explained</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Report written to <code>VECTOR_FORGE_REPORT.md</code></li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> New test vectors saved in standard format (Wycheproof JSON)</li>\n</ul>\n<hr>\n<h2>Integration</h2>\n<table>\n<thead>\n<tr>\n<th>Skill</th>\n<th>Relationship</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>genotoxic</strong> (required for Phase 4)</td>\n<td>Provides graph-informed triage — call graph cuts actionable mutants by 30-50%</td>\n</tr>\n<tr>\n<td><strong>mutation-testing</strong> (mewt/muton)</td>\n<td>Use for Solidity; Vector Forge is language-agnostic</td>\n</tr>\n<tr>\n<td><strong>property-based-testing</strong></td>\n<td>Better than hand-crafted vectors for bitwise mutations in field arithmetic</td>\n</tr>\n<tr>\n<td><strong>testing-handbook-skills</strong> (fuzzing)</td>\n<td>Functions with CC &gt; 10 and surviving mutants need both vectors and fuzz harnesses</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Supporting Documentation</h2>\n<ul>\n<li><strong><a href=\"references/mutation-frameworks.md\">references/mutation-frameworks.md</a></strong> -\nLanguage-specific mutation testing framework setup</li>\n<li><strong><a href=\"references/vector-patterns.md\">references/vector-patterns.md</a></strong> -\nCommon test vector patterns for cryptographic primitives</li>\n<li><strong><a href=\"references/fault-simulation.md\">references/fault-simulation.md</a></strong> -\nLimb-width reimplementation for carry, reduction, and overflow faults</li>\n<li><strong><a href=\"references/report-template.md\">references/report-template.md</a></strong> -\nFull markdown template for the Vector Forge report</li>\n<li><strong><a href=\"references/lessons-learned.md\">references/lessons-learned.md</a></strong> -\nBLS12-381 case study: FFI kill rates, timeout masking, cross-package\nfalse positives, bitwise mutation gaps, and security-critical priorities</li>\n</ul>\n","files":[{"path":"agents/openai.yaml","sizeBytes":241,"isText":true},{"path":"assets/trail-of-bits-mark.svg","sizeBytes":3084,"isText":false},{"path":"references/fault-simulation.md","sizeBytes":5831,"isText":true},{"path":"references/lessons-learned.md","sizeBytes":7785,"isText":true},{"path":"references/mutation-frameworks.md","sizeBytes":20281,"isText":true},{"path":"references/report-template.md","sizeBytes":912,"isText":true},{"path":"references/vector-patterns.md","sizeBytes":9867,"isText":true},{"path":"SKILL.md","sizeBytes":19383,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-17T16:00:40.180188Z","sha256":"6EDC0C60AAE4AB1DF34AC8DEA074D58F26BF4FD71B4F07A82C65225A65EEED9B","sizeBytes":27923},"review":null,"source":{"repositoryUrl":"https://github.com/trailofbits/skills","path":"plugins/trailmark/skills/vector-forge","license":"CC-BY-SA-4.0","commit":"0cc1c73a5e96749ab32d7ea5e14892fafa6972ae","subtreeSha":"170FB7D8D7DC84203A2B0C11BEF40A7E980A65FB538AFAEC8D20D735AF4FF444","lastSyncedAt":"2026-09-25T07:36:46.789003Z"},"reviewedAt":"2026-09-17T16:06:34.635875Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/trailofbits/skills/tree/main/plugins/trailmark/skills/vector-forge"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install trailofbits-skills@llmmart"},{"target":"git","command":"git clone https://github.com/trailofbits/skills.git"}]}