LLM Mart Basic
@llm-mart · Joined Jun 2026
缺陷修复方法论唯一入口:口喷 bug 定位直修;implement-test-review 循环内的反馈修复也走本 skill(循环内分支,原 impl worker 执行)。触发:修 bug / 有个 bug / 报错了 / 行为不对 / fix / /eo-fix。 NOT FOR: 明确的业务变更(走 /eo-change)。
在 /clear 之前生成最小可恢复快照到 tmp/eo/handoff/<topic>.md,让下一个会话载入这一个文件就能从当前节点继续。优先记录任务状态、关键口径、下一步分叉,主动丢弃探索过程。触发:handoff / 存一下进度 / 我要 clear / /eo-handoff。 NOT FOR: 机械压缩对话流(用内置 /compact)。
按 change.md 的 TODO 分批落地代码,批末自验对应 AC。触发:实现 / 写代码 / implement / /eo-implement。 NOT FOR: bug 与反馈修复(口喷 bug 与 test/review/acceptance 循环内反馈一律走 /eo-fix);变更起草(走 /eo-change)。
eo 流程总控:按用户意图圈一段(入口节点 → 出口节点 → 收敛标准),把 eo-change / eo-implement / eo-archive 及可选闸门(eo-change-review / eo-test / eo-review)派发到可插拔执行基底上推进至收敛,窗口化汇报进度。触发:eo-loop / 串起来跑 / 循环推进到收敛 / 总控调度 / /eo-loop。 NOT FOR: 单点动作(直接调对应 eo-* skill);派出去不再监督的完全交接(orca-cli full handoff);bug 口喷(/eo-fix)。
eo-skills 在当前仓库的总入口:生成 .eo-project.json、初始化项目管理侧(roadmap)和代码侧最小骨架(eo-doc/),以及 agent 配置注入。触发:启动项目 / 初始化项目 / 新建项目 / /eo-project-init。
项目记忆的统一写入口:经验教训(lessons/)与关键决策(decisions/),带 INDEX 与检索锚点,供 eo-change / eo-fix / eo-recall 消费。通过 .eo-project.json 定位。触发(仅用户明确要求记录时):把这个坑记下来 / 记条经验 / lesson learned / 把这个决策记下来 / 记录决策 / reindex lessons / reindex decisions / /eo-project-record。NOT FOR: 对话提到踩坑或决策但用户未要求记录;待办类(走 /eo-bac
只读的回忆与解释入口,分层作答带出处。触发:这个功能当时怎么设计的 / 这段逻辑怎么实现的 / 当初为什么这么定 / 帮我回忆 / recall / /eo-recall。 NOT FOR: 修 bug(/eo-fix)、发起变更(/eo-change)、维护文档(/eo-doc-manager)。
按需代码审查:风险信号命中或用户点名时,对已实施代码做 AC 逐条核对 + 代码质量审查,产简版 review.md(P0/P1/P2)。触发:review / 代码审查 / 再找双眼睛看看 / /eo-review。 NOT FOR: change 方案审查(/eo-change-review,代码还没写时用);默认主路(无信号时不强制)。
按需独立测试视角:风险信号命中或用户点名时,对已实现 change 做测试审计、补缺与重验证,产简版 test.md。严禁修改业务代码。触发:找双新眼睛跑测试 / 独立验证 / 给这个 change 补测试 / /eo-test。 NOT FOR: 与 change 无关的日常跑测试或补单测(直接做即可,不产报告);默认主路的自验(归 eo-implement)。
接口级测试时使用——从 OpenAPI/Swagger 文档或用例 Schema 中可自动化的接口用例出发,覆盖参数、边界、鉴权、幂等、并发、错误响应与数据一致性,产出可执行的 API 测试脚本与运行结果;含接口压测承接(k6,类型矩阵轴 1 执行层)。不用于:Web UI 流程(automated-e2e-testing)、手动用例编写(test-case-writing)。
将手动测试用例转为 Playwright E2E 测试并执行时使用;含写自动化前的业务熟悉踩点、Page Object/Helper 编写、执行中的 Bug 证据收集与报告条目记录。不用于:纯 API 接口测试(api-testing)、以理解系统为目的的独立探索会话(exploratory-testing)、已确认 Bug 的根因分析(bug-analysis)。
对已确认的 Bug 做根因定位、影响分析、回归建议时使用——复现 → 读代码定位根因 → 影响五面分析 → 回归建议,条目(根因/影响/Severity 依据/修复建议)追加进测试报告。不用于:仅收集 Bug 证据(automated-e2e-testing / api-testing)、疑似未定性缺陷(test-case-writing 的 Cx 记录)。
qa-skills 共享知识库,安装依赖单元(非触发 skill):承载被其余 10 个 skill 以相对路径引用的方法、规则、模板与脚本(可执行性标准、证据分级、风险模型、类型决策矩阵等)。仅在 qa-skills 系列 skill 工作流中被引用读取;任何具体测试任务都不要独立触发本 skill,独立使用无意义。通过 npx skills 等安装器单独安装其他 qa-skills skill 时,必须同时安装本 skill,否则引用路径断裂。
需求不完整、系统陌生、文档不足时,发起以理解系统/发现风险为目的的独立探索式测试会话时使用——charter 驱动(目标 → 探索 → 记录),产出探索笔记(系统理解/风险清单/测试想法)作为需求建模输入或独立交付。不用于:为写自动化踩点的小规模探索(automated-e2e-testing 工作流零)、按既有用例执行(执行类 skill)。
端到端测试的唯一入口:用户说"帮我测试这个需求/功能"、"把这个功能完整测一遍"时,编排需求理解→测试策略(风险与类型决策)→用例→审查→执行→Bug 分析→回归→报告的完整流水线,产出落盘、可断点续跑。只要单阶段产出(如"帮我审一下这份用例")→ 直接用对应阶段 skill,不用本 skill。
代码变更(diff/Bug 修复/需求变更)后判断应回归哪些测试时使用——沿"改动文件 → 改动函数 → 受影响功能 → 受影响用例"分析链,基于用例 Schema 的追溯映射产出分级回归清单。不用于:用例文件本身的增量修改(test-case-writing)、长期回归策略(test-strategy)。
系统性建模某个需求/系统时使用——从 PRD、设计/API 文档、Bug、Issue、代码中提炼目标、范围、角色、规则、异常、依赖与不明确项,产出结构化需求模型(含澄清记录与用户裁决)。不用于:已有需求模型直接写用例(test-case-writing)、"怎么测"的策略决策(test-strategy)、端到端流水线(qa)。
审查已有测试用例(存量资产、他人编写、AI 产出)的覆盖与可执行性时使用——先建可测点基准,再独立评估覆盖、可执行性与正确性,直接修订用例文件并留审查记录。不用于:从零写用例(test-case-writing)、写时自审(其阶段四)、端到端流水线(qa)。
从需求文档、API 文档、Bug 报告或代码仓库产出可执行的手动测试用例(markmap)时使用——代码优先:索取仓库、先代码审查找潜在 bug 再写用例,并抽取机器可读 Schema。不用于:需求建模(requirement-analysis)、测试策略(test-strategy)、独立审查(test-case-review)、自动化脚本(automated-e2e-testing / api-testing)。
回答"这个功能应该怎么测"时使用——风险评级挂证据(Risk Map),翻译成功能域+类型域两域范围与深度:类型域十轴全轴必答(脚本扫描信号+预填修订),include挂信号、exclude挂理由、full有预算上限。不用于:已有策略直接写用例(test-case-writing)、需求建模(requirement-analysis)、端到端流水线(qa)。
/autopilot
autopilot
Run autonomous hunt loop on a target — scope check → recon → rank surface → hunt → validate → report with configurable checkpoints. Usage: /autopilot target.com [--paranoid|--normal|--yolo]
/chain
chain
Build an exploit chain — given bug A, finds B and C to combine for higher severity and payout. Knows common chain patterns: IDOR→ATO, SSRF→cloud metadata, XSS→ATO, open redirect→OAuth theft, S3→bundle→secret→OAuth. Usage: /chain
/hunt
hunt
Active vulnerability hunting. Two-track dispatcher — asks Red Team vs WAPT, hands off to hunt-dispatch skill and sibling commands. Usage: /hunt target.com | /hunt *.target.com | /hunt targets.txt [--vuln-class X] [--source-code P] [--chrome]
/intel
intel
On-demand intelligence fetch for a target — CVEs, disclosed reports, new features. Pulls NVD/GitHub-Advisory CVEs + bundled disclosed reports + hunt memory context. Usage: /intel target.com
/memory-gc
memory-gc
Inspect or rotate the autopilot ledger JSONL files (findings.jsonl, negatives.jsonl). Caps file size and keeps N rotated backups so memory does not grow unbounded.
/pickup
pickup
Pick up a previous hunt on a target — shows hunt history and untested surface from the autopilot ledger. Usage: /pickup target.com
/recon
recon
Run full recon pipeline on a target — subdomain enum (Chaos API + subfinder), live host discovery (dnsx + httpx), URL crawl (katana + waybackurls + gau), gf pattern classification, nuclei scan. Outputs to recon/<target>/ directory. Usage: /recon target.com
/remember
remember
Optional manual note on a target or the last confirmed finding. Capture is automatic during autopilot; this is for extra context. Usage: /remember
/report
report
Write a submission-ready bug bounty report. Generates H1/Bugcrowd/Intigriti/Immunefi format with CVSS 3.1 score, proof of concept, impact statement, and remediation. Run /validate first. Usage: /report
/scope
scope
Mandatory pre-flight scope check — verify an asset is in scope BEFORE any HTTP touch. Deterministic (deny-wins, default-deny) via engine/scope.py against the engagement's scope.md. Blocks out-of-scope testing. Usage: /scope <asset> [<asset> ...]
/surface
surface
Show ranked attack surface for a target from its recon manifest + hunt memory. Deterministic backing is `cbh surface <target>` (reads recon/<target>/manifest.json); LLM layer adds ledger signal. Usage: /surface target.com
/token-scan
token-scan
Meme coin and token security scan — checks for rug pull vectors (hidden mint, honeypot, fee manipulation, LP lock bypass, authority retention, bonding curve exploits, fake renounce, sandwich amplification). Manual 8-class grep audit (with an optional automated scanner if present). Usage: /token-scan <contract_path_or_dir> [--chain solana]
/triage
triage
Quick 7-Question Gate triage on a finding before writing a report. Kills N/A submissions before they happen. Faster than /validate — for quick go/no-go decisions. Usage: /triage
/validate
validate
Validate a finding — runs 7-Question Gate + 4-gate checklist. Kills weak findings before report writing. Prevents N/A submissions that hurt validity ratio. Usage: /validate
/web3-audit
web3-audit
Smart contract security audit — runs through 10 bug class checklist (accounting desync, access control, incomplete path, off-by-one, oracle errors, ERC4626, reentrancy, flash loan, signature replay, proxy/upgrade). Applies pre-dive kill signals first. Generates Foundry PoC template for confirmed findings. Usage: /web3-audit <contract.sol>
/README
README
Crabbox is a single CLI (`crabbox`). Commands are top-level, not nested under a
/actions
Actions
`crabbox actions` prepares a leased box from your repository's own GitHub
/adapter
Adapter
See [Runtime adapter stack](../features/runtime-adapter-stack.md) for the
/admin
Admin
`crabbox admin` groups trusted operator controls for coordinator-backed leases and the cloud resources behind them. Use it to inspect every lease the broker tracks, reconcile expired leases against live cloud state, force-release or delete a backing server, print provider IAM pol
/artifacts
Artifacts
`crabbox artifacts` turns a desktop lease into durable QA evidence: it collects
Make any song you can imagine
39 views 0 likesLeading AI-powered video generation platform that specializes in creating hyper-realistic talking avatars
37 views 0 likesHermes Agent is an open-source, self-improving autonomous AI agent developed by Nous Research
36 views 0 likesKilo Code is a popular, open-source AI coding agent and "agentic engineering" platform designed to help developers build, refactor, and debug software faster
34 views 0 likesGeneral-purpose agent in one static Go binary. ReAct loop, ACP server for IDEs, OpenAI-compatible REST API with embedded web UI, Telegram gateway, cron schedule…
20 views 0 likesAutonomous agent framework with structured memory, safety hooks, and loop management. Built by the agent that runs on it.
20 views 0 likesTSP自托管、零运维的 A 股「选股 + 监控 + 回测」量化工作台 | 基于 TickFlow 数据源 | LLM能力驱使策略定制+个股分析+复盘 | 自由接入第三方数据源与个性化扩展数据 | 个人开源 ,非TickFlow官方项目
15 views 0 likesCurated, verified Agent Skills powered by ModelStudio.
18 views 0 likesRun Claude Code, Codex, Antigravity, Cursor Agent and OpenCode as one runtime — persistent sessions, multi-agent councils, an OpenAI-compatible endpoint, an MCP…
17 views 0 likespi had nothing (nothing), so I made something (something) — sorry mariozechner-senpai, I went ahead and lovingly soiled your pure pi for you. opinionated fork o…
14 views 0 likesA persistent workspace for development work that self-improves and continues beyond one session.
35 views 0 likesOpen-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.
20 views 0 likes📚 A zero-dependency, git-backed micro-lesson library for AI Agents to asynchronously share and search verified debugging experience. Python stdlib only. | http…
28 views 0 likesDeterministic, local-first memory and guardrails for AI coding agents with no LLM in the hot path.
31 views 0 likesDeterministic spec-orchestration for local LLMs in the pi coding agent — drives prompts through refine→research→grill→compose→critique, with bundled web/docs/fe…
20 views 0 likesNative Safari browser automation for AI agents. 97 tools via AppleScript — zero overhead, keeps logins, runs silently in background. Drop-in alternative to Chro…
34 views 0 likesAgent OS: keep specialist agents in a hub, spin up a temporary orchestrator per task. Local-first, works with any model.
15 views 0 likesGit for agent memory. Branches, diffs, PRs, and rollback for what your agents know.
35 views 0 likesMulti-Provider AI Gateway - No personal logs by design. Model autodiscovery, Failover groups, High availability, Android companion app, and more - "Because we h…
16 views 0 likesProduction-grade MCP server for MikroTik RouterOS with secure AI-native network automation.
31 views 0 likes