LLM Mart Basic
@llm-mart · Joined Jun 2026
缺陷修复方法论唯一入口:口喷 bug 定位直修;implement-test-review 循环内的反馈修复也走本 skill(循环内分支,原 impl worker 执行)。触发:修 bug / 有个 bug / 报错了 / 行为不对 / fix / /eo-fix。 NOT FOR: 明确的业务变更(走 /eo-change)。
在 /clear 之前生成最小可恢复快照到 tmp/eo/handoff/<topic>.md,让下一个会话载入这一个文件就能从当前节点继续。优先记录任务状态、关键口径、下一步分叉,主动丢弃探索过程。触发:handoff / 存一下进度 / 我要 clear / /eo-handoff。 NOT FOR: 机械压缩对话流(用内置 /compact)。
按 change.md 的 TODO 分批落地代码,批末自验对应 AC。触发:实现 / 写代码 / implement / /eo-implement。 NOT FOR: bug 与反馈修复(口喷 bug 与 test/review/acceptance 循环内反馈一律走 /eo-fix);变更起草(走 /eo-change)。
eo 流程总控:按用户意图圈一段(入口节点 → 出口节点 → 收敛标准),把 eo-change / eo-implement / eo-archive 及可选闸门(eo-change-review / eo-test / eo-review)派发到可插拔执行基底上推进至收敛,窗口化汇报进度。触发:eo-loop / 串起来跑 / 循环推进到收敛 / 总控调度 / /eo-loop。 NOT FOR: 单点动作(直接调对应 eo-* skill);派出去不再监督的完全交接(orca-cli full handoff);bug 口喷(/eo-fix)。
eo-skills 在当前仓库的总入口:生成 .eo-project.json、初始化项目管理侧(roadmap)和代码侧最小骨架(eo-doc/),以及 agent 配置注入。触发:启动项目 / 初始化项目 / 新建项目 / /eo-project-init。
项目记忆的统一写入口:经验教训(lessons/)与关键决策(decisions/),带 INDEX 与检索锚点,供 eo-change / eo-fix / eo-recall 消费。通过 .eo-project.json 定位。触发(仅用户明确要求记录时):把这个坑记下来 / 记条经验 / lesson learned / 把这个决策记下来 / 记录决策 / reindex lessons / reindex decisions / /eo-project-record。NOT FOR: 对话提到踩坑或决策但用户未要求记录;待办类(走 /eo-bac
只读的回忆与解释入口,分层作答带出处。触发:这个功能当时怎么设计的 / 这段逻辑怎么实现的 / 当初为什么这么定 / 帮我回忆 / recall / /eo-recall。 NOT FOR: 修 bug(/eo-fix)、发起变更(/eo-change)、维护文档(/eo-doc-manager)。
按需代码审查:风险信号命中或用户点名时,对已实施代码做 AC 逐条核对 + 代码质量审查,产简版 review.md(P0/P1/P2)。触发:review / 代码审查 / 再找双眼睛看看 / /eo-review。 NOT FOR: change 方案审查(/eo-change-review,代码还没写时用);默认主路(无信号时不强制)。
按需独立测试视角:风险信号命中或用户点名时,对已实现 change 做测试审计、补缺与重验证,产简版 test.md。严禁修改业务代码。触发:找双新眼睛跑测试 / 独立验证 / 给这个 change 补测试 / /eo-test。 NOT FOR: 与 change 无关的日常跑测试或补单测(直接做即可,不产报告);默认主路的自验(归 eo-implement)。
接口级测试时使用——从 OpenAPI/Swagger 文档或用例 Schema 中可自动化的接口用例出发,覆盖参数、边界、鉴权、幂等、并发、错误响应与数据一致性,产出可执行的 API 测试脚本与运行结果;含接口压测承接(k6,类型矩阵轴 1 执行层)。不用于:Web UI 流程(automated-e2e-testing)、手动用例编写(test-case-writing)。
将手动测试用例转为 Playwright E2E 测试并执行时使用;含写自动化前的业务熟悉踩点、Page Object/Helper 编写、执行中的 Bug 证据收集与报告条目记录。不用于:纯 API 接口测试(api-testing)、以理解系统为目的的独立探索会话(exploratory-testing)、已确认 Bug 的根因分析(bug-analysis)。
对已确认的 Bug 做根因定位、影响分析、回归建议时使用——复现 → 读代码定位根因 → 影响五面分析 → 回归建议,条目(根因/影响/Severity 依据/修复建议)追加进测试报告。不用于:仅收集 Bug 证据(automated-e2e-testing / api-testing)、疑似未定性缺陷(test-case-writing 的 Cx 记录)。
qa-skills 共享知识库,安装依赖单元(非触发 skill):承载被其余 10 个 skill 以相对路径引用的方法、规则、模板与脚本(可执行性标准、证据分级、风险模型、类型决策矩阵等)。仅在 qa-skills 系列 skill 工作流中被引用读取;任何具体测试任务都不要独立触发本 skill,独立使用无意义。通过 npx skills 等安装器单独安装其他 qa-skills skill 时,必须同时安装本 skill,否则引用路径断裂。
需求不完整、系统陌生、文档不足时,发起以理解系统/发现风险为目的的独立探索式测试会话时使用——charter 驱动(目标 → 探索 → 记录),产出探索笔记(系统理解/风险清单/测试想法)作为需求建模输入或独立交付。不用于:为写自动化踩点的小规模探索(automated-e2e-testing 工作流零)、按既有用例执行(执行类 skill)。
端到端测试的唯一入口:用户说"帮我测试这个需求/功能"、"把这个功能完整测一遍"时,编排需求理解→测试策略(风险与类型决策)→用例→审查→执行→Bug 分析→回归→报告的完整流水线,产出落盘、可断点续跑。只要单阶段产出(如"帮我审一下这份用例")→ 直接用对应阶段 skill,不用本 skill。
代码变更(diff/Bug 修复/需求变更)后判断应回归哪些测试时使用——沿"改动文件 → 改动函数 → 受影响功能 → 受影响用例"分析链,基于用例 Schema 的追溯映射产出分级回归清单。不用于:用例文件本身的增量修改(test-case-writing)、长期回归策略(test-strategy)。
系统性建模某个需求/系统时使用——从 PRD、设计/API 文档、Bug、Issue、代码中提炼目标、范围、角色、规则、异常、依赖与不明确项,产出结构化需求模型(含澄清记录与用户裁决)。不用于:已有需求模型直接写用例(test-case-writing)、"怎么测"的策略决策(test-strategy)、端到端流水线(qa)。
审查已有测试用例(存量资产、他人编写、AI 产出)的覆盖与可执行性时使用——先建可测点基准,再独立评估覆盖、可执行性与正确性,直接修订用例文件并留审查记录。不用于:从零写用例(test-case-writing)、写时自审(其阶段四)、端到端流水线(qa)。
从需求文档、API 文档、Bug 报告或代码仓库产出可执行的手动测试用例(markmap)时使用——代码优先:索取仓库、先代码审查找潜在 bug 再写用例,并抽取机器可读 Schema。不用于:需求建模(requirement-analysis)、测试策略(test-strategy)、独立审查(test-case-review)、自动化脚本(automated-e2e-testing / api-testing)。
回答"这个功能应该怎么测"时使用——风险评级挂证据(Risk Map),翻译成功能域+类型域两域范围与深度:类型域十轴全轴必答(脚本扫描信号+预填修订),include挂信号、exclude挂理由、full有预算上限。不用于:已有策略直接写用例(test-case-writing)、需求建模(requirement-analysis)、端到端流水线(qa)。
/lineage-discovery
Lineage discovery
Discover testnet↔mainnet subnet lineage from repo configs and open a PR for review (pass --dry-run to report only)
/capture
capture
Triage raw inbox notes into reviewed repository destinations without deleting their sources.
/clean-ai-writing
clean-ai-writing
Audit and rewrite content to remove AI writing patterns
/content-shipped
content-shipped
Log a completed piece of content to content/log.md after the user confirms it was published.
/dream-apply
dream-apply
Validate a dream artifact, review each proposal, and apply only individually accepted changes.
/dream
dream
Run a curator pass against the validated memory directory and produce a proposal artifact.
/end
end
End a session — log what happened, update state and the decision log, propose memory updates, and check for uncommitted or unpushed work
/find-context
find-context
Find relevant context files by topic. Use when you need to load files for a topic without a slash command, or when a task spans multiple domains.
/migrate-gemini
migrate-gemini
Inventory and migrate selected Gemini CLI workflows with dry-run review and parity checks.
/mine-gemini-workflows
mine-gemini-workflows
Find repeated workflows in selected Gemini CLI sessions and draft portable skills after review.
/reconcile
reconcile
Scan multi-session drift and offer individually reviewed fixes only after explicit approval.
/recover
recover
Scan orphaned worktrees and stale branches, then offer explicit approval-gated cleanup.
/setup
setup
Guided onboarding or import for durable workspace context
/start
start
Start a session — load state files, flag staleness, and give a briefing on current priorities, deadlines, and blockers
/today
today
Create a morning heartbeat from repository state and update the local heartbeat log.
/update
update
Mid-session checkpoint — append progress to today's session log and update state files if a priority shifted, without ending the session
/distribution-audit
distribution-audit
Maintainer-only. Find every file that would newly ship to adopters, classify each one against the written distribution-boundary categories, default to withhold on no clean match, and ask the maintainer only where the taxonomy does not settle it. Drives the release CLI, which refuses to produce a manifest until every shipping file has an answer.
/gaia-audit
gaia-audit
Audit memory, wiki, and auto-loaded files for duplication, conflicting instructions, and stale content. The default path researches, then asks you a single Apply / Discuss / Decline question; on Apply it applies the report, files any out-of-scope problem as a tech-debt issue, then commits, opens a PR, and merges it on a main-branch run like /update-deps. Pass --apply to re-run the apply-and-publish stage against the most recent report.
/gaia-debt
gaia-debt
Fix the tech-debt backlog, a single issue or a recommended related batch, highest severity then oldest first, on a fresh isolated branch through the audit gate, closing the issue(s) on merge. Pass `list` to see the ordered backlog, `why <issue-number>` to explain the recommendation, or a bare `<issue-number>` to fix that issue directly.
/gaia-fitness
gaia-fitness
Health-check and auto-heal this project's Claude integration, triage, heal, verify, and report an F-to-A+ grade.
Your Personal AI Assistant; easy to install, deploy on your own machine or on the cloud; supports multiple chat apps with easily extensible capabilities.
16 views 0 likesAgent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it.…
17 views 0 likesCurated systems, benchmarks, and papers etc. on memory for LLMs/MLLMs --- long-term context, retrieval, and reasoning.
14 views 0 likes:memo: Vimlike Modal Text Editor in Rust
27 views 0 likesCI-native security testing for MCP servers. Attack simulation, schema drift detection, and health scoring before agents depend on them.
16 views 0 likesHermes Agent memory plugin/provider for scope-aware recall, SQLite truth, LanceDB semantic search, and hybrid retrieval.
15 views 0 likesFor You Agent——AI 时代的个人随身数字人格。把你的模型、AI 账号、技能、提示词和工作方式,带到每一个 AI 工具里。
12 views 0 likesA coding agent: give it a prompt and it reads, writes, runs commands, and searches code in a loop until the work is done, using native tool-calling across OpenA…
14 views 0 likesA secure persistent personal agent server in Rust. One binary, sandboxed execution, multi-provider LLMs, voice, memory, Telegram, WhatsApp, Discord, Teams, and…
14 views 0 likesSelf-evolving agent: grows skill tree from 3.3K-line seed, achieving full system control with 6x less token consumption
14 views 0 likesDeepSeek Harness Desktop (dsh-desktop). EAC: Embracing All Creation (揽尽万象). Bundled Node.js runtime with full dsh-CLI kernel, one-click startup, 10 built-in UI…
14 views 0 likesSee your agent think. Zero-config observability & governance for 26 AI agent runtimes: Claude Code, Cursor, OpenAI Codex, GitHub Copilot, Gemini CLI, Cline, Ope…
13 views 0 likesSave 94% on AI coding tokens. Index your codebase, agents search instead of reading files. Works with Claude Code, Codex, Copilot, Cursor, Gemini CLI. Local MCP…
14 views 0 likesYet another coding agent harness, lightweight and written in go.
14 views 0 likesa coding Agent from pi. ∞ providers, sub-agents, hashline edits, and a permission gate
13 views 0 likesOmnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting…
25 views 0 likes🧠 Leon is your open-source personal assistant.
14 views 0 likesThe Station, an open-world multi-agent environment that models a miniature scientific ecosystem.
14 views 0 likesThe Frontend Stack for Agents & Generative UI. React, Angular, Mobile, Slack, and more. Makers of the AG-UI Protocol
23 views 0 likesVelaTerm = iTerm2 + Codex, The Best Terminal for AI Coding
21 views 0 likes