LLM Mart Basic
@llm-mart · Joined Jun 2026
Use when auditing the developer-facing surface of a CLI, SDK, library, or package: API contracts, errors, public types, onboarding, and config.
Use when recent resolved feedback may reveal a broader recurring defect pattern across the project surface. Not for source-level feedback collection: use feedback-sweep.
Use when one named review viewpoint must run fix cycles until a fresh reviewer finds nothing. Not for multi-viewpoint review or remote, credential, publish, deploy, or irreversible changes.
Use when the user invokes this skill to generate targeted questions proving the author understands the change's codebase effect. Not for reviewing the change: use review.
Use when asked to review a pull request, examine code changes, find bugs, or audit a branch, in standard or depth mode. Not for an iterative review-and-fix loop: use audit-project.
Use when the user wants a per-finding visual walk through a diff or PR. Not for written review reports: use review. Not for codebase tours: use show-me.
Use when implementation must be checked against an authoritative specification, or during PR review for spec drift against checked-in specs. Not for spec updates: use spec-driven-implementation.
Use when the user runs /browser-qa for report-only QA results without entering a fix loop. Not for remote, credential, publish, deploy, or irreversible changes.
Use when asked to reproduce, profile, or verify CLI/TUI behavior. Produces a deterministic transcript or profile proof with session cleanup. Not for CLI design advice, use cli-for-agents.
Use when asked to verify or reproduce browser or Electron UI behavior with before-and-after evidence and no leftover processes. Not for remote, credential, publish, deploy, or irreversible changes.
Use when asked to prove coverage, find missing cases, or enumerate state, decision, requirement, or behavior space. Not for round-based or single-property tests: use askme, property-test-authoring.
Use when a complete product needs production-like acceptance evidence against documented acceptance criteria. Not for single-component evaluation or evaluation without documented criteria.
Use when verification is looping, would re-run untouched code, or duplicates an established proof. Not for tasks that require source or remote-system changes.
Use when a test surface needs behavior-guarding coverage raised to a configured target with mutation kill evidence. Not for line-coverage inflation without mutation proof.
Use when asked to initialize, scope, estimate, configure, validate, or optimize a mewt, muton, or mutation testing campaign before execution. Writes the TOML config. Not for running it: use the mewt CLI.
Use when a mutation campaign leaves surviving mutants needing triage. Classifies each as false-positive, missing-test, genotoxic, or removable. Not for setup: use mutation-campaign-configuration.
Use when a product surface must be tested against extreme or hostile worlds. Not for design disputes: use possible-worlds. Not for remote, credential, publish, deploy, or irreversible changes.
Use when property-based testing, theorem proving, or formal proof tactics require zero unproven properties. Not for remote, credential, publish, deploy, or irreversible changes.
Operate explicit orchestrator, implementer, validator, and scribe roles through a caller-selected agent runtime. Triggers: "agent-native factory", "role-shaped agent panes", "persistent workers".
Use an explicitly selected AGY runtime for one provided packet or fresh validator context. Triggers: "agy", "antigravity", "AGY evidence".
Fourteen posts of being wrong in production, compressed to checkboxes
Healthy nodes, a quiet network, 300 restarts in three days, and a latency budget measured in milliseconds
Discovery worked. Ping worked. Every TCP connection timed out, and later the tunnel only worked when someone had a terminal open.
Every VM came back. The cluster did not. Declarative systems converge on config, and the datapath isn't config.
A surprising share of AI-in-the-terminal failures aren't the AI. They're zsh, and a version of bash from 2006.
A Claude Code plugin turns standalone project configuration into a namespaced, installable extension that teams and communities can update as one unit.
None of the safety came from the model. It came from six boring habits.
Skills package instructions and references. Subagents run work in a separate context and return results. They solve different problems and can be composed deliberately.
Six hours in, one step left, everything green, and the incident that didn't happen
CLAUDE.md carries persistent project context. Skills load reusable procedures when relevant. Separating stable facts from task-specific workflows keeps both easier to maintain.
Twenty minutes recovering secrets that never existed, and the one sentence from a human that ended it
An API request routing a model's tool call through an approval gate to a remote MCP server
31 config keys, two audits, and why the first one was wrong in both directions
The official MCP Registry stores standardized server metadata rather than package code. Publishers verify a namespace, describe installation or remote access, and submit immutable versions.
Everyone looks at the Dockerfile. The file that actually leaked the key was the project file.
Remote MCP authorization uses established OAuth standards, but secure integration still requires issuer validation, least-privilege scopes, protected token handling, and server-side enforcement.
"Copy it over and switch the reference" is two steps, and the outage lives in the one nobody checks
stdio fits local processes and prototypes. Streamable HTTP fits hosted services and shared integrations. The right choice follows where the capability runs and who must reach it.
The most important rule wasn't about what I could change. It was about what I was allowed to display.
Tools perform operations, resources expose readable context, and prompts provide reusable templates. Choosing the correct primitive makes an MCP server easier to understand and govern.
/generate-idl-client
Generate idl client
Generate a typed client from an Anchor or Shank IDL (Codama or Anchor TS)
/migrate-web3
Migrate web3
Migrate TypeScript from @solana/web3.js 1.x to @solana/kit
/plan-feature
Plan feature
Plan a Solana feature before coding: accounts, PDAs, instructions, risks, tests
/product-review
Product review
First-time-user product review: scorecard and fix roadmap; --harsh for a roast
/profile-cu
Profile cu
Measure compute units per instruction and flag the expensive ones
/quick-commit
Quick commit
Format, lint and commit with a conventional message on a kit-named branch
/resync
Resync
Resync external skill submodules to latest upstream versions
/scaffold
Scaffold
Scaffold a Solana project (Anchor, fullstack, frontend, Pinocchio) with the kit
/setup-ci-cd
Setup ci cd
Set up GitHub Actions CI for Solana programs: lint, build, test, audit
/setup-mcp
Setup mcp
Configure MCP server API keys in .env and add the optional MCP servers
/test-and-fix
Test and fix
Run tests, auto-fix fmt and lint, and fix failures until green or stuck
/test-dotnet
Test dotnet
Run C# tests: Unity Test Framework in batchmode, or dotnet test
/test-rust
Test rust
Run Rust tests for programs (LiteSVM, Mollusk, Surfpool, Trident) and backends
/test-ts
Test ts
Run TypeScript tests for programs (Anchor TS, Kit) and dApp frontends
/update
Update
Update solana-ai-kit to latest version from upstream
/write-docs
Write docs
Write docs for a Solana program, SDK or component from its code and IDL
/a
A
Allow all file creation (short for /ar:allow)
/afa
Afa
Set AutoFile policy to allow-all mode (full file creation permissions)
/afj
Afj
Set AutoFile policy to justify-create mode (require justification for new files)
/afs
Afs
Set AutoFile policy to strict-search mode (only modify existing files)
An open-source, privacy-first, self-hosted knowledge workspace where humans and AI agents work together 开源、隐私优先、自托管的知识工作空间,让人与智能体在此协作
14 views 0 likesScale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
17 views 0 likesThe go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
19 views 0 likesTransform and optimize your markdown documentation for Large Language Models (LLMs) and RAG systems. Generate llms.txt automatically.
29 views 0 likes合乎周礼:DeepSeek-powered Zhouli-style Chinese translator, web app, and distributable Skill package.
27 views 0 likesThis is a fork of the https://dockbox.dev project I made
25 views 0 likes🏆 Curated, ranked list of AI agent harnesses (100+) — plus an MCP server, llms.txt & JSON so agents can recommend them too. Rescored weekly.
28 views 0 likesA Python framework for modular, self-contained skill management for machines.
32 views 0 likesADHD — a skill for coding agents. Tree-of-thought with pruning, built on the Claude & Codex Agent SDK. Fans out parallel divergent thoughts under different cogn…
34 views 0 likesAI Agent 驱动的开源可自部署视频工作台:将小说与剧本转为角色、场景、道具资产、分镜、视频和剪映草稿,支持跨镜头一致性、多供应商与费用追踪 | Self-hosted AI video workspace for stories, storyboards and short-form video producti…
15 views 0 likesDeepSeek Harness Desktop App: a local AI desktop workspace for DSH Sessions, projects, files, web research, plugins, and Office artifacts.
13 views 0 likesAutonomous Offensive Security, Bug Bounty & Red Teaming Agent Framework powered by Hermes Agent, specialized reasoning skills, and multi-model LLM orchestration…
14 views 0 likes⌥ Coding agent with the IDE wired in
17 views 0 likesSupercharge AI Agents, Safely
33 views 0 likesX (Twitter) Scraper API and X API Alternative. You do not need an official X developer account. You do not need to connect or use an X account for supported scr…
16 views 0 likesMy AI Stand. Realtime by day, rewriting itself by night. Summon my AI superpower.
14 views 0 likesOpen-source coding agent for your terminal, built in Rust and on a journey of continuous community improvement. Issues and PRs welcome.
15 views 0 likes观澜 / Guanlan:AI Agent 的中文互联网研究、阅读与信源路由工具。
13 views 0 likesMac Agent for macOS 26: the agentic AI harness for your Mac Desktop. Computer use, automation, scripting, coding, and more. Powered by 18+ providers across loca…
15 views 0 likesSemantic version control => entity-level diffs, blame, and impact analysis on top of git. 28 languages via tree-sitter. Built for coding agents.
28 views 0 likes