LLM Mart Basic
@llm-mart · Joined Jun 2026
Investigate a codebase, compare tools or approaches, verify current external or library facts, and research failures before making a substantive technical recommendation. Use for sufficiency and gap reviews; ordinary edits with settled requirements do not need a research stage.
Review authorized application code, configuration, designs or artifacts for concrete authorization, data exposure, injection, secret, dependency and LLM/tool security risks. Use for a requested security review or a change affecting a trust boundary.
Write, revise or review an agent skill (SKILL.md package).
Use when explaining System V AMD64, ARM AAPCS, RISC-V psABI, stack frames, variadic calls, or FFI register rules. Not for the Rust FFI binding layer: use rust-ffi.
Use when the user wants to classify abstractions as useful, bad, or busy and keep one shallow level. Not for tasks requiring source or remote-system changes.
Use when configuring ADC sampling time, DMA-driven ADC, calibration, or DAC channel setup on bare-metal MCUs. Not for the DMA stream itself: use dma-baremetal.
Use when creating AF_XDP sockets, configuring UMEM and XSK rings, writing an XDP redirect program, or choosing copy versus zero-copy mode. Not for full kernel bypass: use dpdk.
Use when a completed session needs an agent-environment retrospective. Not for an engineering retrospective from telemetry: use engineering-retrospective.
Use when asked to audit or repair agent surfaces (plugins, agents, skills, CLAUDE.md/AGENTS.md, docs, prompts, commands, hooks) or improve one skill at depth. Not for agent grading: use skill-doctor.
Use when a redacted, trimmed agent transcript must be appended to a GitHub PR or issue body, with human approval and preview. Not for automated or model-initiated insertion.
Use when the user describes an AI workflow gap or uses an ambiguous cross-session reference such as 'the PR Bob mentioned'. Not for tasks that require source or remote-system changes.
Use when the user requests a deep dive, exploratory analysis, or data analysis on BigQuery. Not for credential, publish, deploy, or irreversible changes.
Use when asked to design or change a public API, route, CLI flag, or module boundary. Not for remote, credential, publish, deploy, or irreversible changes.
Use when non-trivial code needs a design, codebase design or architecture needs improving, or one module needs targeted interface narrowing, seams, or testability. Not for diagrams, deploy, or irreversible changes.
Use when writing or porting AArch64 SIMD to SVE or SVE2: arm_sve.h intrinsics, predicates, vector-length-agnostic loops, auto-vectorization, or SVE registers in GDB. Not for NEON: use simd-intrinsics.
Use when the user knows what they mean but cannot express it completely or clearly. Not for discovery, ideation, or style-only editing: use unslop for style.
Use when asked to run /artifact-arena to generate and judge competing artifact implementations. Not for remote, credential, publish, deploy, or irreversible changes.
Use when eliciting intent/scope/referents or gating long/bundled/high-stakes/hard-to-undo work: exhaustive/collaborative/adversarial/gate/batch/interview/scan/proposal. Not for one fork: use decide.
Use when reading or writing AArch64 or AArch32 Thumb assembly, inline asm in C, AAPCS64 register roles, or NEON and SVE vector code. Not for ABI detail across ISAs: use abi-and-calling-conventions.
Use when reading or writing RV32/RV64 assembly, inline asm in C, the RISC-V psABI, IMAFD extension naming, compressed instructions, or QEMU RISC-V debugging.
Fourteen posts of being wrong in production, compressed to checkboxes
Healthy nodes, a quiet network, 300 restarts in three days, and a latency budget measured in milliseconds
Discovery worked. Ping worked. Every TCP connection timed out, and later the tunnel only worked when someone had a terminal open.
Every VM came back. The cluster did not. Declarative systems converge on config, and the datapath isn't config.
A surprising share of AI-in-the-terminal failures aren't the AI. They're zsh, and a version of bash from 2006.
A Claude Code plugin turns standalone project configuration into a namespaced, installable extension that teams and communities can update as one unit.
None of the safety came from the model. It came from six boring habits.
Skills package instructions and references. Subagents run work in a separate context and return results. They solve different problems and can be composed deliberately.
Six hours in, one step left, everything green, and the incident that didn't happen
CLAUDE.md carries persistent project context. Skills load reusable procedures when relevant. Separating stable facts from task-specific workflows keeps both easier to maintain.
Twenty minutes recovering secrets that never existed, and the one sentence from a human that ended it
An API request routing a model's tool call through an approval gate to a remote MCP server
31 config keys, two audits, and why the first one was wrong in both directions
The official MCP Registry stores standardized server metadata rather than package code. Publishers verify a namespace, describe installation or remote access, and submit immutable versions.
Everyone looks at the Dockerfile. The file that actually leaked the key was the project file.
Remote MCP authorization uses established OAuth standards, but secure integration still requires issuer validation, least-privilege scopes, protected token handling, and server-side enforcement.
"Copy it over and switch the reference" is two steps, and the outage lives in the one nobody checks
stdio fits local processes and prototypes. Streamable HTTP fits hosted services and shared integrations. The right choice follows where the capability runs and who must reach it.
The most important rule wasn't about what I could change. It was about what I was allowed to display.
Tools perform operations, resources expose readable context, and prompts provide reusable templates. Choosing the correct primitive makes an MCP server easier to understand and govern.
/cite
cite
Generate a formatted bibliography from the most recent research session's findings.
/dig
dig
Interactively refine research results by searching deeper into a specific subtopic or channel. Requires an active research session from /tome:research.
/export
export
Export research findings to a format compatible with memory-palace's knowledge-intake skill.
/research
research
Run a multi-source research session searching GitHub, HN, Lobsters, Reddit, arXiv, and Semantic Scholar. Use for multi-channel topic surveys.
/setup
Setup
Configure Azure MCP server with Azure CLI authentication
/load-claude-md
Load claude md
Refresh context with CLAUDE.md instructions
/sync-allowlist
sync-allowlist
Sync allowlist from GitHub repository to user settings
/sync-claude-md
Sync claude md
Sync CLAUDE.md from GitHub repository
/update-readme
Update readme
Update README.md plugin sections and download links
/setup
Setup
Configure GCloud CLI authentication
/setup
Setup
Configure Paper Search MCP (requires Docker)
/functions
Functions
Manage Supabase Edge Functions.
/gen
Gen
Automatically generates type definitions based on your Postgres database schema.
/init
Init
Initialize configurations for Supabase local development.
/link
Link
Link your local development project to a hosted Supabase project.
/login
Login
Connect the Supabase CLI to your Supabase account by logging in with your [personal access token](https://supabase.com/dashboard/account/tokens).
/network-bans
Network bans
Network bans are IPs that get temporarily blocked if their traffic pattern looks abusive (e.g. multiple failed auth attempts).
/projects
Projects
Provides tools for creating and managing your Supabase projects.
/secrets
Secrets
Provides tools for managing environment variables and secrets for your Supabase project.
/start
Start
Starts the Supabase local development stack.
PiG (Pi in Go) is a faithful Go port of upstream Pi, the TypeScript codebase behind the Pi coding agent. It is a parity-bound translation, not a rewrite: upstre…
1 views 0 likesAn AI Agent that lives in your pocket. Local-first and privacy focused.
3 views 0 likesUnofficial skill that teaches coding agents to build with TypeSafe AI's Jev: typed decisions, calibrated confidence, and prior art from 150+ community projects.
6 views 0 likesAdaptive Test-time Learning and Autonomous Specialization
4 views 0 likesPrediction-market trading engine — Wang Transform pricing on 291K+ contracts; paper-traded across Kalshi · Polymarket · Solana DFlow (Jito bundles) · 633 tests
3 views 0 likesKnowledge Management for Humans and Agents
5 views 0 likesOpen-source Claude Cowork / Codex / WorkBuddy alternative — a local-first AI office agent that turns one request into real PPTX, DOCX, XLSX and HTML files. Runs…
5 views 0 likesDeepAgent Code: AI coding agent with persistent memory and control plane
4 views 0 likesAwesome Jev — evidence-graded index of TypeSafe System One: SDKs, MCP tools, agents, apps and open models. 20 languages, rebuilt every 2 hours.
5 views 0 likesCLI for Telegram — agent-friendly, daemon-based, with webhook event push.
5 views 0 likesAI deep-research agent that turns any question into a cited report: plans searches, reads real sources, verifies evidence. Self-hosted, multi-provider, Docker-r…
4 views 0 likesEvent-stream AI Agent framework for building your persona bot 🍊
1 views 0 likesGive the agent a machine. Just not yours. Each AI coding agent gets its own isolated machine with root, Docker, and systemd - active defense detects and stops t…
3 views 0 likesLocal Emperor-style AI agent with Vue WebUI, multi-provider LLMs, streaming chat, tools, skills, memory, and token telemetry.
2 views 0 likesOpen-source AI reverse-engineering agent platform and MCP server for Ghidra, Frida, x64dbg and Rizin — automated PE/APK/binary analysis, CTF and malware researc…
8 views 0 likes"Never send a human to do a machine's job" - Open Source AI hacking agent
3 views 0 likesPrismer Cloud
3 views 0 likesMy Personal Blog (Robotics)
3 views 0 likesTau Coding Agent - like Pi, but twice as much
1 views 0 likesOpen-source alternative to OpenAI Dots: self-hosted AI chat, tools, approvals, connectors, and computer tasks.
0 views 0 likes