LLM Mart Basic
@llm-mart · Joined Jun 2026
Stop everything and surface a rule conflict — persona vs. docs vs. code. Present both sides and the conflict-hierarchy level; the user resolves. No silent reconciliation.
Optional manual drift audit — report stale provenance-tracked docs, then per stale doc offer re-scout, re-baseline, or defer. Not the daily loop; commit-gate auto-heal owns routine maintenance.
Brownfield back-fill — scout an existing codebase and populate .codearbiter/, then lock it initialized.
Investigate-then-decide root-cause analysis for a defect whose cause is unknown. No code changes — exits to /ca-fix, /ca-adr, or a no-action close.
Greenfield decomposition interview — a layered interview that populates .codearbiter/ and locks it initialized.
Verify the active host install, package, command ownership, enforcement, wrapper self-test, and active-dispatch coverage gap. Read-only.
Start a feature: brainstorm a spec, get it approved, then drive it test-first through the pipeline. The one entry to implementation.
Fix a confirmed bug: a failing regression test first, then a minimal fix, then the rest of the tdd gates.
Opt this repo into codeArbiter — scaffold the root-level .codearbiter/ state store.
Read-only 3-metric governance glance — override rate, small-lane rate, sprint low-confidence ratio — each with a trend arrow vs. the prior 20-commit window.
Sanctioned, logged bypass of a gate or hard rule — one audit line, then proceed.
Open a pull request the only sanctioned way — clear every BLOCK-level review finding, then stage the PR. Never a direct write to the default branch.
Zero-onboarding, read-only dry-run of the reviewer fleet against the current uncommitted diff. Predicts reviewers, runs the state-free secret scan, writes nothing.
Trim transcript clutter to extend session lifetime — analyze, prune a copy, or toggle the after-each-turn service. Dry-run by default; gains land at resume/compaction, not the current turn.
SMARTS arbitration — reconcile architectural artifacts against the scaffold and prior decisions; every variance resolved by an explicit, user-attributed choice.
Restructure code with behavioral parity proven through unmodified pre-existing tests, then refactor. No behavior change.
Cut a release the only sanctioned way — SemVer bump from the commit log, a CHANGELOG section, an annotated tag. Takes the declared target's name as its only argument, or --dry-run to preview one with no write. The only path to a version tag.
Review a diff with the reviewer fleet, funneled to one triaged verdict. Targets the current working diff, a path, or an inbound GitHub PR.
Exploratory spike on a throwaway branch — answer a named question with disposable code. Never merges; exits to a findings note or /ca-feature.
Autonomous sprint — one interactive spec gate, then plan-to-PR execution with every auto-decision SMARTS-scored and logged. Hard gates remain true stops.
Fourteen posts of being wrong in production, compressed to checkboxes
Healthy nodes, a quiet network, 300 restarts in three days, and a latency budget measured in milliseconds
Discovery worked. Ping worked. Every TCP connection timed out, and later the tunnel only worked when someone had a terminal open.
Every VM came back. The cluster did not. Declarative systems converge on config, and the datapath isn't config.
A surprising share of AI-in-the-terminal failures aren't the AI. They're zsh, and a version of bash from 2006.
A Claude Code plugin turns standalone project configuration into a namespaced, installable extension that teams and communities can update as one unit.
None of the safety came from the model. It came from six boring habits.
Skills package instructions and references. Subagents run work in a separate context and return results. They solve different problems and can be composed deliberately.
Six hours in, one step left, everything green, and the incident that didn't happen
CLAUDE.md carries persistent project context. Skills load reusable procedures when relevant. Separating stable facts from task-specific workflows keeps both easier to maintain.
Twenty minutes recovering secrets that never existed, and the one sentence from a human that ended it
An API request routing a model's tool call through an approval gate to a remote MCP server
31 config keys, two audits, and why the first one was wrong in both directions
The official MCP Registry stores standardized server metadata rather than package code. Publishers verify a namespace, describe installation or remote access, and submit immutable versions.
Everyone looks at the Dockerfile. The file that actually leaked the key was the project file.
Remote MCP authorization uses established OAuth standards, but secure integration still requires issuer validation, least-privilege scopes, protected token handling, and server-side enforcement.
"Copy it over and switch the reference" is two steps, and the outage lives in the one nobody checks
stdio fits local processes and prototypes. Streamable HTTP fits hosted services and shared integrations. The right choice follows where the capability runs and who must reach it.
The most important rule wasn't about what I could change. It was about what I was allowed to display.
Tools perform operations, resources expose readable context, and prompts provide reusable templates. Choosing the correct primitive makes an MCP server easier to understand and govern.
/health
Health
One-line reliability verdict (OK / DEGRADED / FAILING) for Claude Code usage
/slo
Slo
Print a quick SLO snapshot — completion rate, tool success rate, and error rate
/report
Report
Generate an evidence-backed CCAM executive, cost, reliability, or workflow report.
/runs
Runs
List live and persisted CCAM-launched Claude Code and Codex runs
/find-session
Find session
Search Agent Monitor sessions by cwd, model, or status and print the top matches.
/recent
Recent
List the N most recent Claude Code sessions from the Agent Monitor.
/replay
Replay
Summarize one Agent Monitor session by id — header plus a concise transcript recap.
/dag
Dag
Print the orchestration DAG edges (parent→child subagents) for a session.
/runs
Runs
List recent Workflow-tool fleet runs with status and agent counts.
/workflow
Workflow
Summarize the workflow intelligence for a session — stats, complexity, and top patterns.
/broadcast
broadcast
Run broadcast on a distilled Raw — update related existing pages conversationally
/distill
distill
Distill raw session records into refined Notes and Projects
/evolve
evolve
Save knowledge to the cortex vault (Notes or Projects) and update index
/genesis
genesis
Initialize cortex vault — set up config, vault structure, and rebuild index
/query
query
Manually search the cortex vault (Notes, Projects, Raw) for existing notes
/takeoff
takeoff
Create, resume, or clear a session hand-off baton (one per work line)
/dev
Dev
Command → Agent → Skill orchestration for end-to-end feature development.
/investigate
Investigate
Command → Agent → Skill orchestration for investigating and fixing bugs.
/company
Company
Brief the CEO: coordinate with the peer CTO and start or resume a founder project across 8 departments and 54 skills
/onboard
Onboard
Configure an optional active team profile in three short questions
Open-source, self-hosted AI media server. One server replaces your entire media stack, with a web app, native iPhone app, and an AI agent that gets things done.…
3 views 0 likesA curated list of tools built for Jev — TypeSafe AI's System One model for typed decisions.
3 views 0 likes🦦 Crayotter: A Multimodal AI-Agent for Video-Editing, Video-Composing, and Video Production. Powered by Multimodal LLMs for autonomous Text-to-Video agentic fr…
3 views 0 likesWebextension tool for Odoo
3 views 0 likes