LLM Mart Basic
@llm-mart · Joined Jun 2026
Read a detached Codex bridge job's current state.
Ask Claude Code a read-only question from Codex through the sandbox-external bridge.
Ask Claude Code to implement one bounded write-capable change through the sandbox-external bridge.
Request a read-only adversarial review from Claude Code through the sandbox-external bridge.
Configure, authenticate, repair, and verify the Codex to Claude Code Bridge; safe stages run by default and sensitive stages require separate approval.
Safely manage Codex Rig role-agent shims: doctor, status, install, or remove; one action only.
Audit Codex configuration, workflow, and prompt-efficiency (instruction cost, value-per-token) drift; emit ranked gaps and measurable gates.
Calibrate skills/role cards for leaks/gaps with recall, precision, and confidence-accuracy checks.
Apply selected review fixes; bare PR targets use current online items, while PR +review adds the latest matching artifact.
Close PRs at an evidence gate or review local diffs/PRs with specialists and JSON artifacts.
Implement changes with a linear plan-build-verify workflow and measurable quality gates.
Investigate code debugging and root-cause narrowing; use measurable gates before fixes.
Build/extend grounded Kaggle Jupytext notebooks for training, EDA, inference, or resume workflows, grounding schema and submission format through the authenticated kaggle CLI.
Optimize a measurable metric with bounded iterations, guardrails, and regression gates.
Assess SemVer release readiness with gates/artifacts; never tag, publish, upload, or force-push.
Research docs, papers, or state of the art; provide source-backed recommendations and caveats.
Dry-run active plugin cache drift; refresh/reinstall only with approval; keep shims separate.
Adversarial review — drills to bedrock, treats claims as unproven until evidence. NOT for: plan design (foundry:solution-architect), test coverage (foundry:qa-specialist), config formatting (foundry:curator). TRIGGER: "challenge this", "devil's advocate", "poke holes in". SKIP: w
Content specialist — blog posts, slide decks, social threads, talk abstracts. Reads approved outline, applies four-beat arc. NOT for in-code docs/README/FAQs (foundry:doc-scribe), release notes (oss:release). TRIGGER: "write a blog post", "create slides", "draft a thread". SKIP:
Config quality reviewer. Scope: agents/skills/rules (*.md) — verbosity, duplication, cross-refs, roster overlap; applies fixes. NOT for hooks (foundry:sw-engineer), ADRs (foundry:solution-architect), adversarial challenge (foundry:challenger). TRIGGER: "audit this agent", "review
Fourteen posts of being wrong in production, compressed to checkboxes
Healthy nodes, a quiet network, 300 restarts in three days, and a latency budget measured in milliseconds
Discovery worked. Ping worked. Every TCP connection timed out, and later the tunnel only worked when someone had a terminal open.
Every VM came back. The cluster did not. Declarative systems converge on config, and the datapath isn't config.
A surprising share of AI-in-the-terminal failures aren't the AI. They're zsh, and a version of bash from 2006.
A Claude Code plugin turns standalone project configuration into a namespaced, installable extension that teams and communities can update as one unit.
None of the safety came from the model. It came from six boring habits.
Skills package instructions and references. Subagents run work in a separate context and return results. They solve different problems and can be composed deliberately.
Six hours in, one step left, everything green, and the incident that didn't happen
CLAUDE.md carries persistent project context. Skills load reusable procedures when relevant. Separating stable facts from task-specific workflows keeps both easier to maintain.
Twenty minutes recovering secrets that never existed, and the one sentence from a human that ended it
An API request routing a model's tool call through an approval gate to a remote MCP server
31 config keys, two audits, and why the first one was wrong in both directions
The official MCP Registry stores standardized server metadata rather than package code. Publishers verify a namespace, describe installation or remote access, and submit immutable versions.
Everyone looks at the Dockerfile. The file that actually leaked the key was the project file.
Remote MCP authorization uses established OAuth standards, but secure integration still requires issuer validation, least-privilege scopes, protected token handling, and server-side enforcement.
"Copy it over and switch the reference" is two steps, and the outage lives in the one nobody checks
stdio fits local processes and prototypes. Streamable HTTP fits hosted services and shared integrations. The right choice follows where the capability runs and who must reach it.
The most important rule wasn't about what I could change. It was about what I was allowed to display.
Tools perform operations, resources expose readable context, and prompts provide reusable templates. Choosing the correct primitive makes an MCP server easier to understand and govern.
/release
Release
This is the only sanctioned way a version tag gets created, and it holds no knowledge of any
/review
Review
A read-only review of the current change, dispatched by path against a reviewer-to-path matrix:
/spike
Spike
This is the outlet for "I need to write code to find out" rather than "I know what to build." The
/sprint
Sprint
This is codeArbiter's autonomy mode. It collapses the usual spec-plan-execute cycle into one
/standup
Standup
The daily hygiene checklist, made routine and gated: fetch and offer a fast-forward pull, list
/status
Status
A read-only snapshot of `.codearbiter/` state: the project's `stage:` maturity value, every
/statusline
Statusline
Wires codeArbiter's renderer into `~/.claude/settings.json`, or removes it. A plugin can't own a
/task
Task
The only blessed way to mutate `.codearbiter/open-tasks.md`. It runs a thin writer over pure
/threat-model
Threat model
An opt-in, lightweight STRIDE pass over a design before it's built — invoked deliberately, never
/tribunal
Tribunal
codeArbiter's deepest, most expensive review, convened rarely and never as a required gate. Eleven
/watch
Watch
Watches a pull request's CI checks to completion without polling by hand. The wait is a real
/debug
Debug
A fixture command named debug, colliding with the debug skill/agent.
/prune
Prune
A fixture command named prune, used to regression-test forge badges.
/another
Another
Another command, description only, no argument hint.
/sample
Sample
A sample command for fixtures — single sanctioned path to a thing.
/to-gif
To gif
Convert an animated HTML file (like /fig output) to a perfectly looping high-quality small GIF using Playwright + ffmpeg
/getpix
Getpix
Find a free licensed image on the internet and add it to the project, optimized
/photo-pass
Photo pass
Art-direction photo pass over the site: add photography only where it earns its place
/tree-ring-audit
Tree ring audit
Audit consolidate or forget stale sensitive or superseded Tree Ring Memory entries
/tree-ring-capture
Tree ring capture
Capture a concise validated lesson decision warning or preference in Tree Ring Memory
AI agent orchestration kit for Windows, Linux/MacOS with Codex skills, hooks, routing rules and profiles for Claude, OpenCode, Cursor, Gemini and Windsurf.
1 views 0 likesSkills & reviewer agents for AI-first climate science — built and used by a PhD atmospheric scientist
3 views 0 likes出行路书工作流 skill:联网实查 + 多源交叉验证,产出可核验、能执行的旅行攻略。覆盖吃住行游拍避全维度,附美食情报卡、基准骨架与校验工具,支持一键部署在线版。
3 views 0 likesSelf-hosted AI code reviewer with indexed PR reviews, walkthroughs, vulnerability scanning, dependency graphs, custom rules, and a learning loop.
3 views 0 likes🧬 Extend EvoScientist with Installable Skill & Knowledge Packs
2 views 0 likesGive the agent a machine. Just not yours. Each AI coding agent gets its own isolated machine with root, Docker, and systemd - active defense detects and stops t…
2 views 0 likesAutomatic memory consolidation for OpenClaw agents — like sleep for your AI. Powered by MyClaw.ai
0 views 0 likes🗂 The essential checklist for modern web development, for humans and AI agents
0 views 0 likesDistribb CLI, Claude, Codex, Hermes, OpenClaw skill for AI-powered SEO. Write content with your own AI, publish through Distribb's backlink network.
2 views 0 likesOpen-source AI browser agent — type plain-language commands and Bah operates the web for you. Works with cloud AI (DeepSeek, Mistral, NVIDIA) or local Ollama mo…
3 views 0 likesDeepSeek Harness Desktop (dsh-desktop). EAC: Embracing All Creation (揽尽万象). Bundled Node.js runtime with full dsh-CLI kernel, one-click startup, 10 built-in UI…
2 views 0 likesUse your ChatGPT / Codex subscription with DeepSeek Harness via OAuth, with model access, usage quotas, search, and image generation — no API key or Codex CLI r…
3 views 0 likesLocal-first A-share research workbench for DeepSeek Harness: market dashboards, watchlists, valuation, four investor agents, versioned reports, and continuous p…
2 views 0 likesOpen-source video pipeline (clipper, AI video & more) — scout trends, clip long videos, generate captioned shorts, design motion graphics. Local-first, BYOK. AG…
3 views 0 likesDesktop app for the pi coding agent: streaming timeline, Git review with hunk staging, session tree, native UI for pi extensions. Windows · macOS · Linux. pi 编程…
4 views 0 likesZero-dependency browser video editor that AI agents can drive — JSON timeline, MCP + REST, live-reloading UI
0 views 0 likesPersonal Context Manager for Claude Code. Your life in walnuts.
0 views 0 likesFast way to switch between Claude Code configuration profiles
3 views 0 likesAutoClip|一个链接,一键出片。开源 AI 视频剪辑桌面工具,将播客、访谈、课程等长视频自动剪成短视频,生成字幕、封面和发布文案,适配抖音、小红书、TikTok、Reels 与 YouTube Shorts。Open-source AI video clipping & content repurposing.
3 views 0 likes