LLM Mart Basic
@llm-mart · Joined Jun 2026
Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session. Splits history into an earlier baseline window and a recent window and reports
Compare two sessions side-by-side using Agent Monitor data — per-model token usage (input/output/cache_read/cache_write + compaction baselines), pricing engine cost breakdowns, workflow intelligence (complexity scores, tool flow transitions, subagent effectiveness), session metad
Inspect fired CCAM alerts and manage alert rules for token thresholds, event patterns, inactivity, and status duration. Use when acknowledging alerts, creating or editing a rule, checking cooldowns, or connecting alert rules to webhook targets.
Configure and troubleshoot CCAM Remote Data Sources that collect Claude Code and Codex history over SSH. Use when adding, editing, testing, syncing, or removing a remote machine, verifying provider paths, or deciding whether to retain or purge imported sessions.
Configure and validate CCAM webhook targets across supported chat, incident, automation, and generic providers. Use when listing provider requirements, creating or updating a target, scoping it to alert rules, sending a test notification, reviewing delivery history, or deleting a
Inspect and safely edit the Claude Code and Codex configuration surfaces exposed by CCAM. Use when auditing skills, agents, commands, plugins, marketplaces, MCP servers, hooks, settings, memory, keybindings, profiles, rules, or instruction files, and when a backup-backed allowlis
Import Claude Code or Codex history and move complete CCAM datasets between machines. Use when rescanning provider history, importing a copied directory, uploading JSONL or archives, exporting a backup, restoring it idempotently, or verifying that tokens, workflows, runs, rules,
Inspect and install CCAM monitoring hooks for Claude Code and Codex. Use when onboarding a provider, repairing missing hooks, checking which provider is active, or validating that installation preserved unrelated user hooks.
Configure, launch, validate, and troubleshoot CCAM's comprehensive MCP server for Claude Code, Codex, and other MCP hosts. Use when installing dependencies, building the server, selecting stdio, HTTP, or REPL transport, setting mutation/destructive policy, supplying dashboard aut
Generate a daily standup summary from recent Claude Code sessions — completed work grouped by project (cwd), session costs from the pricing engine, tool invocations, error/compaction/APIError events, and turn velocity metrics from session metadata (turn_count, total_turn_duration
Compile a month-over-month retrospective from Agent Monitor data — sessions, cost, token volumes, completion rate, top projects by working directory, and notable shifts versus the prior month. Uses daily_sessions/daily_events (365d) from analytics, the session list, and the prici
Summarize a sprint's worth of Claude Code activity — sessions grouped by project (cwd), per-model cost breakdown, token efficiency (cache hit rate, compaction baselines), subagent effectiveness from workflow API, velocity metrics (turn_count, turn_duration_ms), and tool diversity
Discover when you are most active and most productive with Claude Code by bucketing sessions and events into hour-of-day and day-of-week bins from their timestamps, then flagging peak versus low-output windows. Uses the session list, per-session events, and analytics daily trends
Compile a weekly productivity report using Agent Monitor data — daily_sessions and daily_events trends, per-session costs from pricing engine, token volumes (input/output/cache_read/cache_write + baselines), tool usage top 20, session completion rates by status, and workflow inte
Analyze workflow patterns using the Agent Monitor's workflow intelligence API — orchestration DAGs, tool flow transitions, subagent effectiveness, model delegation patterns, error propagation by depth, concurrency lanes, compaction impact, and agent co-occurrence. Produces priori
Produce a detailed report on APIError events from Agent Monitor data — counts over time, which sessions and models are affected, and the likely root cause (rate limits, overload/529, or context-window pressure) inferred from each event's summary and data payload. Use when API err
Scan recent Claude Code activity for errors and failure signals across all sessions using Agent Monitor data — APIError events and PreToolUse→PostToolUse gaps (tools that started but never completed) — then group failures by tool and model and rank them by frequency. Use when che
Audit hook delivery health from Agent Monitor data — balance PreToolUse vs PostToolUse (a gap means tools that started but never reported back), detect missing Stop/SubagentStop terminators (sessions/subagents that never closed), and check for stale ingestion (no recent events).
Compare this period's reliability against the prior period using Agent Monitor data — error rate (APIError/total) and tool-failure rate (PreToolUse→PostToolUse gap) — flag any regression where reliability got worse, and optionally wire a persistent alert rule so the dashboard cat
Define and check simple service-level objectives for Claude Code from Agent Monitor data — session completion rate, tool success rate (PostToolUse/PreToolUse), and error rate (APIError/total) — then compare each to its target and report the error budget remaining. Use when report
Fourteen posts of being wrong in production, compressed to checkboxes
Healthy nodes, a quiet network, 300 restarts in three days, and a latency budget measured in milliseconds
Discovery worked. Ping worked. Every TCP connection timed out, and later the tunnel only worked when someone had a terminal open.
Every VM came back. The cluster did not. Declarative systems converge on config, and the datapath isn't config.
A surprising share of AI-in-the-terminal failures aren't the AI. They're zsh, and a version of bash from 2006.
A Claude Code plugin turns standalone project configuration into a namespaced, installable extension that teams and communities can update as one unit.
None of the safety came from the model. It came from six boring habits.
Skills package instructions and references. Subagents run work in a separate context and return results. They solve different problems and can be composed deliberately.
Six hours in, one step left, everything green, and the incident that didn't happen
CLAUDE.md carries persistent project context. Skills load reusable procedures when relevant. Separating stable facts from task-specific workflows keeps both easier to maintain.
Twenty minutes recovering secrets that never existed, and the one sentence from a human that ended it
An API request routing a model's tool call through an approval gate to a remote MCP server
31 config keys, two audits, and why the first one was wrong in both directions
The official MCP Registry stores standardized server metadata rather than package code. Publishers verify a namespace, describe installation or remote access, and submit immutable versions.
Everyone looks at the Dockerfile. The file that actually leaked the key was the project file.
Remote MCP authorization uses established OAuth standards, but secure integration still requires issuer validation, least-privilege scopes, protected token handling, and server-side enforcement.
"Copy it over and switch the reference" is two steps, and the outage lives in the one nobody checks
stdio fits local processes and prototypes. Streamable HTTP fits hosted services and shared integrations. The right choice follows where the capability runs and who must reach it.
The most important rule wasn't about what I could change. It was about what I was allowed to display.
Tools perform operations, resources expose readable context, and prompts provide reusable templates. Choosing the correct primitive makes an MCP server easier to understand and govern.
/do-issue
do-issue
Implement issues (GitHub/GitLab/Bitbucket) using progressive analyze-specify-plan-implement workflow
/fix-pr
fix-pr
Address PR/MR review feedback by reading comments, implementing fixes, and resolving threads. GitHub and GitLab support.
/fix-workflow
fix-workflow
Retrospective analysis and improvement of workflow components with self-evolving patterns
/fixit
fixit
Fix broken functionality from pasted error output, stack traces, or
/git-catchup
git-catchup
Summarize recent git history since a baseline with structured analysis of what changed, why, and what to watch for.
/merge-docs
Merge docs
Consolidate ephemeral LLM-generated markdown into permanent documentation.
/pr-review
pr-review
Review pull requests with scope validation, code analysis, and line comments. Supports GitHub PRs and GitLab MRs.
/prepare-pr
prepare-pr
Prepare a PR end-to-end by updating documentation, running tests, dogfooding checks, and validating with code review.
/resolve-threads
resolve-threads
Batch-resolve unresolved PR/MR review threads via GraphQL API (GitHub/GitLab)
/sync-capabilities
Sync capabilities
Detect and fix drift between plugin.json registrations and capabilities reference documentation
/update-ci
Update ci
Update pre-commit hooks and CI/CD workflows based on recent project changes
/update-dependencies
update-dependencies
Scan and update dependencies across all ecosystems with conflict detection
/update-docs
Update docs
Update project documentation with consolidation, debloating, AI slop detection, capabilities sync, and accuracy verification.
/update-plugins
Update plugins
Audit and sync plugin.json registrations with actual disk contents. Detects missing or stale skills, commands, agents, hooks.
/update-tests
update-tests
Review and update test coverage using TDD/BDD methodology with quality validation. Generates tests for changed code.
/update-tutorial
update-tutorial
Generate or update tutorials with VHS and Playwright recordings
/update-version
Update version
Bump project versions using git-workspace-review and version-updates skills.
/validate-pr
validate-pr
Generate and self-execute a diff-derived test plan for a PR. Reads
/doc-generate
doc-generate
Generate new documentation with human-quality writing.
/doc-polish
doc-polish
Clean up AI-generated content and improve documentation quality.
面向中文开发者的 Claude Code Skills / Agents / Plugins 精选与原创技能库|按场景分类|复制即装|持续更新
17 views 0 likes原生 macOS Git 客户端,以 Agent 驱动仓库管理、审阅与协作。
15 views 0 likesA 股短线复盘看板:涨停池·连板梯队·龙虎榜·板块资金一屏看完,赚钱效应/晋级率/梯队断层/情绪周期等派生指标纯计算直出(不经过 AI),AI 只把数据串成能读的盘面研判。全本地运行,可用 Claude/Codex 订阅免 API key。| A-share short-term daily-review dashbo…
6 views 0 likes禁漫天堂 Agent Skills / AI 原生 JMComic 助手:通过 MCP 与 Skills 将 JMComic 注入你的 AI Agent. / AI-powered JMComic assistant for seamless integration with AI Agents via MCP & S…
13 views 0 likesPairlet (formerly CC Pocket / cc-pocket) — Continue your local AI coding tasks from phone, tablet, or desktop.
16 views 0 likesSelf-hosted Personal AI + agent runtime in .NET (NativeAOT-friendly)
7 views 0 likesA self-hosted knowledge platform for humans and AI agents — publish wikis, blogs, and portable Agent Skills.
14 views 0 likesNotion CLI with AI agent support. Smart queries, Obsidian sync, batch ops, backups, validation and more.
8 views 0 likesAI Coding Agent Multiplexer
16 views 0 likesIndependent executor–verifier orchestration for software changes.
9 views 0 likesAI-native, local-first video editor by Ribbi — runs entirely in the browser. Rust/WASM engine, WebGPU compositing, WebCodecs export; humans and LLM agents edit…
8 views 0 likesDistill your knowledge, memories, and decisions into an open-source, inspectable AI Agent Twin.
11 views 0 likes🔬 Harness Vibe Research with Self-evolving AI Scientists
3 views 0 likesPersonal AI desktop agent for Windows, macOS, Linux, Android & iOS. Set a goal, it works on its own. Teams (pair two desktops, agents + humans), Agent2Agent, Wo…
5 views 0 likesRun agents like a company. AgentOS is the native control plane for OpenClaw — manage agents, tasks, models, context, approvals, and runtime visibility from one…
15 views 0 likesCreate Agentic admin panels faster on TypeScript and Vue.js with AdminForth Framework. Setup main CRUD pages within minutes, extend as you need
9 views 0 likesOpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
18 views 0 likesAn ebook reader with a self-evolving agent: it remembers your reading across books, and plugins extend the reader and the agent alike.
12 views 0 likesA curated list of awesome resources for vibe coding
14 views 0 likesAgentic Voice Notes for iPhone and macOS - Rust, Dioxus, LanceDB + RIG + SQLite
9 views 0 likes