LLM Mart Basic

@llm-mart · Joined Jun 2026

0 Followers 0 Reputation 11456 Contributions
Claude Skill property-based-testing

Use this skill when you need to turn invariants, generation domains, and shrinking strategies into reviewable property-test candidates; triggers include 基于属性的测试 and property-based test design.

0
Claude Skill quality-dashboard-design

Use this skill when you need evidence-bounded quality dashboard audiences, decision questions, panels, drill-downs, freshness, and alert boundaries; triggers include 质量仪表盘 and quality dashboard.

0
Claude Skill quality-debt-analysis

Use this skill when you need evidence-bounded quality-debt items, origins, impact, age, priority, ownership, and paydown tradeoffs; triggers include 质量债务 and quality debt.

0
Claude Skill quality-gate-design

Use this skill when you need evidence-bounded entry criteria, evidence requirements, owners, and exception paths for a delivery or release gate; triggers include 质量门禁 and quality gate.

0
Claude Skill quality-maturity-assessment

Use this skill when you need evidence-bounded quality-practice maturity dimensions, rubric anchors, evidence sufficiency, and improvement gaps; triggers include 质量成熟度 and quality maturity.

0
Claude Skill quality-metrics-design

Use this skill when you need evidence-bounded quality metric definitions, calculation rules, data sources, freshness, and anti-gaming boundaries; triggers include 质量指标 and quality metric.

0
Claude Skill quality-productivity-metrics

Use this skill when you need evidence-bounded quality and delivery metrics, denominators, attribution limits, gaming risk, and the Human-use boundary; triggers include 质量生产力 and quality productivity.

0
Claude Skill quality-risk-analysis

Use this skill when you need to identify and prioritize quality risks from product, change, and evidence inputs; triggers include quality risk analysis.

0
Claude Skill rag-quality-testing

Use this skill when you need evidence-bounded grounding, relevance, completeness, citation support, abstention, and answer-level evidence in RAG outputs; triggers include RAG 质量 and RAG quality.

0
Claude Skill rag-retrieval-testing

Use this skill when you need evidence-bounded query variants, chunking, filters, recall/precision proxies, ranking, freshness, and retrieval evidence; triggers include 检索结果 and retrieval result.

0
Claude Skill recovery-testing

Use this skill when you need evidence-bounded recovery-testing analysis and validation preparation; triggers include 恢复测试 and recovery-testing.

0
Cursor Skill scomp-link

End-to-end ML toolkit with 26 CLI commands. Use when training models, tuning hyperparameters, detecting data drift, generating HTML reports with charts, profiling datasets, detecting anomalies, forecasting time series, checking fairness, or serving models as REST APIs. Prefer ove

0
Claude Skill api-fetch-wrapper

Wrap a public HTTP API (Open-Meteo weather as the demo) with credential handling, error normalisation, and a single retry on transient network failures. Demonstrates the production-shaped baseline for any "skill that calls an external service" — env-based secrets, structured erro

0
Claude Skill csv-processor

Read a CSV file from disk, compute per-column min/mean/max for every numeric column, emit the result as JSON. Stdlib-only Python; no pandas, no numpy. Demonstrates the simplest possible "give me a file path, get back structured analysis" skill — a deliberate baseline for any skil

0
Claude Skill text-summarizer

Summarise a chunk of text down to roughly `length` words using the agent's configured LLM provider. Input shape `{ text: string, length?: number }` on stdin, JSON; output shape `{ summary: string }` on stdout, JSON. Minimal: ~50 lines, no streaming, no retries — a deliberate base

0
Claude Skill chrono-ai-service-manual

Unified operational manual for AI agents driving the Chrono AI service stack — NyxID (identity, services, orgs, OAuth clients, proxy) AND Ornn (skill lifecycle — search, pull, install, execute, build, upload, share). One skill, two halves, one identity bootstrap, one set of failu

0
Claude Skill ornn-agent-manual-cli

The manual an AI agent loads to operate Ornn — the model-agnostic skill-lifecycle API (an npm-style registry + CLI for agent skills) — via the NyxID CLI (`nyxid proxy request ornn-api …`). Load and follow this skill WHENEVER the user asks to do anything with Ornn skills or skills

0
Claude Skill ornn-agent-manual-http

Operational manual for AI agents using the Ornn skill-lifecycle API via direct HTTPS with a NyxID bearer token (`curl -H "Authorization: Bearer $TOKEN" …`). Once loaded, the host agent can search / pull / execute / build / upload / share skills end-to-end. Authoritative contract

0
Claude Skill explore

Use this whenever you need to know what is actually in a database, warehouse, or DuckDB file before you trust it: ranked inventory of what exists, column profiles, PII detection, grain and data-quality problems, verified join inference, Mermaid ER diagrams, guarded ad-hoc SQL pro

0
Claude Skill maintain

Use this to keep a dbt project and its semantic layer correct as the warehouse and the business change, including a semantic layer that is native Apache Ossie documents rather than dbt. It detects drift on four axes and proposes the fix: schema drift (source columns and tables ad

0
The agent is running in *your* shell

A surprising share of AI-in-the-terminal failures aren't the AI. They're zsh, and a version of bash from 2006.

devops ai-agents automation shell
Sep 26
How to create and share a Claude Code plugin

A Claude Code plugin turns standalone project configuration into a namespaced, installable extension that teams and communities can update as one unit.

agents security skill-md claude-code
Sep 25
The scaffolding that made it safe

None of the safety came from the model. It came from six boring habits.

git devops ai-agents claude-code
Sep 25
Claude Code skills vs. subagents: when to use each

Skills package instructions and references. Subagents run work in a separate context and return results. They solve different problems and can be composed deliberately.

agent-skills claude-code context
Sep 24
Knowing when to stop

Six hours in, one step left, everything green, and the incident that didn't happen

prompt-engineering ai ai-agents sre
Sep 24
CLAUDE.md vs. skills: where should Claude Code instructions live?

CLAUDE.md carries persistent project context. Skills load reusable procedures when relevant. Separating stable facts from task-specific workflows keeps both easier to maintain.

agent-skills claude-skills configuration context
Sep 23
Those are the other app's keys

Twenty minutes recovering secrets that never existed, and the one sentence from a human that ended it

security devops ai ai-agents
Sep 23
How to use remote MCP servers with the OpenAI Responses API

An API request routing a model's tool call through an approval gate to a remote MCP server

security mcp integrations open-api
Sep 22
The coverage audit before you delete the safety net

31 config keys, two audits, and why the first one was wrong in both directions

security devops ai-agents secrets-management
Sep 22
How to publish an MCP server to the official MCP Registry

The official MCP Registry stores standardized server metadata rather than package code. Publishers verify a namespace, describe installation or remote access, and submit immutable versions.

security mcp
Sep 21
Rotating a leaked credential, in the right order

Everyone looks at the Dockerfile. The file that actually leaked the key was the project file.

security devops ai-agents containers
Sep 21
MCP authentication explained: OAuth, scopes, and safe token handling

Remote MCP authorization uses established OAuth standards, but secure integration still requires issuer validation, least-privilege scopes, protected token handling, and server-side enforcement.

security mcp
Sep 20
Byte-identical or bust

"Copy it over and switch the reference" is two steps, and the outage lives in the one nobody checks

security kubernetes verification
Sep 20
MCP stdio vs. Streamable HTTP: which transport should you use?

stdio fits local processes and prototypes. Streamable HTTP fits hosted services and shared integrations. The right choice follows where the capability runs and who must reach it.

security mcp
Sep 19
Never let the AI print a secret

The most important rule wasn't about what I could change. It was about what I was allowed to display.

security kubernetes devops ai-agents
Sep 19
MCP tools vs. resources vs. prompts: when to use each

Tools perform operations, resources expose readable context, and prompts provide reusable templates. Choosing the correct primitive makes an MCP server easier to understand and govern.

security mcp
Sep 18
How to test an MCP server with MCP Inspector

Use MCP Inspector to connect to local or remote servers, inspect capabilities, call tools, read resources, test prompts, and diagnose failures before release.

debugging security mcp
Sep 17
How to build an MCP server in TypeScript: step-by-step

Build an MCP server in TypeScript with focused tools, validated schemas, local and remote transports, Inspector tests, and production security controls.

security ai mcp
Sep 15
What is an MCP server? A practical guide

An MCP server exposes tools, resources, or prompts through a standard protocol so an AI application can discover and use external capabilities.

agents ai agent-skills claude-skills
Sep 11
How to vet AI agent skills before installing them

Treat an AI agent skill as both an instruction package and a software dependency: inspect what it says, what it runs, what it can access, and how it updates.

agents security ai-agents ai-agent-skills
Sep 10
/adr Adr

Author a numbered, dated, user-attributed Architecture Decision Record under .codearbiter/decisions/.

0
/audit Audit

Assemble the governance record for a range — commits, overrides, ADRs, sprint auto-decisions, open questions, checkpoint findings — into one dated audit packet. Read-only.

0
/btw Btw

Lightweight Q&A about the project — answer from context and return, no routing, no state change.

0
/checkpoint Checkpoint

Periodic multi-reviewer sweep of the whole codebase — surfaces a triaged checkpoint report.

0
/chore Chore

Sanctioned lane for non-behavioral work — docs-only edits, dependency bumps, reverts. Type-scaled gates; no TDD demanded of prose.

0
/cleanup Cleanup

Finish an already-merged branch — classify the leftover artifacts, return to a fast-forwarded default checkout, and delete the merged local branch. Every discard confirmed per item; containment proven, never assumed.

0
/commands Commands

Show the codeArbiter command catalog — the public command list and what each routes to.

0
/commit Commit

Run the full commit gate — the only sanctioned path to a git commit.

0
/conflict Conflict

Stop everything and surface a rule conflict — persona vs. docs vs. code. Present both sides and the conflict-hierarchy level; the user resolves. No silent reconciliation.

0
/context-check Context check

Optional manual drift audit — report stale provenance-tracked docs, then per stale doc offer re-scout, re-baseline, or defer. Not the daily loop; commit-gate auto-heal owns routine maintenance.

0
/create-context Create context

Brownfield back-fill — scout an existing codebase and populate .codearbiter/, then lock it initialized.

0
/debug Debug

Investigate-then-decide root-cause analysis for a defect whose cause is unknown. No code changes — exits to /ca:fix, /ca:adr, or a no-action close.

0
/decompose Decompose

Greenfield decomposition interview — a layered interview that populates .codearbiter/ and locks it initialized.

0
/doctor Doctor

Verify the active host install, package, command ownership, enforcement, and harmless live-fire probe. Read-only.

0
/feature Feature

Start a feature: brainstorm a spec, get it approved, then drive it test-first through the pipeline. The one entry to implementation.

0
/fix Fix

Fix a confirmed bug: a failing regression test first, then a minimal fix, then the rest of the tdd gates.

0
/init Init

Opt this repo into codeArbiter — scaffold the root-level .codearbiter/ state store.

0
/metrics Metrics

Read-only 3-metric governance glance — override rate, small-lane rate, sprint low-confidence ratio — each with a trend arrow vs. the prior 20-commit window.

0
/new-skill New skill

Author a new codeArbiter skill: prove the gap is real, get the spec approved, then write it.

0
/override Override

Sanctioned, logged bypass of a gate or hard rule — one audit line, then proceed.

0
Podlite

Implementation of Podlite markup language

12 views 0 likes
Nomi

Open-source AI video workbench. Bring any model or your local ComfyUI, and let Claude Code / Codex / Cursor direct it over MCP — storyboard, references, generat…

14 views 0 likes
Ai Agent Book

《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码

8 views 0 likes
Distilly

Distilly — Distill how they think into reusable Skills for any Agent or Bot. Formerly Colleague Skill(原同事 Skill).

15 views 0 likes
Knote

本地优先的类飞书 Markdown 编辑器,内置可审改的 AI 助手 | Local-first WYSIWYG Markdown editor with a reviewable AI agent

13 views 0 likes
Uniterm

A lightweight all-in-one terminal with 20+ protocols — SSH, RDP, SFTP, databases, Kubernetes and more. With a built-in autonomous AI Agent that plans and runs m…

27 views 0 likes
Alife

一款专注于桌宠方向的AIAgent。特点是一键安装、功能齐全、极低开销、完全暴露上下文、全功能插件化、AI自主插件开发、永久唯一会话、类游戏引擎交互策略。具有极高的扩展性和拟人程度上限,非常适合想长期培养和自定义需求高的用户。

10 views 0 likes
Google Workspace Mcp

Control Gmail, Google Calendar, Docs, Sheets, Slides, Chat, Forms, Tasks, Search & Drive with AI - Comprehensive Google Workspace MCP Server & CLI Tool

12 views 0 likes
Goutoujunshi

一个先接住情绪、再分析关系并给出可执行策略的 Codex 恋爱军师,内置心理、法律、社会、人文、哲学、婚姻家庭与性学知识库,支持多元关系。

10 views 0 likes
Siclaw

AI-powered SRE platform — read-only infrastructure diagnostics with deep investigation, security governance, and team collaboration

27 views 0 likes
Cyrene Agent

An open-source AI desktop companion inspired by Cyrene, combining immersive Chat, personalized long-term memory, and an agentic Work mode.

12 views 0 likes
Bitterbot Desktop

A local-first AI agent with persistent memory, emotional intelligence, and a peer-to-peer skills economy.

14 views 0 likes
Agentconnect

The open-source, multi-agent alternative to Claude Tag. @ any agent, wherever work happens, your agents work alongside your team and each other, learning as the…

27 views 0 likes
Julius

Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds

13 views 0 likes
Firecrawl Mcp Server

🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.

14 views 0 likes
Open Claude In Chrome

Claude in Chrome, reverse-engineered and open-source. No domain blocklist. Any Chromium browser. Same 18 MCP tools, same performance.

13 views 0 likes
LetsFG

Agent-native flight & hotel search and booking — MCP server, CLI, and Python/JS SDKs. Hundreds of airlines plus the major booking sites, with per-flight reliabi…

19 views 0 likes
Pinvou Agent

Open-source desktop AI agent for tools, files, knowledge, workflows, and real deliverables.

12 views 0 likes
Career Ops

Open-source AI job search: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs l…

24 views 0 likes
Job Application Agent

Privacy-first Codex skill for discovering, validating, completing, and tracking job applications

13 views 0 likes