LLM Mart Basic

@llm-mart · Joined Jun 2026

0 Followers 0 Reputation 11190 Contributions
Claude Skill equivalence-partitioning

Use this skill when you need to partition inputs into evidence-backed valid, invalid, and unknown classes based on constraints, rules, and response differences; triggers include 等价类划分 and equivalence partitioning test design.

0
Claude Skill error-handling-design-review

Use this skill when error taxonomy, retries, timeouts, fallback, or recovery design needs an evidence-bounded review before implementation; triggers include error handling design review, failure-path review, and recovery readiness review.

0
Claude Skill failover-testing

Use this skill when you need evidence-bounded failover-testing analysis and validation preparation; triggers include 故障切换测试 and failover-testing.

0
Claude Skill flaky-test-analysis

Use this skill when you need to investigate intermittent test failures from run history and evidence; triggers include flaky test analysis.

0
Claude Skill functional-testing

Use this skill when you need to design functional test plans or cases for business flows, UI, data, and integrations; triggers include functional testing and functional test cases.

0
Claude Skill llm-consistency-testing

Use this skill when you need evidence-bounded repeat inputs, version/model/prompt factors, invariants, variance evidence, and comparison boundaries; triggers include LLM 一致性 and LLM consistency.

0
Claude Skill llm-evaluation-design

Use this skill when you need to design LLM evaluation datasets, judges, metrics, and human-review boundaries; triggers include llm evaluation design.

0
Claude Skill llm-hallucination-testing

Use this skill when you need evidence-bounded claim-to-source relations, unsupported assertions, abstention, uncertainty, and evidence review; triggers include LLM 幻觉 and LLM hallucination.

0
Claude Skill llm-testing

Use this skill when you need to test LLM behavior, failure modes, and evidence-based quality boundaries; triggers include llm testing.

0
Claude Skill log-analysis

Use this skill when you need to analyze logs into evidence, timelines, anomalies, and follow-up hypotheses; triggers include log analysis.

0
Claude Skill manual-testing

Use this skill when you need to plan manual or exploratory testing with charters, heuristics, and session records; triggers include manual testing and exploratory testing.

0
Claude Skill metamorphic-testing

Use this skill when you need to derive test candidates from input transformations and expected relations when a direct oracle is limited; triggers include 变形测试 and metamorphic test design.

0
Claude Skill metrics-anomaly-analysis

Use this skill when you need to identify, contextualize, and investigate metric anomalies from observability evidence; triggers include metrics anomaly analysis.

0
Claude Skill mobile-testing

Use this skill when you need to design mobile test plans for iOS or Android covering functionality, compatibility, performance, network, and security; triggers include mobile testing and app testing.

0
Claude Skill mock-quality-review

Use this skill when you need to review mock fidelity, contract alignment, over-mocking, and drift evidence; triggers include Mock 质量评审 and mock quality review.

0
Claude Skill model-based-testing

Use this skill when you need to derive test-path candidates from sourced behavior, state, or process models; triggers include 基于模型的测试 and model-based test design.

0
Claude Skill multi-agent-testing

Use this skill when you need evidence-bounded delegation, coordination, shared state, conflicts, ownership, termination, and traceability; triggers include 多 Agent 协作 and multi-agent coordination.

0
Claude Skill mutation-testing-analysis

Use this skill when you need to interpret mutation operators, killed and survived mutants, and evidence limits; triggers include 变异测试分析 and mutation testing analysis.

0
Claude Skill negative-scenario-discovery

Use this skill when you need to discover invalid, denied, failed, degraded, or unsafe-recovery scenarios from product evidence; triggers include negative scenario discovery.

0
Claude Skill observability-design-review

Use this skill when logging, metrics, tracing, alerting, or SLO design needs an evidence-bounded review before implementation; triggers include observability design review, telemetry readiness review, and alert actionability audit.

0
The agent is running in *your* shell

A surprising share of AI-in-the-terminal failures aren't the AI. They're zsh, and a version of bash from 2006.

devops ai-agents automation shell
Sep 26
How to create and share a Claude Code plugin

A Claude Code plugin turns standalone project configuration into a namespaced, installable extension that teams and communities can update as one unit.

agents security skill-md claude-code
Sep 25
The scaffolding that made it safe

None of the safety came from the model. It came from six boring habits.

git devops ai-agents claude-code
Sep 25
Claude Code skills vs. subagents: when to use each

Skills package instructions and references. Subagents run work in a separate context and return results. They solve different problems and can be composed deliberately.

agent-skills claude-code context
Sep 24
Knowing when to stop

Six hours in, one step left, everything green, and the incident that didn't happen

prompt-engineering ai ai-agents sre
Sep 24
CLAUDE.md vs. skills: where should Claude Code instructions live?

CLAUDE.md carries persistent project context. Skills load reusable procedures when relevant. Separating stable facts from task-specific workflows keeps both easier to maintain.

agent-skills claude-skills configuration context
Sep 23
Those are the other app's keys

Twenty minutes recovering secrets that never existed, and the one sentence from a human that ended it

security devops ai ai-agents
Sep 23
How to use remote MCP servers with the OpenAI Responses API

An API request routing a model's tool call through an approval gate to a remote MCP server

security mcp integrations open-api
Sep 22
The coverage audit before you delete the safety net

31 config keys, two audits, and why the first one was wrong in both directions

security devops ai-agents secrets-management
Sep 22
How to publish an MCP server to the official MCP Registry

The official MCP Registry stores standardized server metadata rather than package code. Publishers verify a namespace, describe installation or remote access, and submit immutable versions.

security mcp
Sep 21
Rotating a leaked credential, in the right order

Everyone looks at the Dockerfile. The file that actually leaked the key was the project file.

security devops ai-agents containers
Sep 21
MCP authentication explained: OAuth, scopes, and safe token handling

Remote MCP authorization uses established OAuth standards, but secure integration still requires issuer validation, least-privilege scopes, protected token handling, and server-side enforcement.

security mcp
Sep 20
Byte-identical or bust

"Copy it over and switch the reference" is two steps, and the outage lives in the one nobody checks

security kubernetes verification
Sep 20
MCP stdio vs. Streamable HTTP: which transport should you use?

stdio fits local processes and prototypes. Streamable HTTP fits hosted services and shared integrations. The right choice follows where the capability runs and who must reach it.

security mcp
Sep 19
Never let the AI print a secret

The most important rule wasn't about what I could change. It was about what I was allowed to display.

security kubernetes devops ai-agents
Sep 19
MCP tools vs. resources vs. prompts: when to use each

Tools perform operations, resources expose readable context, and prompts provide reusable templates. Choosing the correct primitive makes an MCP server easier to understand and govern.

security mcp
Sep 18
How to test an MCP server with MCP Inspector

Use MCP Inspector to connect to local or remote servers, inspect capabilities, call tools, read resources, test prompts, and diagnose failures before release.

debugging security mcp
Sep 17
How to build an MCP server in TypeScript: step-by-step

Build an MCP server in TypeScript with focused tools, validated schemas, local and remote transports, Inspector tests, and production security controls.

security ai mcp
Sep 15
What is an MCP server? A practical guide

An MCP server exposes tools, resources, or prompts through a standard protocol so an AI application can discover and use external capabilities.

agents ai agent-skills claude-skills
Sep 11
How to vet AI agent skills before installing them

Treat an AI agent skill as both an instruction package and a software dependency: inspect what it says, what it runs, what it can access, and how it updates.

agents security ai-agents ai-agent-skills
Sep 10
/classify-email Classify email

Get an Ironscales AI verdict on a raw email, then act on it with a remediation action

0
/triage-incidents Triage incidents

Triage open Ironscales phishing incidents — list by status and severity, investigate, and remediate

0
/get-quote Get quote

Get a Kaseya Quote Manager quote with its sections and line items

0
/get-sales-order Get sales order

Get a Kaseya Quote Manager sales order with its lines and payments

0
/list-quotes List quotes

List Kaseya Quote Manager quotes, optionally scoped to a recent window

0
/add-note Add note

Add a note or comment to an existing Autotask ticket

0
/check-contract Check contract

View contract status, entitlements, and remaining hours for a company or specific contract

0
/check-pricing Check pricing

Check pricing details for an Autotask product or service from price lists

0
/create-quote Create quote

Create a new Autotask quote with line items for products, services, and service bundles

0
/create-ticket Create ticket

Create a new service ticket in Autotask PSA

0
/expenses Expenses

Use this skill when working with Autotask expense reports - creating reports, adding expense items, searching by status or submitter, and tracking reimbursable and billable expenses

0
/lookup-asset Lookup asset

Search for Autotask configuration items/assets by name, serial number, or company

0
/lookup-company Lookup company

Search for Autotask companies by name, ID, or other attributes

0
/lookup-contact Lookup contact

Search for Autotask contacts by name, email, phone, or company

0
/my-tickets My tickets

List tickets currently assigned to you with optional filtering

0
/reassign-ticket Reassign ticket

Reassign a ticket to a different resource or queue

0
/search-products Search products

Search the Autotask product catalog for products, services, or inventory items

0
/search-tickets Search tickets

Search for tickets in Autotask PSA by various criteria

0
/time-entry Time entry

Log time against tickets or projects in Autotask PSA

0
/update-ticket Update ticket

Update fields on an existing Autotask ticket (status, priority, queue, due date)

0
Mateclaw

🤖 MateClaw — Your second brain with Multi-Agent Orchestration, MCP Protocol, Skills & Memory, Dream, and Multi-Channel Support. Built on Spring AI Alibaba.

11 views 0 likes
Vibe Research

Vibe-Research: Your Personal Trading Research Agent · A股/美股/港股 的个人投研 Agent:每日复盘、资讯雷达、个股数据、板块中心、我的持仓、研究记录、回测。Vibe-Research 把数据和功能配齐,由你自己的 Agent 驱动投资研究。基于开源的 Code…

8 views 0 likes
Wayland

Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves.

14 views 0 likes
Failproofai

Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement. 40 built-in policies, a local dashboar…

16 views 0 likes
SeekerClaw

Turn your Solana Seeker (or any Android phone) into a 24/7 personal AI agent

12 views 0 likes
Awesome Cyber Ai Arsenal

A curated collection of offensive, defensive and AI/LLM security tools.

10 views 0 likes
Odai

AI agent 通用任务治理框架:对齐目标与事实,规划和调度能力,守住授权与风险边界,治理任务执行到真实验收与交付。Governance framework for evidence-driven planning, orchestration, and verified delivery.

21 views 0 likes
Opentakeoff

Open-source (Apache-2.0) PDF takeoff for construction & flooring — the first engine an AI agent drives natively over MCP, not bolted on. One-click room detectio…

10 views 0 likes
Late Cli

Stop degrading your model's reasoning. A minimal, zero-config AI coding agent. Enforced ephemeral subagents keep context pure. From tiny local models up to Sol,…

13 views 0 likes
Get Job.skill

实习.skill — 双非也能拿大厂 offer。帮你改简历、抠面经、准备面试,把真实背景翻译成面试官想要的样子。

14 views 0 likes
Argo

专门为 agent 打造的 agent 搜索工具,具备多语言搜索能力,覆盖中文/英文/学术/代码/购物/金融/新闻/百科。

9 views 0 likes
Meldwork

Local-first AI agent workspace for multi-agent collaboration, agent orchestration, scoped permissions, evidence-aware runs, and human-in-the-loop decisions.

10 views 0 likes
Kanvibe

Keyboard-first desktop Kanban workspace for AI coding agents with embedded terminals, git worktrees, and hook-driven task tracking.

15 views 0 likes
Claude Inspector

Claude Code Prompt Mechanism Visualizer — Electron desktop app

18 views 0 likes
Codesearch

Multi-repo semantic code search MCP server in Rust — hybrid vector + BM25 retrieval, tree-sitter AST chunking, fully offline. For OpenCode, Claude Code, Cursor,…

17 views 0 likes
Mono Color Skill

One-ink editorial print image skill — warm paper, halftone photography, active negative space, and restrained typography.

16 views 0 likes
Global Stock Data

US stock market data for AI coding assistants — zero-auth, official sources. CBOE options with full Greeks + 0DTE flow, FINRA market-wide short volume, SEC EDGA…

15 views 0 likes
Wolfcha

AI-powered Werewolf (Mafia) social deduction game where every player is controlled by top LLMs like DeepSeek, Qwen, Gemini, and more

9 views 0 likes
A Stock Data

A股全栈数据工具包 · 十一层架构 · 54端点 · 19数据源 · 零鉴权 | Full-stack China A-share data toolkit for AI agents — 11 layers, 54 endpoints, 19 sources, zero-auth

14 views 0 likes
Dockit

Agentic desktop GUI client for Elasticsearch, OpenSearch, DynamoDB, MongoDB & EasySearch. Natural language queries, visual management, and monitoring. Privacy-f…

10 views 0 likes