LLM Mart Basic
@llm-mart · Joined Jun 2026
Use this skill to assign consistent risk labels to a Legal or HR matter — severity ratings, privilege and privacy sensitivity labels, retaliation and discrimination risk labels, matter-type classes, escalation-gate triggers, and the audit-log schema. It standardizes the vocabular
Use this skill when a Legal or HR matter must be classified and routed to the right specialist agent, when a matter crosses both domains and needs parallel review, or when Legal and HR agents disagree and the conflict must be resolved. It defines routing rules, the overlap handof
Use this skill when Microsoft 365 or Dynamics 365 licenses have been assigned and the organisation needs to measure adoption, instrument value outcomes, identify waste, and reclaim inactive licenses before purchasing more. Orchestrates microsoft-business-impact-value-realization-
Use this skill when a NetSuite matter must be classified and routed to the right specialist agent, when a matter crosses multiple NetSuite domains and needs parallel review, when a live-account mutation intent must be gated, or when specialist agents disagree and the conflict mus
Use this skill to coordinate the order-to-cash process across Dynamics 365 Supply Chain Management, Finance, and Sales, covering confirmed order through fulfillment, invoicing, accounts receivable, and cash collection. It defines stage ownership, gate conditions, agent handoff ru
Use this skill to orchestrate the procure-to-pay (source-to-pay) process across Dynamics 365 Supply Chain Management, Finance, and compliance-aware separation-of-duties governance. It coordinates the journey from purchase requisition through purchase order approval, goods or serv
Use this skill to review the cross-tier seams of revenue-critical journeys — checkout, payment submission, account creation, and login — for idempotency of money-moving and account-creating requests, server-side re-validation of client-enforced rules, webhook duplicate/out-of-ord
Use this skill when a Salesforce specialist agent must hand a matter to another agent and the context, uncertainty, evidence quality, privilege posture, and privacy posture must survive the handoff intact. Defines the shared salesforce-case-capsule — a controlled, auditable excha
Use this skill when a Salesforce data exposure event has been detected or is strongly suspected. Triggers include: guest-user data exposure via Experience Cloud, cross-org data sync without a Data Processing Agreement, regulated-data sync in Marketing Cloud without a consent map,
Use this skill when any proposed mutation to a live Salesforce production org must be evaluated before execution. This is a refusal-by-default gate: if any required precondition is missing, the skill stops and refuses. Required preconditions are target_org_identity, environment_t
Use this skill when a Salesforce matter must be assigned a standardized matter type, risk tier, or escalation gate before routing or handoff. Defines all matter types (org-config, automation, code, integration, security/IAM, data, sales/CPQ, service/SLA, experience-cloud, marketi
Use this skill when a Salesforce matter must be classified and routed to the right specialist agent, when a matter crosses multiple Salesforce domains and needs parallel review, or when specialist agents disagree and the conflict must be resolved. It defines routing rules per mat
Use this skill to statically review AI/BI Genie agent and dashboard design: agent scoping (30-table limit), instructions and trusted assets, metric-view correctness, dashboard limits and rendering, benchmark design and honest accuracy reading, and the critical 'Individual data' v
Use this skill to review Databricks data protection and privacy design for regulatory alignment and least-privilege enforcement: row filters and column masks, ABAC policies, data classification, deletion and GDPR erasure mechanics, Delta Sharing egress, residency and Geo constrai
Use this skill to design and verify data quality expectations, table constraints, Lakehouse Monitoring, freshness detection, event-log interrogation, quality SLAs, and downstream quality signaling for Lakeflow pipelines. Reads pipeline source, table schema, expectations, monitor
Use this skill to review a Declarative Automation Bundle configuration, authentication setup, and deployment flow against production readiness criteria: bundle structure, deployment modes, run-as identity boundaries, variable resolution timing, OAuth and environment-variable auth
Use this skill to statically review Databricks cost and cost-attribution: system.billing.usage and system.billing.list_prices for correct joins, custom-tag-based attribution with coverage-confidence reporting, DBU uptime charging semantics, serverless versus classic cost comparis
Use this skill to review generative-AI agent design on Databricks: Mosaic AI Agent Framework and ResponsesAgent interface, Databricks AI Search index variant and sync-mode choice, retrieval and context engineering, MCP server category and trust boundaries, external model-provider
Use this skill to review generative-AI evaluation, tracing, and observability design on Databricks: MLflow Tracing instrumentation and span design, trace storage and governance, `mlflow.genai.evaluate()` harness design, the judge-versus-scorer distinction, built-in judge selectio
Use this skill to review Databricks identity and network security design for proper admin separation, SCIM/federation configuration, credential hygiene, and network boundary enforcement: admin roles, service-principal posture, OAuth vs PAT, token lifecycle, IP access lists, serve
Fourteen posts of being wrong in production, compressed to checkboxes
Healthy nodes, a quiet network, 300 restarts in three days, and a latency budget measured in milliseconds
Discovery worked. Ping worked. Every TCP connection timed out, and later the tunnel only worked when someone had a terminal open.
Every VM came back. The cluster did not. Declarative systems converge on config, and the datapath isn't config.
A surprising share of AI-in-the-terminal failures aren't the AI. They're zsh, and a version of bash from 2006.
A Claude Code plugin turns standalone project configuration into a namespaced, installable extension that teams and communities can update as one unit.
None of the safety came from the model. It came from six boring habits.
Skills package instructions and references. Subagents run work in a separate context and return results. They solve different problems and can be composed deliberately.
Six hours in, one step left, everything green, and the incident that didn't happen
CLAUDE.md carries persistent project context. Skills load reusable procedures when relevant. Separating stable facts from task-specific workflows keeps both easier to maintain.
Twenty minutes recovering secrets that never existed, and the one sentence from a human that ended it
An API request routing a model's tool call through an approval gate to a remote MCP server
31 config keys, two audits, and why the first one was wrong in both directions
The official MCP Registry stores standardized server metadata rather than package code. Publishers verify a namespace, describe installation or remote access, and submit immutable versions.
Everyone looks at the Dockerfile. The file that actually leaked the key was the project file.
Remote MCP authorization uses established OAuth standards, but secure integration still requires issuer validation, least-privilege scopes, protected token handling, and server-side enforcement.
"Copy it over and switch the reference" is two steps, and the outage lives in the one nobody checks
stdio fits local processes and prototypes. Streamable HTTP fits hosted services and shared integrations. The right choice follows where the capability runs and who must reach it.
The most important rule wasn't about what I could change. It was about what I was allowed to display.
Tools perform operations, resources expose readable context, and prompts provide reusable templates. Choosing the correct primitive makes an MCP server easier to understand and govern.
/attach
Attach
`crabbox attach` follows the recorded events of an active coordinator run and
/azure
Azure
`crabbox azure` groups Azure provider setup commands. It currently has a single
/bench
Bench
`crabbox bench` records and reports local benchmark timing observations. It is a
/cache
Cache
`crabbox cache` inspects, purges, or warms package and build caches on a
/capsule
Capsule
`crabbox capsule` captures, replays, and tracks lightweight failure capsules.
/checkpoint
Checkpoint
Save the state of a lease, then restore it onto another box or fork it into a
/claims
Claims
`crabbox claims list` prints the lease claims stored on the current machine. It
/cleanup
Cleanup
`crabbox cleanup` sweeps direct-provider machines and local provider state that
/code
Code
`crabbox code` bridges a Linux lease's `code-server` workspace into the
/config
Config
`crabbox config` inspects and updates user configuration. It has three
/connect
Connect
`crabbox connect` resolves a lease and opens an interactive SSH session to it.
/cp
Cp
`crabbox cp` copies files or directories between the host and a Crabbox-owned
/desktop
Desktop
`crabbox desktop` drives a visible desktop session on a lease that was warmed
/doctor
Doctor
`crabbox doctor` runs a preflight before you commit to a long workflow. It is
/egress
Egress
`crabbox egress` gives a lease mediated outbound network: a lease-local browser
/events
Events
`crabbox events` prints the broker's event log for a recorded run.
/heartbeat
Heartbeat
`crabbox heartbeat` refreshes the idle deadline for one owned lease and prints
/history
History
`crabbox history` lists recorded remote command runs from the broker. Each run is
/image
Image
`crabbox image` holds the trusted-operator controls for provider base images:
/init
Init
`crabbox init` onboards the current repository: it writes the minimal config
Your Personal AI Assistant; easy to install, deploy on your own machine or on the cloud; supports multiple chat apps with easily extensible capabilities.
16 views 0 likesAgent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it.…
17 views 0 likesCurated systems, benchmarks, and papers etc. on memory for LLMs/MLLMs --- long-term context, retrieval, and reasoning.
14 views 0 likes:memo: Vimlike Modal Text Editor in Rust
27 views 0 likesCI-native security testing for MCP servers. Attack simulation, schema drift detection, and health scoring before agents depend on them.
16 views 0 likesHermes Agent memory plugin/provider for scope-aware recall, SQLite truth, LanceDB semantic search, and hybrid retrieval.
15 views 0 likesFor You Agent——AI 时代的个人随身数字人格。把你的模型、AI 账号、技能、提示词和工作方式,带到每一个 AI 工具里。
12 views 0 likesA coding agent: give it a prompt and it reads, writes, runs commands, and searches code in a loop until the work is done, using native tool-calling across OpenA…
14 views 0 likesA secure persistent personal agent server in Rust. One binary, sandboxed execution, multi-provider LLMs, voice, memory, Telegram, WhatsApp, Discord, Teams, and…
14 views 0 likesSelf-evolving agent: grows skill tree from 3.3K-line seed, achieving full system control with 6x less token consumption
14 views 0 likesDeepSeek Harness Desktop (dsh-desktop). EAC: Embracing All Creation (揽尽万象). Bundled Node.js runtime with full dsh-CLI kernel, one-click startup, 10 built-in UI…
14 views 0 likesSee your agent think. Zero-config observability & governance for 26 AI agent runtimes: Claude Code, Cursor, OpenAI Codex, GitHub Copilot, Gemini CLI, Cline, Ope…
13 views 0 likesSave 94% on AI coding tokens. Index your codebase, agents search instead of reading files. Works with Claude Code, Codex, Copilot, Cursor, Gemini CLI. Local MCP…
14 views 0 likesYet another coding agent harness, lightweight and written in go.
14 views 0 likesa coding Agent from pi. ∞ providers, sub-agents, hashline edits, and a permission gate
13 views 0 likesOmnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting…
25 views 0 likes🧠 Leon is your open-source personal assistant.
14 views 0 likesThe Station, an open-world multi-agent environment that models a miniature scientific ecosystem.
14 views 0 likesThe Frontend Stack for Agents & Generative UI. React, Angular, Mobile, Slack, and more. Makers of the AG-UI Protocol
23 views 0 likesVelaTerm = iTerm2 + Codex, The Best Terminal for AI Coding
21 views 0 likes