Claude
Skill
ai-skill-builder
Guides creation, audit, and improvement of portable Agent Skills using the shared
Virus-scanned
Reviewed automatically before listing.
Download
ahundt-autorun-plugins_autorun_skills_ai-skill-builder-6fb6027.zip · 98 KB
Install
skills CLI
npx skills add https://github.com/ahundt/autorun/tree/main/plugins/autorun/skills/ai-skill-builder
Claude Code
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ahundt-autorun@llmmart
Git
git clone https://github.com/ahundt/autorun.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ahundt/autorun collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
AI Skill Builder
Files (autorun)
-
agents
-
openai.yaml 203 B
interface: display_name: "AI Skill Builder" short_description: "Create, audit, and improve portable Agent Skills" default_prompt: "Use $ai-skill-builder to create or audit a portable agent skill."
-
-
references
-
examples
-
SKILL-template.md 4 KB
--- name: your-skill-name-here description: This skill should be used when the user wants to "trigger phrase 1", "trigger phrase 2", or needs guidance on [domain]. Brief capability summary. metadata: version: 0.1.0 --- # Skill Name <purpose> [LEVEL 1: THE HOOK - 50-100 words] Brief introduction that includes: - Clear value proposition (what outcome does this achieve?) - Who it's for (target users) - When to use it (triggering scenarios) To start: [first action to take] **Invoke with:** `/your-skill-name` or ask about [trigger scenario] </purpose> <workflow> ## How It Works [LEVEL 2: THE WORKFLOW - 200-400 words] Brief overview of the process in 3-5 clear steps: ### Step 1: [Phase Name] - What happens in this step - What inputs are needed - What outputs are produced - Definition of Done: [what must be true before Step 2 starts] ### Step 2: [Phase Name] - What happens in this step - What inputs are needed - What outputs are produced - Definition of Done: [what must be true before Step 3 starts] ### Step 3: [Phase Name] - What happens in this step - What inputs are needed - What outputs are produced - Definition of Done: [what must be true before the skill reports success] **Definition of Done for the whole workflow**: [the checkable end state] --- ## Detailed Workflow [LEVEL 3: COMPREHENSIVE DOCUMENTATION - No word limit] ### Prerequisites - Required tools/dependencies - Required knowledge/skills - Required access/permissions ### Complete Step-by-Step Guide #### Step 1: [Detailed Phase Name] **Purpose**: [Why this step matters] **Process**: 1. [Detailed sub-step 1] 2. [Detailed sub-step 2] 3. [Detailed sub-step 3] **Inputs**: - Input 1: [Description, format, example] - Input 2: [Description, format, example] **Outputs**: - Output 1: [Description, format, example] - Output 2: [Description, format, example] **Common Issues**: - Issue 1: [Problem and solution] - Issue 2: [Problem and solution] #### Step 2: [Continue for all steps...] ### Configuration Options **Option 1: [Name]** - Description: [What it does] - Default: [Default value] - Valid values: [Acceptable inputs] - Example: `[Example usage]` ### Error Handling **Error 1: [Error name/message]** - Cause: [Why this happens] - Solution: [How to fix it] - Prevention: [How to avoid it] ### Advanced Usage [Optional advanced features, edge cases, power user tips] </workflow> <examples> ## Examples ### Example 1: [Common Use Case] **Scenario**: [Describe the situation] **Input**: ``` [Actual input example] ``` **Process**: 1. [What happens step by step] 2. [With actual values/outputs] **Output**: ``` [Actual output example] ``` **Result**: [Outcome achieved] ### Example 2: [Edge Case] [Repeat structure] </examples> <success_criteria> ## Success Metrics ### Quantitative - [Metric 1]: [Baseline] → [With skill] ([Improvement %]) - [Metric 2]: [Baseline] → [With skill] ([Improvement %]) ### Qualitative - [Quality aspect 1]: [How it improves] - [Quality aspect 2]: [How it improves] </success_criteria> <troubleshooting> ## Troubleshooting ### Common Issues **Issue**: [Problem description] - **Symptoms**: [How you know this is happening] - **Cause**: [Root cause] - **Solution**: [Step-by-step fix] - **Prevention**: [How to avoid in future] </troubleshooting> <resources> ## Resources ### Documentation - [Link to relevant docs] - [Link to API references] ### Examples - [Link to example projects] - [Link to sample outputs] ### Community - [Link to support channels] - [Link to issue tracker] --- ## Changelog Record the version in `metadata.version` above and the release history in `references/changelog.md`. Readers using the skill do not need it; readers changing it do. --- ## License & Attribution [License information if applicable] [Attribution to sources, tools, or inspirations] </resources> <!-- Every major section above sits inside a balanced, descriptive XML region (purpose, workflow, examples, success_criteria, troubleshooting, resources). Rename regions to fit the task; a `## ` heading outside every region fails scripts/audit-skill.sh. -->
-
-
ai-skill-builder-guide.md 35.6 KB
# AI Skill Builder Guide — Extracted from Anthropic PDF **Source**: The Complete Guide to Building Skills for Claude (January 2026) --- ## Page 1 The Complete Guide to Building Skills for Claude --- ## Page 2 Contents Introduction 3 Fundamentals 4 Planning and design 7 Testing and iteration 14 Distribution and sharing 18 Patterns and troubleshooting 21 Resources and references 28 2 --- ## Page 3 Introduction A skill is a set of instructions - packaged as a simple folder - that teaches Claude Two Paths Through This Guide how to handle specific tasks or workflows. Skills are one of the most powerful Building standalone skills? Focus on Fundamentals, Planning and Design, and ways to customize Claude for your specific needs. Instead of re-explaining your category 1-2. Enhancing an MCP integration? The "Skills + MCP" section and preferences, processes, and domain expertise in every conversation, skills let you category 3 are for you. Both paths share the same technical requirements, but teach Claude once and benefit every time. you choose what's relevant to your use case. Skills are powerful when you have repeatable workflows: generating frontend What you'll get out of this guide: By the end, you'll be able to build a functional designs from specs, conducting research with consistent methodology, creating skill in a single sitting. Expect about 15-30 minutes to build and test your first documents that follow your team's style guide, or orchestrating multi-step working skill using the skill-creator. processes. They work well with Claude's built-in capabilities like code execution Let's get started. and document creation. For those building MCP integrations, skills add another powerful layer helping turn raw tool access into reliable, optimized workflows. This guide covers everything you need to know to build effective skills - from planning and structure to testing and distribution. Whether you're building a skill for yourself, your team, or for the community, you'll find practical patterns and real-world examples throughout. What you'll learn: • Technical requirements and best practices for skill structure • Patterns for standalone skills and MCP-enhanced workflows • Patterns we've seen work well across different use cases • How to test, iterate, and distribute your skills Who this is for: • Developers who want Claude to follow specific workflows consistently 3 • Power users who want Claude to follow specific workflows • Teams looking to standardize how Claude works across their organization --- ## Page 4 Chapter 1 Fundamentals 4 --- ## Page 5 Chapter 1 Fundamentals What is a skill? Composability A skill is a folder containing: Claude can load multiple skills simultaneously. Your skill should work well alongside others, not assume it's the only capability available. • SKILL.md (required): Instructions in Markdown with YAML frontmatter • scripts/ (optional): Executable code (Python, Bash, etc.) Portability • references/ (optional): Documentation loaded as needed Skills work identically across Claude.ai, Claude Code, and API. Create a skill once • assets/ (optional): Templates, fonts, icons used in output and it works across all surfaces without modification, provided the environment supports any dependencies the skill requires. Core design principles For MCP Builders: Skills + Connectors Progressive Disclosure 💡 Building standalone skills without MCP? Skip to Planning and Design - you can Skills use a three-level system: always return here later. • First level (YAML frontmatter): Always loaded in Claude's system prompt. If you already have a working MCP server, you've done the hard part. Skills are Provides just enough information for Claude to know when each skill should the knowledge layer on top - capturing the workflows and best practices you be used without loading all of it into context. already know, so Claude can apply them consistently. • Second level (SKILL.md body): Loaded when Claude thinks the skill is relevant to the current task. Contains the full instructions and guidance. The kitchen analogy • Third level (Linked files): Additional files bundled within the skill directory that Claude can choose to navigate and discover only as needed. MCP provides the professional kitchen: access to tools, ingredients, and equipment. This progressive disclosure minimizes token usage while maintaining specialized expertise. Skills provide the recipes: step-by-step instructions on how to create something valuable. 5 --- ## Page 6 Together, they enable users to accomplish complex tasks without needing to figure out every step themselves. How they work together: MCP (Connectivity) Skills (Knowledge) Connects Claude to your service Teaches Claude how to use your service (Notion, Asana, Linear, etc.) effectively Provides real-time data access and tool Captures workflows and best practices invocation What Claude can do How Claude should do it Why this matters for your MCP users Without skills: • Users connect your MCP but don't know what to do next • Support tickets asking "how do I do X with your integration" • Each conversation starts from scratch • Inconsistent results because users prompt differently each time • Users blame your connector when the real issue is workflow guidance With skills: • Pre-built workflows activate automatically when needed • Consistent, reliable tool usage • Best practices embedded in every interaction • Lower learning curve for your integration 6 --- ## Page 7 Chapter 2 Planning and design 7 --- ## Page 8 Chapter 2 Planning and design Start with use cases Common skill use case categories Before writing any code, identify 2-3 concrete use cases your skill should enable. At Anthropic, we’ve observed three common use cases: Good use case definition: Category 1: Document & Asset Creation Used for: Creating consistent, high-quality output including documents, Use Case: Project Sprint Planning presentations, apps, designs, code, etc. Trigger: User says "help me plan this sprint" or "create sprint tasks" Real example: frontend-design skill (also see skills for docx, pptx, xlsx, and Steps: ppt) 1. Fetch current project status from Linear (via MCP) "Create distinctive, production-grade frontend interfaces with high design 2. Analyze team velocity and capacity 3. Suggest task prioritization quality. Use when building web components, pages, artifacts, posters, or 4. Create tasks in Linear with proper labels and estimates applications." Result: Fully planned sprint with tasks created Key techniques: • Embedded style guides and brand standards Ask yourself: • Template structures for consistent output • What does a user want to accomplish? • Quality checklists before finalizing • What multi-step workflows does this require? • No external tools required - uses Claude's built-in capabilities • Which tools are needed (built-in or MCP?) • What domain knowledge or best practices should be embedded? 8 --- ## Page 9 Category 2: Workflow Automation Define success criteria Used for: Multi-step processes that benefit from consistent methodology, How will you know your skill is working? including coordination across multiple MCP servers. These are aspirational targets - rough benchmarks rather than precise Real example: skill-creator skill thresholds. Aim for rigor but accept that there will be an element of vibes-based assessment. We are actively developing more robust measurement guidance and "Interactive guide for creating new skills. Walks the user through use case tooling. definition, frontmatter generation, instruction writing, and validation." Quantitative metrics: Key techniques: • Skill triggers on 90% of relevant queries • Step-by-step workflow with validation gates – How to measure: Run 10-20 test queries that should trigger your skill. Track • Templates for common structures how many times it loads automatically vs. requires explicit invocation. • Built-in review and improvement suggestions • Completes workflow in X tool calls • Iterative refinement loops – How to measure: Compare the same task with and without the skill enabled. Count tool calls and total tokens consumed. Category 3: MCP Enhancement • 0 failed API calls per workflow – How to measure: Monitor MCP server logs during test runs. Track retry rates Used for: Workflow guidance to enhance the tool access an MCP server provides. and error codes. Real example: sentry-code-review skill (from Sentry) Qualitative metrics: "Automatically analyzes and fixes detected bugs in GitHub Pull Requests using • Users don't need to prompt Claude about next steps Sentry's error monitoring data via their MCP server." – How to assess: During testing, note how often you need to redirect or clarify. Ask beta users for feedback. Key techniques: • Coordinates multiple MCP calls in sequence • Workflows complete without user correction • Embeds domain expertise – How to assess: Run the same request 3-5 times. Compare outputs for structural consistency and quality. • Provides context users would otherwise need to specify • Consistent results across sessions • Error handling for common MCP issues – How to assess: Can a new user accomplish the task on first try with minimal guidance? 9 --- ## Page 10 Technical requirements YAML frontmatter: The most important part The YAML frontmatter is how Claude decides whether to load your skill. Get this File structure right. your-skill-name/ Minimal required format ├── SKILL.md # Required - main skill file ├── scripts/ # Optional - executable code --- │ ├── process_data.py # Example name: your-skill-name │ └── validate.sh # Example description: What it does. Use when user asks to [specific ├── references/ # Optional - documentation phrases]. │ ├── api-guide.md # Example --- │ └── examples/ # Example └── assets/ # Optional - templates, etc. └── report-template.md # Example That's all you need to start. Field requirements Critical rules name (required): SKILL.md naming: • kebab-case only • Must be exactly SKILL.md (case-sensitive) • No spaces or capitals • No variations accepted (SKILL.MD, skill.md, etc.) • Should match folder name Skill folder naming: description (required): • Use kebab-case: notion-project-setup ✅ • MUST include BOTH: • No spaces: Notion Project Setup ❌ – What the skill does • No underscores: notion_project_setup ❌ – When to use it (trigger conditions) • No capitals: NotionProjectSetup ❌ • Under 1024 characters • No XML tags (< or >) No README.md: • Include specific tasks users might say • Don't include README.md inside your skill folder • Mention file types if relevant • All documentation goes in SKILL.md or references/ • Note: when distributing via GitHub, you'll still want a repo-level README for human users — see Distribution and Sharing. 10 --- ## Page 11 license (optional): Writing effective skills • Use if making skill open source The description field • Common: MIT, Apache-2.0 According to Anthropic's engineering blog: "This metadata...provides just compatibility (optional) enough information for Claude to know when each skill should be used without • 1-500 characters loading all of it into context." This is the first level of progressive disclosure. • Indicates environment requirements: e.g. intended product, required system packages, network access needs, etc. Structure: metadata (optional): [What it does] + [When to use it] + [Key capabilities] • Any custom key-value pairs • Suggested: author, version, mcp-server • Example: Examples of good descriptions: ```yaml metadata: author: ProjectHub # Good - specific and actionable version: 1.0.0 mcp-server: projecthub description: Analyzes Figma design files and generates ``` developer handoff documentation. Use when user uploads .fig files, asks for "design specs", "component documentation", or Security restrictions "design-to-code handoff". Forbidden in frontmatter: # Good - includes trigger phrases description: Manages Linear project workflows including sprint • XML angle brackets (< >) planning, task creation, and status tracking. Use when user • Skills with "claude" or "anthropic" in name (reserved) mentions "sprint", "Linear tasks", "project planning", or asks to "create tickets". Why: Frontmatter appears in Claude's system prompt. Malicious content could inject instructions. # Good - clear value proposition description: End-to-end customer onboarding workflow for PayFlow. Handles account creation, payment setup, and subscription management. Use when user says "onboard new customer", "set up subscription", or "create PayFlow account". 11 --- ## Page 12 Examples of bad descriptions: Example: # Too vague python scripts/fetch_data.py --project-id PROJECT_ID description: Helps with projects. Expected output: [describe what success looks like] # Missing triggers description: Creates sophisticated multi-page documentation (Add more steps as needed) systems. # Too technical, no user triggers Examples description: Implements the Project entity model with hierarchical relationships. Example 1: [common scenario] User says: "Set up a new marketing campaign" Writing the main instructions Actions: After the frontmatter, write the actual instructions in Markdown. 1. Fetch existing campaigns via MCP 2. Create new campaign with provided parameters Recommended structure: Result: Campaign created with confirmation link Adapt this template for your skill. Replace bracketed sections with your specific content. (Add more examples as needed) --- Troubleshooting name: your-skill description: [--.] Error: [Common error message] --- Cause: [Why it happens] # Your Skill Name Solution: [How to fix] -# Instructions (Add more error cases as needed) --# Step 1: [First Major Step] Clear explanation of what happens. 12 --- ## Page 13 Best Practices for Instructions Reference bundled resources clearly Be Specific and Actionable Before writing queries, consult `references/api-patterns.md` for: ✅ Good: - Rate limiting guidance - Pagination patterns Run `python scripts/validate.py --input {filename}` to check - Error codes and handling data format. If validation fails, common issues include: - Missing required fields (add them to the CSV) Use progressive disclosure - Invalid date formats (use YYYY-MM-DD) Keep SKILL.md focused on core instructions. Move detailed documentation to `references/` and link to it. (See Core Design Principles for how the three- ❌ Bad: level system works.) Validate the data before proceeding. Include error handling -# Common Issues --# MCP Connection Failed If you see "Connection refused": 1. Verify MCP server is running: Check Settings > Extensions 2. Confirm API key is valid 3. Try reconnecting: Settings > Extensions > [Your Service] > Reconnect 13 --- ## Page 14 Chapter 3 Testing and iteration 14 --- ## Page 15 Chapter 3 Testing and iteration Skills can be tested at varying levels of rigor depending on your needs: Recommended Testing Approach • Manual testing in Claude.ai - Run queries directly and observe behavior. Fast Based on early experience, effective skills testing typically covers three areas: iteration, no setup required. • Scripted testing in Claude Code - Automate test cases for repeatable 1. Triggering tests validation across changes. Goal: Ensure your skill loads at the right times. • Programmatic testing via skills API - Build evaluation suites that run systematically against defined test sets. Test cases: Choose the approach that matches your quality requirements and the visibility • ✅ Triggers on obvious tasks of your skill. A skill used internally by a small team has different testing needs • ✅ Triggers on paraphrased requests than one deployed to thousands of enterprise users. • ❌ Doesn't trigger on unrelated topics Example test suite: Pro Tip: Iterate on a single task before expanding Should trigger: We’ve found that the most effective skill creators iterate on a single challenging - "Help me set up a new ProjectHub workspace" task until Claude succeeds, then extract the winning approach into a skill. This - "I need to create a project in ProjectHub" leverages Claude’s in-context learning and provides faster signal than broad - "Initialize a ProjectHub project for Q4 planning" testing. Once you have a working foundation, expand to multiple test cases for Should NOT trigger: coverage. - "What's the weather in San Francisco?" - "Help me write Python code" - "Create a spreadsheet" (unless ProjectHub skill handles sheets) 15 --- ## Page 16 2. Functional tests With skill: Goal: Verify the skill produces correct outputs. - Automatic workflow execution - 2 clarifying questions only Test cases: - 0 failed API calls • Valid outputs generated - 6,000 tokens consumed • API calls succeed • Error handling works Using the skill-creator skill • Edge cases covered The skill-creator skill - available in Claude.ai via plugin directory or Example: download for Claude Code - can help you build and iterate on skills. If you have an MCP server and know your top 2–3 workflows, you can build and test a Test: Create project with 5 tasks functional skill in a single sitting - often in 15–30 minutes. Given: Project name "Q4 Planning", 5 task descriptions When: Skill executes workflow Creating skills: Then: • Generate skills from natural language descriptions - Project created in ProjectHub - 5 tasks created with correct properties • Produce properly formatted SKILL.md with frontmatter - All tasks linked to project • Suggest trigger phrases and structure - No API errors Reviewing skills: • Flag common issues (vague descriptions, missing triggers, structural 3. Performance comparison problems) • Identify potential over/under-triggering risks Goal: Prove the skill improves results vs. baseline. • Suggest test cases based on the skill's stated purpose Use the metrics from Define Success Criteria. Here's what a comparison might look like. Iterative improvement: • After using your skill and encountering edge cases or failures, bring those Baseline comparison: examples back to skill-creator • Example: "Use the issues & solution identified in this chat to improve how the Without skill: skill handles [specific edge case]" - User provides instructions each time - 15 back-and-forth messages - 3 failed API calls requiring retry - 12,000 tokens consumed 16 --- ## Page 17 To use: Execution issues: • Inconsistent results "Use the skill-creator skill to help me build a skill for • API call failures [your use case]" • User corrections needed Note: skill-creator helps you design and refine skills but does not execute Solution: Improve instructions, add error handling automated test suites or produce quantitative evaluation results. Iteration based on feedback Skills are living documents. Plan to iterate based on: Undertriggering signals: • Skill doesn't load when it should • Users manually enabling it • Support questions about when to use it Solution: Add more detail and nuance to the description - this may include keywords particularly for technical terms Overtriggering signals: • Skill loads for irrelevant queries • Users disabling it • Confusion about purpose Solution: Add negative triggers, be more specific 17 --- ## Page 18 Chapter 4 Distribution and sharing 18 --- ## Page 19 Chapter 4 Distribution and sharing Skills make your MCP integration more complete. As users compare connectors, Using skills via API those with skills offer a faster path to value, giving you an edge over MCP-only alternatives. For programmatic use cases - such as building applications, agents, or automated workflows that leverage skills - the API provides direct control over skill management and execution. Current distribution model (January 2026) Key capabilities: How individual users get skills: • `/v1/skills` endpoint for listing and managing skills 1. Download the skill folder • Add skills to Messages API requests via the `container.skills` parameter 2. Zip the folder (if needed) • Version control and management through the Claude Console 3. Upload to Claude.ai via Settings > Capabilities > Skills • Works with the Claude Agent SDK for building custom agents 4. Or place in Claude Code skills directory When to use skills via the API vs. Claude.ai: Organization-level skills: • Admins can deploy skills workspace-wide (shipped December 18, 2025) Use Case Best Surface • Automatic updates End users interacting with skills directly Claude.ai / Claude Code • Centralized management Manual testing and iteration during development Claude.ai / Claude Code An open standard Individual, ad-hoc workflows Claude.ai / Claude Code We've published Agent Skills as an open standard. Like MCP, we believe skills should be portable across tools and platforms - the same skill should work Applications using skills programmatically API whether you're using Claude or other AI platforms. That said, some skills are designed to take full advantage of a specific platform's capabilities; authors can Production deployments at scale API note this in the skill's compatibility field. We've been collaborating with members of the ecosystem on the standard, and we're excited by early adoption. Automated pipelines and agent systems API 19 --- ## Page 20 Note: Skills in the API require the Code Execution Tool beta, which provides the secure environment skills need to run. - Select the skill folder (zipped) For implementation details, see: 3. Enable the skill: - Toggle on the [Your Service] skill • Skills API Quickstart - Ensure your MCP server is connected • Create Custom skills 4. Test: • Skills in the Agent SDK - Ask Claude: "Set up a new project in [Your Service]" Recommended approach today Positioning your skill Start by hosting your skill on GitHub with a public repo, clear README (for How you describe your skill determines whether users understand its value and human visitors — this is separate from your skill folder, which should not contain actually try it. When writing about your skill—in your README, documentation, a README.md), and example usage with screenshots. Then add a section or marketing - keep these principles in mind. to your MCP documentation that links to the skill, explains why using both together is valuable, and provides a quick-start guide. Focus on outcomes, not features: 1. Host on GitHub ✅ Good: – Public repo for open-source skills – Clear README with installation instructions "The ProjectHub skill enables teams to set up complete project – Example usage and screenshots workspaces in seconds — including pages, databases, and templates — instead of spending 30 minutes on manual setup." 2. Document in Your MCP Repo – Link to skills from MCP documentation – Explain the value of using both together ❌ Bad: – Provide quick-start guide 3. Create an Installation Guide "The ProjectHub skill is a folder containing YAML frontmatter and Markdown instructions that calls our MCP server tools." -# Installing the [Your Service] skill 1. Download the skill: Highlight the MCP + skills story: - Clone repo: `git clone https:-/github.com/yourcompany/ skills` - Or download ZIP from Releases "Our MCP server gives Claude access to your Linear projects. Our skills teach Claude your team's sprint planning workflow. 2. Install in Claude: Together, they enable AI-powered project management." - Open Claude.ai > Settings > skills - Click "Upload skill" 20 --- ## Page 21 Chapter 5 Patterns and troubleshooting 21 --- ## Page 22 Chapter 5 Patterns and troubleshooting These patterns emerged from skills created by early adopters and internal teams. Pattern 1: Sequential workflow orchestration They represent common approaches we've seen work well, not prescriptive templates. Use when: Your users need multi-step processes in a specific order. Example structure: Choosing your approach: Problem-first vs. tool-first Think of it like Home Depot. You might walk in with a problem - "I need to fix a -# Workflow: Onboard New Customer kitchen cabinet" - and an employee points you to the right tools. Or you might pick out a new drill and ask how to use it for your specific job. --# Step 1: Create Account Call MCP tool: `create_customer` Skills work the same way: Parameters: name, email, company • Problem-first: "I need to set up a project workspace" → Your skill orchestrates --# Step 2: Setup Payment the right MCP calls in the right sequence. Users describe outcomes; the skill Call MCP tool: `setup_payment_method` handles the tools. Wait for: payment method verification • Tool-first: "I have Notion MCP connected" → Your skill teaches Claude the optimal workflows and best practices. Users have access; the skill provides --# Step 3: Create Subscription expertise. Call MCP tool: `create_subscription` Parameters: plan_id, customer_id (from Step 1) Most skills lean one direction. Knowing which framing fits your use case helps you choose the right pattern below. --# Step 4: Send Welcome Email Call MCP tool: `send_email` Template: welcome_email_template Key techniques: • Explicit step ordering • Dependencies between steps • Validation at each stage • Rollback instructions for failures 22 --- ## Page 23 Pattern 2: Multi-MCP coordination Pattern 3: Iterative refinement Use when: Workflows span multiple services. Use when: Output quality improves with iteration. Example: Design-to-development handoff Example: Report generation --# Phase 1: Design Export (Figma MCP) -# Iterative Report Creation 1. Export design assets from Figma 2. Generate design specifications --# Initial Draft 3. Create asset manifest 1. Fetch data via MCP 2. Generate first draft report --# Phase 2: Asset Storage (Drive MCP) 3. Save to temporary file 1. Create project folder in Drive 2. Upload all assets --# Quality Check 3. Generate shareable links 1. Run validation script: `scripts/check_report.py` 2. Identify issues: --# Phase 3: Task Creation (Linear MCP) - Missing sections 1. Create development tasks - Inconsistent formatting 2. Attach asset links to tasks - Data validation errors 3. Assign to engineering team --# Refinement Loop --# Phase 4: Notification (Slack MCP) 1. Address each identified issue 1. Post handoff summary to #engineering 2. Regenerate affected sections 2. Include asset links and task references 3. Re-validate 4. Repeat until quality threshold met Key techniques: --# Finalization 1. Apply final formatting • Clear phase separation 2. Generate summary • Data passing between MCPs 3. Save final version • Validation before moving to next phase • Centralized error handling Key techniques: • Explicit quality criteria • Iterative improvement • Validation scripts • Know when to stop iterating 23 --- ## Page 24 Pattern 4: Context-aware tool selection Pattern 5: Domain-specific intelligence Use when: Same outcome, different tools depending on context. Use when: Your skill adds specialized knowledge beyond tool access. Example: File storage Example: Financial compliance -# Smart File Storage -# Payment Processing with Compliance --# Decision Tree --# Before Processing (Compliance Check) 1. Check file type and size 1. Fetch transaction details via MCP 2. Determine best storage location: 2. Apply compliance rules: - Large files (>10MB): Use cloud storage MCP - Check sanctions lists - Collaborative docs: Use Notion/Docs MCP - Verify jurisdiction allowances - Code files: Use GitHub MCP - Assess risk level - Temporary files: Use local storage 3. Document compliance decision --# Execute Storage --# Processing Based on decision: IF compliance passed: - Call appropriate MCP tool - Call payment processing MCP tool - Apply service-specific metadata - Apply appropriate fraud checks - Generate access link - Process transaction ELSE: --# Provide Context to User - Flag for review Explain why that storage was chosen - Create compliance case --# Audit Trail Key techniques: - Log all compliance checks - Record processing decisions • Clear decision criteria - Generate audit report • Fallback options • Transparency about choices Key techniques: • Domain expertise embedded in logic • Compliance before action • Comprehensive documentation • Clear governance 24 --- ## Page 25 Troubleshooting # Wrong name: My Cool Skill Skill won't upload Error: "Could not find SKILL.md in uploaded folder" # Correct name: my-cool-skill Cause: File not named exactly SKILL.md Solution: Skill doesn't trigger • Rename to SKILL.md (case-sensitive) • Verify with: ls -la should show SKILL.md Symptom: Skill never loads automatically Error: "Invalid frontmatter" Fix: Cause: YAML formatting issue Revise your description field. See The Description Field for good/bad examples. Common mistakes: Quick checklist: • Is it too generic? ("Helps with projects" won't work) # Wrong - missing delimiters • Does it include trigger phrases users would actually say? name: my-skill • Does it mention relevant file types if applicable? description: Does things Debugging approach: # Wrong - unclosed quotes name: my-skill Ask Claude: "When would you use the [skill name] skill?" Claude will quote the description: "Does things description back. Adjust based on what's missing. # Correct Skill triggers too often --- name: my-skill Symptom: Skill loads for unrelated queries description: Does things --- Solutions: 1. Add negative triggers Error: "Invalid skill name" description: Advanced data analysis for CSV files. Use for Cause: Name has spaces or capitals statistical modeling, regression, clustering. Do NOT use for simple data exploration (use data-viz skill instead). 25 --- ## Page 26 2. Be more specific Instructions not followed Symptom: Skill loads but Claude doesn't follow instructions # Too broad description: Processes documents Common causes: 1. Instructions too verbose # More specific description: Processes PDF legal documents for contract review – Keep instructions concise – Use bullet points and numbered lists – Move detailed reference to separate files 3. Clarify scope 2. Instructions buried – Put critical instructions at the top description: PayFlow payment processing for e-commerce. Use – Use ## Important or ## Critical headers specifically for online payment workflows, not for general – Repeat key points if needed financial queries. 3. Ambiguous language # Bad MCP connection issues Make sure to validate things properly Symptom: Skill loads but MCP calls fail # Good CRITICAL: Before calling create_project, verify: Checklist: - Project name is non-empty 1. Verify MCP server is connected - At least one team member assigned – Claude.ai: Settings > Extensions > [Your Service] - Start date is not in the past – Should show "Connected" status Advanced technique: For critical validations, consider bundling a script 2. Check authentication – API keys valid and not expired that performs the checks programmatically rather than relying on language – Proper permissions/scopes granted instructions. Code is deterministic; language interpretation isn't. See the Office – OAuth tokens refreshed skills for examples of this pattern. 4. Model "laziness" Add explicit encouragement: 3. Test MCP independently – Ask Claude to call MCP directly (without skill) – "Use [Service] MCP to fetch my projects" -# Performance Notes – If this fails, issue is MCP not skill - Take your time to do this thoroughly - Quality is more important than speed 4. Verify tool names - Do not skip validation steps – Skill references correct MCP tool names – Check MCP server documentation Note: Adding this to user prompts is more effective than in SKILL.md – Tool names are case-sensitive 26 --- ## Page 27 Large context issues Symptom: Skill seems slow or responses degraded Causes: • Skill content too large • Too many skills enabled simultaneously • All content loaded instead of progressive disclosure Solutions: 1. Optimize SKILL.md size – Move detailed docs to references/ – Link to references instead of inline – Keep SKILL.md under 5,000 words 2. Reduce enabled skills – Evaluate if you have more than 20 - 50 skills enabled simultaneously – Recommend selective enablement – Consider skill "packs" for related capabilities 27 --- ## Page 28 Chapter 6 Resources and references 28 --- ## Page 29 Chapter 6 Resources and references If you're building your first skill, start with the Best Practices Guide, then Tools and Utilities reference the API docs as needed. skill-creator skill: Official Documentation • Built into Claude.ai and available for Claude Code • Can generate skills from descriptions Anthropic Resources: • Reviews and provides recommendations • Best Practices Guide • Use: "Help me build a skill using skill-creator" • Skills Documentation • API Reference Validation: • MCP Documentation • skill-creator can assess your skills • Ask: "Review this skill and suggest improvements" Blog Posts: • Introducing Agent Skills Getting Support • Engineering Blog: Equipping Agents for the Real World For Technical Questions: • Skills Explained • General questions: Community forums at the Claude Developers Discord • How to Create Skills for Claude • Building Skills for Claude Code For Bug Reports: • Improving Frontend Design through Skills • GitHub Issues: anthropics/skills/issues • Include: Skill name, error message, steps to reproduce Example skills Public skills repository: • GitHub: anthropics/skills • Contains Anthropic-created skills you can customize 29 --- ## Page 30 Before upload Reference A: Quick Tested triggering on obvious tasks checklist Tested triggering on paraphrased requests Verified doesn't trigger on unrelated topics Use this checklist to validate your skill before and after upload. If you want Functional tests pass a faster start, use the skill-creator skill to generate your first draft, then run Tool integration works (if applicable) through this list to make sure you haven't missed anything. Compressed as .zip file Before you start After upload Identified 2-3 concrete use cases Test in real conversations Tools identified (built-in or MCP) Monitor for under/over-triggering Reviewed this guide and example skills Collect user feedback Planned folder structure Iterate on description and instructions During development Update version in metadata Folder named in kebab-case SKILL.md file exists (exact spelling) YAML frontmatter has --- delimiters name field: kebab-case, no spaces, no capitals description includes WHAT and WHEN No XML tags (< >) anywhere Instructions are clear and actionable Error handling included Examples provided References clearly linked 30 --- ## Page 31 Security notes Reference B: YAML frontmatter Allowed: • Any standard YAML types (strings, numbers, booleans, lists, objects) • Custom metadata fields Required fields • Long descriptions (up to 1024 characters) --- Forbidden: name: skill-name-in-kebab-case description: What it does and when to use it. Include specific • XML angle brackets (< >) - security restriction trigger phrases. • Code execution in YAML (uses safe YAML parsing) --- • Skills named with "claude" or "anthropic" prefix (reserved) All optional fields name: skill-name description: [required description] license: MIT # Optional: License for open-source allowed-tools: "Bash(python:*) Bash(npm:*) WebFetch" # Optional: Restrict tool access metadata: # Optional: Custom fields author: Company Name version: 1.0.0 mcp-server: server-name category: productivity tags: [project-management, automation] documentation: https:-/example.com/docs support: support@example.com 31 --- ## Page 32 Reference C: Complete skill examples For full, production-ready skills demonstrating the patterns in this guide: • Document Skills - PDF, DOCX, PPTX, XLSX creation • Example Skills - Various workflow patterns • Partner Skills Directory - View skills from various partners such as Asana, Atlassian, Canva, Figma, Sentry, Zapier, and more These repositories stay up-to-date and include additional examples beyond what's covered here. Clone them, modify them for your use case, and use them as templates. 32 --- ## Page 33 claude.ai -
best-practices.md 14.1 KB
# Claude Skill Best Practices Quick reference guide for building effective Claude skills based on Anthropic's official methodology. ## File Naming Rules (CRITICAL) ### ✅ Correct - **File name**: `SKILL.md` (exactly, case-sensitive) - **Folder name**: `api-test-generator` (kebab-case) - **Location**: `~/.claude/skills/api-test-generator/SKILL.md` ### ❌ Incorrect - ❌ `README.md` (Claude doesn't read this) - ❌ `skill.md` (wrong case) - ❌ `SKILL.MD` (wrong extension case on some systems) - ❌ `api_test_generator` (underscores) - ❌ `apiTestGenerator` (camelCase) - ❌ `API-Test-Generator` (capitals in folder) ## YAML Frontmatter Requirements ### Minimal Required ```yaml --- name: skill-name description: This skill should be used when the user wants to "trigger phrase 1", "trigger phrase 2", or needs help with [domain]. --- ``` ### Extended — every optional field ```yaml --- name: your-skill-name description: What it does and when. Use when user says "trigger phrase 1", "trigger phrase 2". license: MIT # Optional: open-source license (MIT, Apache-2.0, etc.) allowed-tools: "Bash(python:*) WebFetch" # Optional: restrict which tools the skill can use compatibility: Claude Code # Optional: 1-500 chars; environment requirements metadata: # Optional: custom key-value pairs author: Your Name version: 1.0.0 # Version belongs inside metadata, never at the top level mcp-server: your-server # If skill requires a specific MCP server category: productivity tags: [automation, workflow] dependencies: tool-name,mcp-server-name documentation: https://example.com/docs --- ``` `license`, `compatibility`, and `allowed-tools` are optional in the portable specification and accepted differently by different hosts. `allowed-tools` is experimental and its accepted shape varies. Name the host and version you tested rather than assuming portability. ### Description Field Constraints - **Under 1024 characters** (hard limit — longer descriptions are truncated) - **No XML angle brackets** (`<` or `>`) in any frontmatter field — a host restriction from Anthropic's skill guide and the Claude Code docs, not a YAML rule; the body must carry balanced semantic XML regions (SKILL-REQ004 in `SKILL.md`) - **Skill name must be kebab-case** (e.g., `my-skill`) — no spaces, no capitals ## Description Writing Formula ### Two Audiences, Two Formats The `description` field in YAML frontmatter and a GitHub README.md serve different audiences and require different language: | Location | Audience | Purpose | Language style | |----------|----------|---------|---------------| | `description` field in SKILL.md | Claude (AI) | Auto-activation: pattern-match user queries | Trigger phrases | | `README.md` at repo root | Humans | Installation decision: "should I install this?" | Outcome-focused | ### ❌ Wrong for description field — outcome-focused (misses trigger matching) ```yaml description: Generate API test suites 87% faster than manual writing ``` ### ✅ Correct for description field — trigger-phrase format ```yaml # Format A (plugin-dev style): description: This skill should be used when the user wants to "generate API tests", "create a test suite", "write tests for my endpoints", or needs help with API test generation. # Format B (Anthropic PDF style — capability + triggers): description: Generates API test suites from OpenAPI specs. Use when user asks for "generate API tests", "create a test suite from my spec", or "automate endpoint testing". ``` ### ✅ Correct for README.md — outcome-focused (for human readers) ```markdown ## Why Use This Skill? Generate production-ready API tests 87% faster than writing them manually. ``` ### ❌ Wrong for description field — feature-focused ```yaml description: Uses OpenAPI parser with Jinja2 templates to generate Jest tests ``` **Rule**: `description` field → what users SAY → trigger phrases. GitHub README → what users ACHIEVE → outcome-focused. ## Progressive Disclosure Levels ### Level 1: The Hook (50-100 words) **Purpose**: Quick decision - "Is this for me?" **Include**: - Clear value proposition - Target user - Triggering scenario - Exact trigger phrase **Omit**: - Technical details - How it works internally - Configuration options - Edge cases ### Level 2: The Workflow (200-400 words) **Purpose**: Understanding - "How does this work?" **Include**: - 3-5 numbered steps - Numbered items wherever order matters or a reader must cite one; bullets only for unordered sets - Input → Output per step - A Definition of Done per step: the checkable condition that must hold before the next one starts. Not a wall-clock estimate — how long a step takes depends on who or what runs it. **Omit**: - Implementation details - Error handling - Advanced configuration - Troubleshooting ### Level 3: Comprehensive (No limit) **Purpose**: Reference - "How do I handle X?" **Include**: - Complete technical details - All configuration options - Error messages and solutions - Edge cases and examples - Advanced usage patterns **Structure**: 1. Prerequisites 2. Detailed steps 3. Configuration 4. Error handling 5. Examples 6. Troubleshooting ## Skill Categories ### Category 1: Document & Asset Creation **Pattern**: Input → Analysis → Generation → Output **Examples**: - Generate documentation from code - Create test suites from specs - Build diagrams from descriptions - Generate reports from data **Structure**: ``` Input Requirements → Analysis Phase → Generation Phase → Validation → Output ``` ### Category 2: Workflow Automation **Pattern**: Task → Orchestration → Execution → Validation **Examples**: - Deploy to production - Run data pipelines - Execute health checks - Coordinate multi-step processes **Structure**: ``` Pre-flight Checks → Sequential/Parallel Steps → Error Handling → Status Report ``` ### Category 3: MCP Enhancement **Pattern**: MCP Tools → Composition → Intelligence Layer → Enhanced Output **Examples**: - Combine database + API tools - Add semantic search over filesystem - Cache slow MCP operations - Create composite MCP operations **Structure**: ``` MCP Tool Discovery → Tool Composition → Add AI Layer → Return Results ``` ## Folder Structure Patterns ### Minimal (Document Creation) ``` skill-name/ └── SKILL.md ``` ### Standard (With Scripts) ``` skill-name/ ├── SKILL.md └── scripts/ ├── generate.py └── validate.sh ``` ### Complete (Full Featured) ``` skill-name/ ├── SKILL.md ├── scripts/ │ ├── deploy.sh │ └── rollback.sh ├── references/ │ ├── api-docs.md │ └── examples.md └── assets/ ├── config.json └── template.yaml ``` ### Annotated (Claude Code standalone example) ``` ~/.claude/skills/your-skill-name/ ├── SKILL.md # Required — loaded when skill triggers (<5k words) ├── references/ # Docs Claude loads into context as needed │ ├── detailed-guide.md # schemas, API docs, policies, detailed workflows │ └── examples/ # working code users copy (subdirectory of references/) │ └── working-example.sh ├── scripts/ # Executables (run without loading into context) │ └── validate.sh └── assets/ # Files used IN skill output (not loaded to context) └── template.html ``` ## Testing Framework ### 1. Triggering Tests **Purpose**: Verify Claude detects the skill correctly ```markdown Test 1: Exact trigger Input: "/skill-name" Expected: Skill activates Test 2: Natural language Input: "Help me [task description]" Expected: Skill activates Test 3: Similar but wrong Input: "[Related but different task]" Expected: Skill does NOT activate ``` ### 2. Functional Tests **Purpose**: Verify skill works correctly ```markdown Test 1: Happy path Input: [Standard valid input] Expected Output: [Correct result] Success Criteria: [Measurable outcome] Test 2: Edge case Input: [Minimal/maximal/unusual input] Expected Output: [Handled gracefully] Success Criteria: [No errors, sensible result] Test 3: Error case Input: [Invalid input] Expected Output: [Clear error message] Success Criteria: [Helpful guidance provided] ``` ### 3. Performance Tests **Purpose**: Verify skill provides value ```markdown Metric 1: Time Baseline: [Manual time] With Skill: [Automated time] Improvement: [Percentage reduction] Metric 2: Quality Baseline: [Manual quality metric] With Skill: [Automated quality metric] Improvement: [Improvement description] Metric 3: Consistency Baseline: [Variation in manual process] With Skill: [Standardization achieved] Improvement: [Consistency improvement] ``` ## Common Antipatterns ### Antipattern 1: The Wall of Text **Problem**: Everything in one giant block **Solution**: Use progressive disclosure ``` Level 1 (Hook) → Level 2 (Workflow) → Level 3 (Details) ``` ### Antipattern 2: The Feature List **Problem**: Describing what it has, not what it achieves **Solution**: Name the outcome — in the repo README and in the SKILL.md body. The `description` field is the one place this does not apply: there, outcome language displaces the trigger phrases a host matches on. See "Two Audiences, Two Formats" above. ``` ❌ "Has integration with 5 APIs" ✅ "Sync data across 5 platforms automatically" ``` ### Antipattern 3: The Assumption Trap **Problem**: Assuming user knows context **Solution**: State prerequisites explicitly ``` Prerequisites: - Docker installed - AWS credentials configured - Node.js 18+ ``` ### Antipattern 4: The Mystery Box **Problem**: No examples of actual usage **Solution**: Include concrete examples ``` Example Input: [Actual input] Example Output: [Actual output] Result: [Outcome achieved] ``` ### Antipattern 5: The Untested Skill **Problem**: Publishing without validation **Solution**: Test before releasing, in the order SKILL.md's Step 4 gives ``` 1. Triggering tests 2. Functional tests 3. Performance tests 4. Compatibility tests (every named host and version) 5. Forward tests (realistic positive, negative, ambiguous, and adversarial prompts) ``` Then collect user feedback. ### Antipattern 6: The Unmeasured Rewrite **Problem**: A revision adds material and no task outcome improves **Solution**: Run the same tasks against the version being replaced; keep the revision only if it does them better ``` ❌ "The audit passes and it reads better now" ✅ "Same 5 tasks, 4 correct before, 5 correct after" ``` ## Success Criteria Patterns The strings below are shapes to fill from your own measurements, not results anyone recorded. A number earns its place once it names four things: the measured event, the workload and conditions, the number with its unit, and which direction counts as better. Record the manual baseline before building, so the improvement has something to be relative to — `references/testing.md` covers the measurement framework. ### Quantitative Metrics - **Time Reduction**: "75% faster than manual process" - **Error Reduction**: "90% fewer deployment failures" - **Cost Savings**: "Save $5K/month in manual work" - **Scale Improvement**: "Handle 10x more requests" - **Quality Increase**: "85% test coverage vs 60% manual" ### Qualitative Metrics - **Consistency**: "Standardized across 12 teams" - **Best Practices**: "Follows industry standards automatically" - **Accessibility**: "Non-experts can use effectively" - **Maintainability**: "Reduced code complexity by 40%" - **Reliability**: "Zero-downtime deployments" ## Distribution Checklist ### GitHub Repository - [ ] Clear README with installation steps - [ ] LICENSE file (MIT recommended) - [ ] Example use cases documented - [ ] Screenshots/GIFs if applicable - [ ] CHANGELOG for version tracking ### Documentation - [ ] Installation guide tested on clean system - [ ] Prerequisites clearly listed - [ ] Common issues documented - [ ] Example usage included - [ ] Support channel identified ### Community - [ ] Announcement in Claude Discord - [ ] Post in relevant forums/communities - [ ] Blog post or tutorial (optional) - [ ] Response plan for issues/questions ### Maintenance - [ ] Issue tracking enabled - [ ] Update plan defined - [ ] Support commitment stated - [ ] Deprecation path considered ## Quick Reference Commands ### Create New Skill ```bash mkdir -p ~/.claude/skills/my-new-skill cd ~/.claude/skills/my-new-skill touch SKILL.md ``` ### Validate Structure ```bash # Check file exists ls ~/.claude/skills/my-skill/SKILL.md # Check frontmatter head -5 ~/.claude/skills/my-skill/SKILL.md ``` ### Test Skill Discovery ```bash # Restart Claude Code # Then try trigger phrase /my-skill ``` ## Word Count Targets ### Skill Sections - **Level 1 Hook**: 50-100 words - **Level 2 Workflow**: 200-400 words - **Level 3 Details**: No limit (comprehensive) ### Description Field - **YAML description**: 1-3 sentences, long enough to carry 4-8 quoted trigger phrases - Focus: what the skill does, then the phrases users say. Use Format A or B above. - Avoid: outcome-only positioning ("87% faster"). That belongs in the repo README. - Hard limit 1024 characters. `scripts/audit-skill.sh` fails a description with no quoted phrases. ### Step Descriptions - **Step title**: 3-5 words - **Step description**: 20-50 words - **Definition of Done**: one checkable condition ## Version Control Tips ### Semantic Versioning - **v1.0.0**: Initial release - **v1.1.0**: New features (backward compatible) - **v1.0.1**: Bug fixes - **v2.0.0**: Breaking changes ### Changelog Format ```markdown ## v1.1.0 - 2026-02-15 ### Added - New configuration option for custom templates - Support for Python 3.12 ### Fixed - Error handling for missing dependencies - Typo in step 3 instructions ### Changed - Improved performance by 25% ``` ## Resources ### Official - [Claude Code skills documentation](https://code.claude.com/docs/en/skills) - [MCP Protocol](https://modelcontextprotocol.io) - [Agent Skills specification](https://agentskills.io/specification) ### Community - [Anthropic on GitHub](https://github.com/anthropics) ### Tools - Claude Code CLI - MCP Inspector - skill-creator (built-in) -
categories.md 9.7 KB
# Skill Categories — In-Depth Guide From Anthropic's "The Complete Guide to Building Skills for Claude" (January 2026), pages 8-9. Three official skill categories determine the architecture and workflow of any skill. Choose the category that best matches the primary input-to-output transformation. --- ## Category 1: Document & Asset Creation **Core pattern**: Structured INPUT → Analysis → Generation → Structured OUTPUT ### When to Use Use Category 1 when the skill takes data, specifications, or content as input and produces a document, code file, diagram, or report as output. The key characteristic: the output is a durable artifact the user keeps or deploys. ### Characteristics - **Input**: Specs, requirements, data, raw content, templates - **Process**: Parse input → extract structure → apply templates → validate → format - **Output**: Document, code, diagram, report, test suite, deployment manifest - **Time profile**: Predictable; scales with input size ### Examples | Use Case | Input | Output | |----------|-------|--------| | API test generator | OpenAPI spec | Jest/Pytest test suite | | Meeting notes summarizer | Audio transcript | Structured meeting notes + action items | | PR description generator | Git diff | Pull request description | | API documentation builder | Source code | Markdown documentation | | Deployment manifest creator | Service config | Kubernetes YAML | | Report generator | CSV data | Formatted PDF/Markdown report | ### Recommended Structure ``` my-doc-skill/ ├── SKILL.md # Workflow overview ├── references/ │ ├── output-templates.md # Template formats for outputs │ ├── examples/ │ │ └── example-output.md # Real example of expected output │ └── validation-rules.md # What makes output valid └── scripts/ └── validate-output.sh # Structural validation of output ``` ### Best Practices 1. **Define input format precisely**: List required fields, accepted file types, example inputs 2. **Provide output examples**: Show a complete, working example of the expected output 3. **Include validation steps**: How to verify the generated artifact is correct 4. **Support multiple output formats**: e.g., Jest vs Pytest, Markdown vs PDF 5. **Handle missing or partial inputs gracefully**: What to skip vs what to require ### SKILL.md Level 2 Template (Workflow) ```markdown ## How It Works Generate [output type] from [input type] in 3 steps: ### Step 1: Parse Input - Read and validate the [input type] - Extract [key fields] from the structure - Identify [special conditions or options] ### Step 2: Generate [Output] - Create [output structure] from extracted data - Apply [templates or patterns] to each [element] - Add [supporting elements] (auth, error handling, etc.) ### Step 3: Validate and Format - Check output for [correctness criteria] - Format to [output spec] - Output to [location] **Definition of Done**: [the specific, checkable end state for this skill] ``` --- ## Category 2: Workflow Automation **Core pattern**: Task Parameters → Orchestration → Sequential/Parallel Execution → Status Report ### When to Use Use Category 2 when the skill coordinates multiple steps, tools, or services to complete a multi-phase process. The key characteristic: the skill manages state across steps and handles failures at each stage. ### Characteristics - **Input**: Task parameters, configuration, context about what to do - **Process**: Pre-flight checks → execute phases → verify results → handle errors → report - **Output**: Completed workflow + status report (what succeeded, what failed, next steps) - **Time profile**: Variable; depends on external services and error recovery ### Examples | Use Case | Input | Output | |----------|-------|--------| | Deploy to production | App name + version | Deployment confirmation + health check | | Data migration pipeline | Source + destination config | Migration report with row counts | | Code review workflow | PR number | Review comments + approval decision | | Multi-service health check | Service list | Health dashboard + alert summary | | Database schema migration | Migration file | Applied changes + rollback script | | Release automation | Release config | Tagged release + changelog + notification | ### Recommended Structure ``` my-workflow-skill/ ├── SKILL.md # Workflow phases + error handling ├── references/ │ ├── phases.md # Detailed description of each phase │ ├── error-recovery.md # What to do when each phase fails │ └── examples/ │ └── example-run.md # Sample successful + failed run └── scripts/ ├── pre-flight-check.sh # Validate preconditions ├── execute-phase-N.sh # Phase execution scripts └── rollback.sh # Undo changes if workflow fails ``` ### Best Practices 1. **Define pre-flight checks explicitly**: What must be true before starting (credentials, config, etc.) 2. **Make each phase atomic**: Either fully succeeds or leaves system unchanged 3. **Add progress indicators**: Users need to know what's happening during long operations 4. **Always provide rollback capability**: Every destructive step needs an undo 5. **Structure error messages with next actions**: Not "failed" but "failed at step 3, run rollback.sh" ### SKILL.md Level 2 Template (Workflow) ```markdown ## How It Works [Describe task] in [N] phases: ### Phase 1: Pre-flight Checks - Verify [credentials/config/dependencies] - Check [required state] is correct - Confirm [safety conditions] before proceeding ### Phase 2: [Primary Action] - [Main operation step 1] - [Main operation step 2] - Verify [intermediate result] ### Phase 3: Validate and Report - Check [success criteria] - Generate status report - If failure: provide rollback instructions **On failure**: Run `scripts/rollback.sh` to restore previous state **Definition of Done**: [the specific, checkable end state. Not a wall-clock estimate.] ``` --- ## Category 3: MCP Enhancement **Core pattern**: MCP Tool Outputs → Composition Layer → AI Orchestration → Enhanced Results ### When to Use Use Category 3 when the skill adds intelligence or coordination on top of existing MCP server capabilities. The key characteristic: the skill makes MCP tools smarter by combining them, adding caching, or applying domain knowledge that the raw tools lack. ### Characteristics - **Input**: MCP tool configurations, user queries, tool outputs - **Process**: Discover available tools → compose operations → apply intelligence → cache results - **Output**: Richer results than any single MCP tool could produce alone - **Time profile**: Varies; caching dramatically improves repeat queries ### Examples | Use Case | Input | Output | |----------|-------|--------| | Smart file search | User query | Relevant files with semantic context | | BigQuery assistant | Natural language question | SQL query + results + interpretation | | Multi-database sync | Sync config | Sync report with conflict resolution | | Composite API orchestrator | Business operation | Coordinated API calls + unified result | | Caching layer for slow tools | Tool + query | Cached result with freshness indicator | | Schema-aware query builder | Table name + intent | Validated SQL with schema checks | ### Recommended Structure ``` my-mcp-skill/ ├── SKILL.md # MCP dependencies + orchestration logic ├── references/ │ ├── mcp-integration.md # How to configure required MCP servers │ ├── schema.md # Domain schema (e.g., database tables) │ └── examples/ │ └── example-queries.md # Real queries and expected results ``` ### Best Practices 1. **Document MCP dependencies explicitly**: List every MCP server the skill requires with setup instructions 2. **Handle MCP tool failures gracefully**: What to do if a required tool is unavailable 3. **Cache expensive operations**: Use filesystem or memory to avoid redundant MCP calls 4. **Document which tools are optional vs required**: Skill should degrade gracefully if optional tools are absent 5. **Test without real MCP servers**: Provide mock data in references/examples/ for development ### SKILL.md Level 2 Template (Workflow) ```markdown ## How It Works Enhance [MCP capability] with [intelligence layer]: ### Required MCP Servers - **[server-name]**: [What it provides] — Install: `[install command]` - **[server-name]**: [What it provides] — Install: `[install command]` ### Step 1: Discover and Configure (auto) - Detect available MCP tools from [server list] - Load [schema/config] from `references/schema.md` - Initialize cache if available ### Step 2: Orchestrate Query (1-5 sec) - Decompose user query into [tool-specific operations] - Execute [tool A] for [data type A] - Combine with [tool B] for [data type B] ### Step 3: Apply Intelligence and Return (1-2 sec) - Apply [domain knowledge] to interpret results - Format response with [context and explanations] - Cache result with [TTL] for repeat queries ``` --- ## Choosing Between Categories | Question | If YES → | |----------|----------| | Does output live in a file/repo after the skill runs? | Category 1 | | Does the skill execute real actions (deploy, migrate, delete)? | Category 2 | | Does the skill primarily coordinate or enhance MCP tools? | Category 3 | | Does the skill call external APIs or services? | Usually Category 2 | | Does the skill generate code the user will use? | Category 1 | | Does the skill need rollback capabilities? | Category 2 | **When in doubt**: Category 1 if the output is a document, Category 2 if the output is a completed action. --- ## Source Anthropic, "The Complete Guide to Building Skills for Claude," January 2026, pages 8-9. -
changelog.md 28.5 KB
# ai-skill-builder changelog The revision history of the ai-skill-builder package itself. Nothing here describes the skill you are building. Release history first, then the detail behind each release, newest first. ## Release history ### v1.3.0 - 2026-08-15 - Absorbed the `engineer-agent-skills` package: its thirteen P0 requirements (SKILL-REQ001–013) now sit in a `<requirements>` region of `SKILL.md`, its compatibility matrix, claim audit, and validation receipt are `references/portability-and-claim-audit.md`, and its `agents/openai.yaml` shape is shipped here. Listed the package as superseded so nobody installs both. - Made semantic XML body regions mandatory (SKILL-REQ004): every major section inside a balanced, descriptive tag on its own line, Markdown inside, code in fences. `SKILL.md`, `scripts/scaffold-skill.sh` output, and `references/examples/SKILL-template.md` now follow it. - `scripts/audit-skill.sh`: new section 4 fails a body with no region, an unbalanced or mis-nested tag, an HTML presentational tag used as a region, or a `## ` heading outside every region; warns about prose outside every region; ignores frontmatter and fenced code. Added reader item 8 for what it cannot decide (whether a tag name describes its content). Later sections renumbered 5–8. Corrected the frontmatter angle-bracket FAIL to cite the host restriction rather than "forbidden in YAML", which was false. - Step 1 classifies authoring, execution, and evaluation; Step 4 adds host validators, per-host validation receipts, and forward tests; Step 5 states installer ownership and the delivery report. Distribution channels moved verbatim to `references/distribution.md`; the annotated folder tree moved verbatim to `references/best-practices.md`. - `references/sources.md`: added GitHub Copilot, Microsoft Agent Framework, and the Codex `quick_validate.py` accepted-key list; recorded the discrepancy between the Anthropic guide's "no XML tags anywhere" checklist line and its frontmatter-only Reference B. - Description gains "create an agent skill", "audit a SKILL.md", "test skill discovery", "install skills across harnesses" and names semantic XML regions, frontmatter, scripts, arguments, and security among the topics it covers. ### v1.2.1 - 2026-08-13 - Added Antipattern 6, The Unmeasured Rewrite, to `references/best-practices.md`: run the same tasks against the version being replaced and keep the revision only if it does them better. A passing audit had been read as a quality verdict, which the audit itself denies. - Stated the numbering rule in the same file: numbered items where order matters or a reader must cite one, bullets only for unordered sets. Only "3-5 numbered steps" had been written. ### v1.2.0 - 2026-08-03 - Audited every file against every other file and against `scripts/audit-skill.sh`; fixed the contradictions found. Detail below. - Split `audit-skill.sh` output into decided, proxy, and needs-a-reader checks, and scored only the decided ones. - Replaced wall-clock phase estimates with a Definition of Done throughout. ### v1.1.0 - 2026-03-05 - Added YAML frontmatter (enables Claude auto-detection) - Rewrote body to imperative form throughout - Integrated IMPROVEMENTS.md self-critique (moved to this file) - Corrected description templates: trigger-phrase format for SKILL.md, outcome-focused for README - Trimmed SKILL.md from 932 → ~350 lines; moved detail to references/ - Corrected directory taxonomy: references/examples/ for standalone skills (Anthropic's recommended structure) - Corrected README.md rule: enforce "no README inside skill folder" with distribution exception - Added description field constraints: 1024-char hard limit, no angle brackets, kebab-case names - Added Additional Resources section linking all references/ files and scripts/ - Created references/patterns.md with 5 advanced patterns + troubleshooting from Anthropic's guide - Confirmed skill categories are from Anthropic's official guide — no "unofficial" label ### v1.0.0 - Initial release based on Anthropic's Complete Guide - Complete 4-phase methodology - Progressive disclosure templates - Testing framework - Distribution strategies - Common pitfalls guide ## TODO — empirical regression guard for over-engineered skills - Investigate the deep-superset regression: its `SKILL.md` grew from 167 lines / 1,129 words at `a3e571c` to 715 lines / 4,991 words at `01cc1d9`; structural audits still passed while cold reads found executable and sequencing defects, and the worktree was rolled back to the pre-cluster behavior. - Thread the needle: retain the smallest coherent portable core and only add a guard when it maps to a reproduced failure. Validate the real task empirically with failing fixtures, focused scripts, state/exit-status checks, cold reads, and longer-term dogfood/recurrence evidence; a clean structural audit is not a behavior verdict. - Add a refinement checklist and regression fixture that reject complexity without measured outcome gain, preserve the last known-good baseline, and make rollback/reassessment explicit. ## v1.2.0 detail — internal-contradiction audit Every file in the package was read in full and checked against the other files and against `scripts/audit-skill.sh`. **All line numbers below are as of v1.1.0, before these edits.** They locate each defect in the version that carried it; `git show` the v1.1.0 tree to follow them. This edit shifted every file, so the same numbers do not point at the same text in v1.2.0. ### Guidance that contradicted the validator or the rest of the package 1. Outcome-focused `description` prescribed at 8 sites, which `audit-skill.sh` fails for having no quoted trigger phrases, and which `SKILL.md:238`, `best-practices.md:62`, and `troubleshooting.md:51` already call wrong: `best-practices.md:405`, `refining-skills.md:127`, `:137`, `:240`, `:260`, `:282`, `:565`, and `changelog.md:352`. `refining-skills.md:137` offered as its ✅ result the exact string `best-practices.md:64` labels ❌. 2. Top-level `version` field prescribed at 3 sites, which `audit-skill.sh` fails: `troubleshooting.md:18`, `distribution.md:40`, `:45`. Four other sites already said `metadata.version`. ### Broken pointers and commands 3. A `SKILL.md` bullet pointed at a file that is not part of the package. The pointer was removed; the package cites only files it ships or public sources. 4. `refining-skills.md:494` pointed at `../templates/SKILL-template.md`. Corrected to `examples/SKILL-template.md`. 5. `discovery.md:55` pointed at `references/skill-categories.md`. Corrected to `categories.md`. 6. `troubleshooting.md:288` called `open('~/.claude/skills/…')`. Python does not expand `~`, so the diagnostic raised `FileNotFoundError` on every run. Rewritten to take the path as an argument. ### Numbers without a source 7. `testing.md:204` and `:225` required "at least two of ≥50% time reduction, ≥10 percentage points quality improvement…". No primary source establishes cross-skill thresholds. Replaced with criteria relative to the skill's own measured baseline. 8. Invented before/after figures removed from `refining-skills.md:224`, `:365`, `:369`, `:582`, `:627` and from the ROI, Projected Impact, Impact Measurements, and Success Metrics sections of this file. The v1.1 file inventory is now labelled a dated snapshot. ### Claims, links, counts 9. `distribution.md:15` claimed "any Claude Code installation can use any skill". Discovery roots, `allowed-tools` handling, plugin versus standalone placement, and reload behavior differ by host. Narrowed to tested hosts and versions. 10. `best-practices.md:442` cited `agent-skills.dev` for the standard while `SKILL.md:22` and `sources.md:18` cite `agentskills.io`. Unified. Bare `docs.anthropic.com` replaced with the Claude Code skills page; a dead Discord invite removed. 11. Word limits appeared as 2,000, 3,000, and 5,000 in different files. Reconciled to the 2,000-word guideline and 5,000-word hard limit, with the specification's 5,000-token and 500-line recommendation recorded in `sources.md`. 12. `refining-skills.md:113` instructed `rm`. Replaced with `trash`, which preserves the deletion record. 13. `discovery.md` and `SKILL.md` described a "20-question" guide containing 22 questions. ### Conformance of ai-skill-builder to the rules it teaches 14. Added `metadata.version`, required by `SKILL.md:101` and absent from ai-skill-builder's own frontmatter since v1.1. 15. Dropped `Glob` and `Grep` from `allowed-tools`. `Glob` appeared nowhere in the package except that line; `Grep` only inside a hypothetical example in `patterns.md`. 16. Brought `SKILL.md` back under the 2,000-word guideline by moving the optional-frontmatter field list, the trigger-fix detail, and the GitHub setup steps into the references that already covered them. 17. Rewrote the reference index so each entry states when to read the file and what it produces, rather than summarising its contents. ### Added 18. Target-host discovery in Step 1: host, version, discovery root, invocation form, and reload behavior, with unknowns marked as unknown. 19. Compatibility tests as a fourth test type in Step 4. 20. A Common Pitfalls row for unverified compatibility and performance claims. 21. A pointer to the `engineer-agent-skills` skill for cross-host claim matrices, per-host validation receipts, and install ownership, rather than duplicating that material here. ### Second pass Found by re-reading the whole package after the edits above, so these are defects the first pass introduced or walked past. 22. `scripts/audit-skill.sh` still repeated an unpublished local measurement claim after item 3 removed the package pointer. The useful checks stayed; the unsupported claim was removed from distributable references. 23. Four markdown files mis-rendered because a same-length code fence was nested inside another: `troubleshooting.md` State Tracking, `distribution.md` README template, and a stray unclosed fence in the extracted guide. Everything after the break rendered as the wrong kind of content. `audit-skill.sh` now checks fence balance and same-length nesting. 24. `references/examples/SKILL-template.md` — the file every new skill is copied from — still carried `(X minutes)` per step and a `**Total Time**` line, so item 6's Definition of Done policy would have been undone by the first skill built from the template. 25. Two checks classified as PASS were heading and regex matches: the Level 3 detail heading and the description's technical identifiers. Both are now proxies. 26. The `SKILL.md` reference index used bare filenames (`research.md`), which the link checker does not resolve and which leave the directory to be inferred. Restored to `references/` paths. 27. `sources.md` cited Format A and Format B by line number into `best-practices.md`; both numbers were already stale. Replaced with the section name. 28. `scripts/scaffold-skill.sh` generated a skill that failed `scripts/audit-skill.sh`. Its frontmatter carried an outcome-focused TODO description with no quoted trigger phrases (a FAIL) and no `metadata.version` beside the `## Version History` heading it also generated (a second FAIL). It emitted `(X minutes)` per step, a `**Total Time**` line, and an `[X%] faster` metric — everything items 6 and 7 removed elsewhere. Running the scaffolder and auditing the result is now the check that keeps the two scripts agreeing; a fresh scaffold reports zero FAILs and two warnings for its own unfilled TODOs. 29. `scaffold-skill.sh` used `${SKILL_NAME^}` for the generated H1. That form needs bash 4; macOS ships bash 3.2.57 as `/bin/bash`, where it is `bad substitution` and the script exits 1. It ran here only because Homebrew bash 5.3 precedes `/bin/bash` on PATH. Rewritten with `tr` and `awk`, and verified with `/bin/bash` directly. 30. `scaffold-skill.sh` ran `rm -rf` on an existing skill directory after a single y/N prompt, in a package that tells authors to use `trash` and whose audit fails a skill for showing `rm` on a real path. Now uses `trash`, or a timestamped `mv` when `trash` is absent. 31. `scaffold-skill.sh` wrote a `README.md` into `scripts/`, `references/`, and `assets/`, in a package whose stated rule is that a skill folder carries no `README.md`, and where every file under `references/` is loadable context. The one in `references/` cost tokens to say "add documentation here". The guidance now prints to the terminal. 32. `scaffold-skill.sh` hardcoded `~/.claude/skills` as the only output root, and pointed at the same dead `templates/SKILL-template.md` path item 4 corrected elsewhere. `SKILLS_DIR` is now overridable, matching the target-host discovery step item 18 added. 33. Two Common Pitfalls rows contradicted the row item 20 added, from two lines away. The ✅ column for missing success criteria read "Reduce API test writing time by 75%", a bare percentage with no baseline; and the ✅ column for no testing still listed three test types after item 19 made it four. Both corrected. 34. `SKILL.md` described `audit-skill.sh` as producing "scored output (0-100%)" after item 25 made the number a decided-check score. It now says what the number covers and that the gate is zero FAILs. 35. `SKILL.md`'s opening claimed Antigravity support. `sources.md` cites host documentation for Claude Code, Codex, and Qwen Code and nothing for Antigravity, so the sentence was an instance of the pitfall item 20 added three sections below it. It now names the three sourced hosts and tells the reader to verify any other. ## v1.1.0 development notes (written 2026-02-15, released 2026-03-05) A critique of v1.0 and a record of what v1.1 changed in response. Kept for provenance. Where it disagrees with the guidance in the package today, the package is current and this section is not. --- ## Critical Analysis of Initial Version (v1.0) ### ❌ Major Gaps Identified **1. No Refinement Pathway** (CRITICAL GAP) - **Problem**: Only covered creating NEW skills, ignored improving existing ones - **Impact**: Users with old skills had no migration path - **Real-world scenario**: Someone with `my_api_test/README.md` has no way to upgrade - **Severity**: High - Most users refine more than they create **2. Missing Source Links** (DOCUMENTATION GAP) - **Problem**: No attribution to Anthropic's guide, no reference links - **Impact**: Users can't verify methodology or dive deeper - **Missing links**: PDF guide, Anthropic docs, MCP, community resources - **Severity**: Medium - Reduces credibility and learning ability **3. No Validation Tools** (AUTOMATION GAP) - **Problem**: No way to audit if skills follow best practices - **Impact**: Users don't know if their skills are correct - **Missing**: Automated checker for file structure, naming, frontmatter - **Severity**: High - Manual validation is error-prone **4. Inappropriate Tool Usage** (IMPLEMENTATION ERROR) - **Problem**: Examples used `nano`, `vim` (human text editors) - **Impact**: Claude can't use these - uses Read/Edit/Write tools instead - **Context**: Skills are for Claude to use, not humans - **Severity**: Medium - Confusing and technically incorrect **5. No Migration Guide** (USABILITY GAP) - **Problem**: No path from old skill standards to new ones - **Impact**: Existing skill authors stuck with old patterns - **Missing**: Before/after examples, step-by-step upgrade process - **Severity**: Medium - Blocks adoption of new standards **6. Limited Real Examples** (LEARNING GAP) - **Problem**: Hypothetical examples only, no real skill references - **Impact**: Users can't see actual working implementations - **Missing**: Links to community skills, real-world patterns - **Severity**: Low - Learning is slower but possible **7. No Performance Tracking** (MEASUREMENT GAP) - **Problem**: No framework for measuring refinement impact - **Impact**: Can't validate improvements worked - **Missing**: Before/after metrics, success measurement - **Severity**: Medium - Can't prove value of changes --- ## Improvements Implemented (v1.1) ### ✅ Major Additions **1. Complete Refinement Workflow** - **File**: `references/refining-skills.md` (1,902 words) - **Content**: - 5-step refinement process (Audit → Prioritize → Fix → Validate → Document) - Common scenarios (migration, enhancement, automation) - Real before/after examples with impact metrics - Tool usage guide (Read/Edit/Write, not nano) - Continuous improvement framework - **Impact**: Fills critical gap in Anthropic's guide **2. Automated Audit Script** - **File**: `scripts/audit-skill.sh` (executable) - **Capabilities**: - File structure validation (SKILL.md, kebab-case) - YAML frontmatter checking (name, description) - Progressive disclosure detection (3 levels) - Content quality analysis (examples, metrics) - Scoring system (0-100%) - Actionable fix recommendations - **Impact**: automates the structural checks that were being done by eye **3. Comprehensive Source Documentation** - **File**: `references/sources.md` (1,115 words) - **Content**: - Primary source: Anthropic's PDF with full citation - Official docs: Claude, MCP, Agent Skills Standard - Community resources: Discord, GitHub - Related tools and technologies - Learning resources (prompt engineering, markdown) - Testing tools (ShellCheck, markdownlint) - Progressive disclosure theory (Nielsen Norman Group) - Outcome-focused design (Jobs to Be Done) - Complete URL reference table - Citation formats (APA, Chicago) - **Impact**: Full transparency and verifiability **4. Fixed Tool Usage Throughout** - **Changed**: All examples from `nano/vim` → Claude tools - **Examples now show**: - `Read:` for examining files - `Edit:` for precise changes - `Write:` for creating files - `Bash:` for file operations - **Impact**: Technically accurate for Claude's use **5. Enhanced Main Documentation** - **File**: `SKILL.md` updated (3,281 words) - **Additions**: - Major "Refining Existing Skills" section - Source attribution at top - Links to refinement guide - Tool usage examples (Claude-appropriate) - Migration scenarios - Performance tracking examples - **Impact**: Now covers full lifecycle (create + refine) **6. Improved README** - **File**: `README.md` updated (1,398 words) - **Additions**: - Refinement capabilities highlighted - Source links section - Audit script documentation - Example 3: Refining existing skill - Fixed tool usage in examples - Version history with improvements listed - **Impact**: Clear discovery of new capabilities --- ## Files Created/Updated Summary ### New Files (v1.1) 1. `scripts/audit-skill.sh` - Automated validation 2. `references/refining-skills.md` - Complete refinement guide 3. `references/sources.md` - All source links 4. `IMPROVEMENTS.md` - This document ### Updated Files (v1.1) 1. `SKILL.md` - Added refinement section, source links 2. `README.md` - Added refinement capabilities, sources (later removed; see v1.2.0) 3. `references/examples/SKILL-template.md` - Already good (no changes needed) 4. `references/best-practices.md` - Already comprehensive (no changes needed) 5. `scripts/scaffold-skill.sh` - Already functional (no changes needed) ### Total documentation as of v1.1 (2026-03-05) - **Word count**: 9,510 words - **Files**: 8 (3 markdown docs, 3 reference docs, 2 scripts) - **Coverage**: Create + Refine + Sources + Tools This inventory is a v1.1 snapshot and is not maintained. For the current file list run `bash scripts/audit-skill.sh <skill>`. --- ## Gap Analysis: What Was Missing ### From Anthropic's Guide **Guide Covered**: - ✅ 4-phase creation methodology - ✅ Progressive disclosure structure - ✅ Skill categories - ✅ Testing framework - ✅ Distribution strategies **Guide Missed**: - ❌ Refining existing skills - ❌ Migration from old standards - ❌ Automated validation tools - ❌ Performance tracking - ❌ Continuous improvement ### Our Implementation **Now Includes**: - ✅ All content from guide - ✅ Refinement workflow (original) - ✅ Audit automation (original) - ✅ Migration guides (original) - ✅ Source attribution (original) - ✅ Performance tracking (original) **Total Coverage**: Guide + 6 major additions --- ## What changed between v1.0 and v1.1 ### Countable | Thing counted | v1.0 | v1.1 | |---|---|---| | Files in the package | 5 | 8 | | Scripts | 1 (scaffold) | 2 (scaffold, audit) | | Reference documents | 1 | 3 | Percentage deltas were reported here through v1.1. They divided counts of different things by each other and are gone; the counts themselves are above, as a dated v1.1 snapshot. ### Qualitative Improvements **Coverage**: - v1.0: Creation workflow only (~50% of skill lifecycle) - v1.1: Complete lifecycle (create + refine + validate) **Usability**: - v1.0: Manual validation required - v1.1: Automated audit with scoring **Accuracy**: - v1.0: Mixed tool usage (nano/vim inappropriate for Claude) - v1.1: All examples use Claude tools (Read/Edit/Write) **Verifiability**: - v1.0: No source links - v1.1: Complete source attribution with URLs --- ## Lessons Learned ### What Worked Well Initially 1. **Progressive disclosure structure** - Correctly implemented from guide 2. **Scaffolding automation** - Good use of bash scripting 3. **Template provision** - Helpful starting point 4. **File naming rules** - Comprehensive and accurate ### What Needed Improvement 1. **Lifecycle coverage** - Too focused on creation, ignored refinement 2. **Source attribution** - No links to verify methodology 3. **Automation** - Manual validation is error-prone 4. **Tool usage** - Confused human and AI tool usage 5. **Real examples** - Hypothetical only, no real references ### Design Decisions Made **Decision 1: Separate refinement guide** - **Rationale**: Refinement is complex enough for dedicated doc - **Alternative**: Could have embedded in SKILL.md - **Chose**: Separate file for clarity - **Impact**: Better organization, easier to find **Decision 2: Bash audit script** - **Rationale**: Fast, portable, no dependencies - **Alternative**: Could use Python for richer checks - **Chose**: Bash for simplicity - **Impact**: Works immediately, easy to understand **Decision 3: Comprehensive source documentation** - **Rationale**: Transparency and verifiability - **Alternative**: Could just link to Anthropic guide - **Chose**: Complete source catalog - **Impact**: Users can verify and dive deeper --- ## Self-Critique Summary ### v1.0 **Strengths**: - Accurate methodology from Anthropic guide - Good template structure - Useful scaffolding automation **Weaknesses**: - Covered creating a skill but not improving one - No source attribution - No validation tools - Inappropriate tool examples ### v1.1 **Strengths**: - Complete lifecycle (create + refine) - Full source attribution - Automated validation - Technically accurate tool usage - Comprehensive documentation **Remaining Gaps**: - Could add more real-world skill examples - Could integrate with Claude marketplace (when available) - Could add skill performance analytics - Could add collaborative refinement features --- ## Comparison to Anthropic Guide ### What We Preserved - ✅ 4-phase methodology - ✅ Progressive disclosure (3 levels) - ✅ Skill categories (3 types) - ✅ Testing framework - ✅ Phase estimates (v1.1 kept the guide's wall-clock figures; v1.2.0 replaced them with a Definition of Done per phase) - ✅ Success criteria patterns - ✅ Common pitfalls - ✅ Distribution strategies ### What We Enhanced - ➕ Complete refinement workflow - ➕ Automated audit tooling - ➕ Migration guides - ➕ Source attribution - ➕ Performance tracking - ➕ Tool usage corrections - ➕ Before/after examples - ➕ Continuous improvement framework ### Why Enhancements Were Needed **Anthropic's guide** (excellent for creation): - Target: Creating new skills from scratch - Audience: Developers starting fresh - Scope: Design → Implementation → Testing → Distribution **Real-world needs** (include refinement): - Reality: Most skills need improvement over time - Audience: Developers maintaining existing skills - Scope: Full lifecycle including evolution **Our additions** address the gap between "how to build" and "how to maintain." --- ## Testing This Skill Itself ### Applied Own Methodology **Audit Results**: ```bash bash ./scripts/audit-skill.sh \ ~/.claude/skills/ai-skill-builder ``` **Gate**: zero FAILs. See `references/distribution.md` on why the printed percentage is not a quality verdict. **Checks**: - ✅ SKILL.md exists (correct name) - ✅ kebab-case folder name - ✅ YAML frontmatter present - ✅ Progressive disclosure structure - ✅ Examples included - ✅ Success metrics documented - ✅ No TODOs or placeholders - ✅ Scripts directory with tools - ✅ References directory with guides ### Dogfooding Results **This skill follows its own guidance**: - Progressive disclosure: 3 levels ✅ - Description in capability-plus-quoted-trigger-phrase format: ✅ - Source attribution: ✅ - Examples: ✅ - Testing framework: ✅ - Continuous improvement: ✅ --- ## Recommendations for Future Versions ### v1.2 (Next Minor Release) - Add real-world skill examples (links to quality community skills) - Create skill performance analytics tool - Add collaborative refinement guide (team workflows) - Integrate with emerging Claude marketplace ### v2.0 (Next Major Release) - Interactive web-based skill builder - AI-powered skill suggestion based on use case - Skill dependency management - Automated testing harness - Skill version migration tool --- ## Key Learnings ### On Critique Process 1. **First implementation is never complete** - Critique reveals gaps 2. **User perspective matters** - "Help me improve" is as important as "Help me build" 3. **Source attribution is essential** - Verifiability builds trust 4. **Automation reduces errors** - Manual validation misses issues 5. **Tool accuracy matters** - Claude uses different tools than humans ### On Skill Development 1. **Progressive disclosure works** - Users find what they need quickly 2. **Outcome-focus resonates** - Users care about results, not features 3. **Examples are essential** - Abstract descriptions don't teach 4. **Testing catches issues** - Untested skills create bad experiences 5. **Iteration improves quality** - V1.1 >> V1.0 with focused critique ### On Documentation 1. **Complete source attribution** - Always link to original sources 2. **Separate concerns** - Refinement deserves own guide 3. **Tool-appropriate examples** - Match tool usage to audience (Claude vs humans) 4. **Comprehensive coverage** - Better to be thorough than brief 5. **Self-critique documents** - Show your thinking and improvements --- ## Conclusion ### What We Built A comprehensive skill that: - Teaches Anthropic's official methodology (v1.0) - Adds refinement workflow for existing skills (v1.1 NEW) - Provides automation tools (scaffolding, audit) (v1.1 ENHANCED) - Includes complete source attribution (v1.1 NEW) - Uses technically accurate tool examples (v1.1 FIXED) ### Why It Matters **Before**: Users could create skills but not improve them **After**: Users can create, refine, validate, and continuously improve skills **Before**: No way to know if skills follow best practices **After**: Automated audit with scoring and actionable fixes **Before**: No source links for verification **After**: Complete source catalog with URLs and citations ### What v1.1 added **Lifecycle**: creation only → creation plus refinement **Automation**: 1 script (scaffold) → 2 (scaffold, audit) **Sourcing**: no citations → every methodology claim attributed in `sources.md` **Tool usage in examples**: `nano`/`vim` → Read, Edit, Write, Bash The self-assigned grades and coverage percentages that stood here were not measurements of anything and have been removed. --- ## Sources for This Critique **Methodology**: - Anthropic's Guide: Original best practices - User feedback: "critique your work and improve it" - Challenge Mode (CLAUDE.md): "Never assume, always verify" - Concrete spec (CLAUDE.md): Specific, measurable, testable **Tools Used**: - Read: Examined existing implementation - Edit: Made precise improvements - Write: Created new documentation - Bash: Tested scripts and structure - Critical thinking: Identified gaps and solutions **References**: - Anthropic's Complete Guide to Building Skills for Claude (Jan 2026) - User's CLAUDE.md development guidelines - Real-world skill development experience - Software engineering best practices --- The v1.1 development notes above were last edited 2026-02-15 and are not maintained. The current package version is the `metadata.version` field in `SKILL.md`. -
discovery.md 8.8 KB
# Guided Discovery — Interactive Skill Requirements Q&A Use this 22-question guide when starting a new skill from scratch or when requirements are unclear. Work through each phase in order. Skip a question only when it is clearly not applicable. --- ## Phase 1: Discovery (Questions 1–5) **Purpose**: Understand the problem and who benefits. **1. What problem does this skill solve?** Ask for a concrete description of the pain point. Avoid abstract answers. - ✅ "Writing API test suites takes 2-3 hours per endpoint and our team does 50 new endpoints per sprint" - ❌ "Testing is hard and takes time" Follow up: "What do you do today instead of using this skill?" **2. Who will use it?** Be specific about the user role and technical level. - Developers (what language/stack?) - Designers (what tools?) - Data scientists (what workflow?) - Non-technical users (what context?) Follow up: "What would a typical user already know before using this skill?" **3. What are 2-3 concrete use cases?** Concrete = you could watch someone do it. Gather actual examples. For each use case, capture: - What triggers this task - What inputs the user provides - What outputs they receive - How they use the output **4. What does success look like?** Set measurable success criteria before writing a line of SKILL.md. | Metric type | Example | |-------------|---------| | Time reduction | "From 2 hours to 15 minutes per API" | | Error reduction | "90% fewer missing test cases" | | Consistency | "Same structure across all 12 teams" | | Accessibility | "Junior devs can do it without help" | **5. Which skill category best fits?** Review `categories.md` and choose: - **Category 1**: Primary output is a document or code artifact - **Category 2**: Primary output is a completed multi-step process - **Category 3**: Primary output enhances or orchestrates MCP tools --- ## Phase 2: Design (Questions 6–10) **Purpose**: Translate use cases into skill structure. **6. What is the skill name?** Rules: - kebab-case only: `api-test-generator` not `apiTestGenerator` or `api_test_generator` - Describes function, not features: `meeting-notes-summarizer` not `gpt4-enhanced-transcriber` - No brand names unless truly required **7. What trigger phrases will users say?** The `description` field in YAML frontmatter is what Claude pattern-matches against user queries. Collect 4-8 trigger phrases users would naturally say when they need this skill. Examples for an API test generator: - "Generate API tests from my OpenAPI spec" - "Create a test suite for my REST API" - "Write tests for all my endpoints" - "Help me generate test coverage for my API" **8. What are the main workflow steps?** Break the skill into 3-5 phases. For each phase: - Name (action verb + noun): "Parse Spec", "Generate Tests", "Validate Output" - Definition of Done: what must be true before the next phase starts - What goes in, what comes out **9. What inputs does the skill need from the user?** List every piece of information the skill needs: | Input | Required? | Format | Example | |-------|-----------|--------|---------| | OpenAPI spec | Yes | YAML/JSON file path | `./api-spec.yaml` | | Test framework | No (default: Jest) | String | `pytest` | | Output directory | No (default: `./tests/`) | Path | `./src/__tests__/` | **10. What outputs does the skill produce?** Describe the output artifacts: - File type and location - Content structure - How the user will use each output --- ## Phase 3: Implementation (Questions 11–14) **Purpose**: Decide what goes in each directory. **11. What scripts should be bundled?** Scripts are appropriate when: - The same code would be rewritten every time the skill is used - The operation is deterministic and can be tested independently - The script runs faster via CLI than via Claude-generated code Examples: validators, generators, format converters, test runners. **12. What reference documentation should be bundled?** References are appropriate when: - There is a schema, API spec, or policy that Claude needs to consult - The documentation is large enough to justify keeping out of SKILL.md - The information changes independently of the workflow Examples: database schemas, API documentation, company policies, domain knowledge. **13. What examples should be bundled?** Examples go in `references/examples/` and are appropriate when: - There is a template file users will copy and adapt - A complete working example helps Claude understand the expected output format - Real-world samples prevent ambiguity about what "correct" looks like **14. What assets should be bundled?** Assets go in `assets/` and are appropriate when: - The skill produces output that includes non-text files (images, fonts, icons) - There is an HTML/React boilerplate the skill pastes into output - There is a binary template (PowerPoint, PDF) the skill modifies --- ## Phase 4: Testing (Questions 15–18) **Purpose**: Define testable criteria before writing SKILL.md. **15. What phrases should trigger the skill?** Write the exact triggering test cases: ```markdown Test 1: Slash command trigger Phrase: "/your-skill-name" Expected: Skill activates Test 2: Natural language trigger (most common form) Phrase: "[most natural way users will ask]" Expected: Skill activates Test 3: Closely related but different request Phrase: "[similar but out-of-scope request]" Expected: Skill does NOT activate ``` **16. What defines a passing functional test?** For each use case from Question 3, write: ```markdown Test: [Use Case Name] Input: [Exact input provided] Expected output: [Specific, verifiable result] Success criteria: [How to confirm it worked] ``` **17. What performance baseline are you improving on?** Measure before building: | Metric | Manual baseline | Target with skill | |--------|----------------|-------------------| | Time | [How long does it take today?] | [Target time] | | Quality | [Current error rate/coverage?] | [Target metric] | | Consistency | [How variable is manual process?] | [Standardized metric] | **18. Who will provide user feedback before release?** Identify 1-2 people who represent the target user: - Who will test the skill in a real scenario - What format feedback will be collected (written, meeting, async) - What criteria must pass before release --- ## Phase 5: Distribution (Questions 19–22) **Purpose**: Prepare for sharing and maintenance. **19. Where will the skill live for distribution?** - GitHub repository (recommended): `github.com/username/skill-name` - Internal company repository - Local only (no distribution planned) **20. What does the GitHub README need?** The repo-level README.md (OUTSIDE the skill folder) should include: - Outcome-focused positioning: "Generate tests 87% faster" (this is marketing, not SKILL.md) - Installation instructions (clone-to-install pattern) - Compatibility requirements (Claude Code version, MCP servers, etc.) - Example use cases with screenshots if applicable **21. What is the initial version number?** Use semantic versioning: - `0.1.0`: Early/experimental skill, API may change - `1.0.0`: Stable skill, ready for broad use - `1.x.0`: New features, backward compatible - `2.0.0`: Breaking changes to inputs or outputs **22. What is the support plan?** Before releasing publicly: - GitHub issues enabled? - Response time expectation (days, weeks, best-effort)? - Who owns updates if requirements change? --- ## Discovery Summary Template After completing all 22 questions, fill in this template to confirm you have everything needed: ```yaml skill_plan: name: "your-skill-name" # kebab-case, from Q6 category: "Category N: [Name]" # from Q5 trigger_phrases: # from Q7 (4-8 phrases) - "natural language phrase 1" - "natural language phrase 2" use_cases: # from Q3 (2-3 cases) - "Case 1: [concrete scenario]" - "Case 2: [concrete scenario]" success_criteria: # from Q4 quantitative: "Reduce X from Y to Z" qualitative: "[Quality improvement]" inputs: # from Q9 required: ["input1", "input2"] optional: ["input3 (default: value)"] outputs: # from Q10 - "output1: [file/format/location]" bundled_resources: # from Q11-14 scripts: ["name: purpose"] references: ["name: content description"] examples: ["name: what it shows"] assets: ["name: what it is"] testing: # from Q15-17 trigger_test: "Exact phrase that must activate skill" negative_test: "Phrase that must NOT activate skill" functional_test: "Concrete input → expected output" performance_baseline: "Manual: X min → Target: Y min" distribution: repo: "github.com/username/skill-name" version: "0.1.0" ``` -
distribution.md 8.5 KB
# Skill Distribution Guide How to package, document, and publish Claude Code skills for others to install. --- ## Current Distribution Landscape (2026) Skills are distributed directly via GitHub — no central marketplace exists yet. Installation is a simple git clone into `~/.claude/skills/`. This means: - **No approval process**: Publish when you're ready - **Version control**: GitHub handles versioning and changelogs - **Discovery**: Word of mouth, Claude Discord, social sharing - **Compatibility**: state the hosts and versions you tested. Discovery roots, `allowed-tools` handling, plugin versus standalone placement, and reload behavior all differ by host. ### Three distribution channels (choose one or all) **A. Claude.ai / Claude Code (individual install)** ```bash # Clone into Claude Code skills directory: cd ~/.claude/skills && git clone https://github.com/username/your-skill-name # Or: download ZIP → upload in Claude.ai Settings > Capabilities > Skills ``` **B. Organization-wide deployment** (admins only, shipped Dec 2025) Admins can deploy skills workspace-wide via Claude.ai admin console — automatic updates, centralized management. Users get the skill without any install step. **C. Programmatic / API** Add skills to Messages API requests via `container.skills` parameter. Use the `/v1/skills` endpoint to manage skills. Works with the Claude Agent SDK for building custom agents. Other hosts install from their own skill roots (`~/.agents/skills`, `~/.codex/skills`, `~/.qwen/skills`, and so on); the installer that places the skill must own exact generated paths, preserve manual edits, support multiple configured roots, and remove only its own links or files (SKILL-REQ010). --- ## Step 1: Prepare the Skill for Distribution Before creating a GitHub repo, verify: ```bash # 1. Correct file naming ls ~/.claude/skills/your-skill-name/SKILL.md # must exist ls ~/.claude/skills/your-skill-name/README.md # must NOT exist (inside skill folder) # 2. Correct folder naming (kebab-case) ls -d ~/.claude/skills/your-skill-name # no underscores, no capitals # 3. Run the audit script bash ~/.claude/skills/ai-skill-builder/scripts/audit-skill.sh \ ~/.claude/skills/your-skill-name # Gate: zero FAILs. The percentage it prints is a decided-check score — the share of # mechanically decidable checks that passed — so it says nothing about content quality. # Read the proxy lines and the "Needs a reader" list by hand before distributing. ``` ### Distribution Checklist - [ ] `SKILL.md` has YAML frontmatter with `name` and `description` - [ ] `description` uses trigger-phrase format (what users SAY to activate) - [ ] Folder is kebab-case, no README.md inside the skill folder - [ ] Triggering tests pass (skill activates on expected phrases) - [ ] Functional tests pass (at least happy path + 1 edge case) - [ ] `metadata.version` is set to `0.1.0` or higher (a top-level `version` fails `audit-skill.sh`) --- ## Step 2: Create the GitHub Repository The GitHub repo structure should follow this layout: ``` your-skill-repo/ ← GitHub repo root ├── README.md ← Human-facing docs (outcome-focused, installation, examples) └── your-skill-name/ ← actual skill folder (no README.md inside) ├── SKILL.md ├── references/ │ └── ... └── scripts/ └── ... ``` **Why this layout?** - `README.md` at repo root serves as the GitHub landing page for humans - The skill folder inside contains only what Claude needs - Users clone the repo root, getting the skill folder at the right nesting **Alternative for single-skill repos** (when the repo IS the skill): ``` your-skill-repo/ ← GitHub repo root = skill folder ├── README.md ← OK here: repo root is outside the skill's context ├── SKILL.md ├── references/ └── scripts/ ``` This works but the README.md rule applies inside the skill folder — since here they coincide, it's the exception documented in the Anthropic PDF. --- ## Step 3: Write the README.md The repo-level README.md is for **human readers deciding whether to install**. Use outcome-focused language here — this is where "generate tests 87% faster" belongs. ### README.md Template ````markdown # [Skill Name] > [One outcome-focused sentence: "Generate X in Y% less time"] [2-3 sentences describing the problem this skill solves and who benefits most.] ## Installation ```bash cd ~/.claude/skills git clone https://github.com/username/your-skill-name ``` Verify installation: ```bash ls ~/.claude/skills/your-skill-name/SKILL.md ``` Restart Claude Code, then activate with: `/your-skill-name` ## Requirements - Claude Code [version or "latest"] - [Any required MCP servers with install links] - [Any required CLI tools] ## What It Does [2-3 bullet points with concrete outcomes] - ✅ [Outcome 1 with metric] - ✅ [Outcome 2 with metric] - ✅ [Outcome 3 with metric] ## Usage Examples **[Example 1 scenario]:** ``` /your-skill-name [typical usage] ``` **[Example 2 scenario]:** ``` [Natural language trigger phrase] ``` ## Changelog **v0.1.0** - [Date] - Initial release ## License [MIT / Apache-2.0 / etc.] ```` ### README.md Content Rules | Include | Exclude | |---------|---------| | Outcome-focused description | Feature lists ("uses GPT-4, Jinja2, OpenAPI parser") | | Specific improvement metrics | Vague claims ("saves time", "improves quality") | | Prerequisites with install links | Implementation details | | Working install commands | Marketing language without evidence | | Real usage examples | Changelog entries (put in CHANGELOG.md) | --- ## Step 4: Publishing and Positioning ### Positioning Language (for README and community posts) Outcome-focused language drives adoption. Show what users achieve, not what the skill uses. **Pattern**: `[Action verb] [what] [quantifiable improvement]` ```markdown ✅ "Generate API test suites 87% faster than writing them manually" ✅ "Deploy to production in 10 minutes instead of 2 hours" ✅ "Reduce missing test cases by 90% with spec-driven generation" ❌ "Uses OpenAPI parser and Jinja2 templates" ❌ "AI-powered test generation" ❌ "Smart and efficient API testing" ``` ### Community Sharing Current best channels (2026): 1. **Claude Discord** — `#skills` or relevant domain channel; share install command + outcome 2. **GitHub** — Star and fork encourage discovery; good README drives organic sharing 3. **Reddit** — r/ClaudeAI, r/artificial, domain-specific subreddits 4. **Twitter/X** — Demo GIF + install command gets traction 5. **Domain communities** — Dev forums, Slack communities in the skill's target domain ### Support Commitment State your support policy in the README: ```markdown ## Support Issues: https://github.com/username/your-skill-name/issues Response time: [Best effort / within 1 week / actively maintained through YYYY] ``` Be honest. An unsupported skill that works reliably is better than a maintained skill that breaks. --- ## Step 5: Versioning and Updates ### Semantic Versioning ``` v1.0.0 ← stable, ready for production │ │ └─ patch: bug fixes, typo corrections │ └─── minor: new features, backward compatible (new workflow steps, new references/) └───── major: breaking changes (renamed triggers, removed workflow steps, changed output format) ``` ``` v0.x.0 ← experimental, API may change ``` ### Changelog Keep a `CHANGELOG.md` or update `references/changelog.md` with each release: ```markdown ## v1.1.0 — 2026-03-15 ### Added - Support for Python test frameworks (pytest, unittest) - `references/python-examples/` with working test files ### Fixed - Triggering phrase "write API tests" now activates correctly - Auth test generation works with Bearer tokens ### Changed - Output directory default changed from `./tests/` to `./src/__tests__/` (upgrade note: update your `.gitignore` if needed) ``` ### Notifying Users of Updates Users who installed via git clone can update with: ```bash cd ~/.claude/skills/your-skill-name git pull ``` Add this to your README's "Updating" section. --- ## GitHub Repository Setup Checklist - [ ] Repo created with correct name (matches skill folder name) - [ ] `README.md` at repo root with outcome-focused description - [ ] Installation instructions tested on a fresh machine (or fresh `~/.claude/skills/`) - [ ] `LICENSE` file added (MIT is common for skills) - [ ] GitHub Issues enabled for support - [ ] `CHANGELOG.md` or `references/changelog.md` started - [ ] Initial tag created: `git tag v0.1.0 && git push --tags` -
patterns.md 8.4 KB
# Advanced Skill Patterns Five advanced patterns from Anthropic's "The Complete Guide to Building Skills for Claude" (January 2026, Chapter 5). Apply these when the basic 4-phase workflow is insufficient. For fixing broken skills, see `troubleshooting.md` in this directory. --- ## Pattern 1: Sequential Workflow Orchestration **When to use**: Tasks with 3+ phases where each phase's output feeds the next, and failures at any phase must halt the entire workflow. ### Structure ``` Phase 1 → validate output → Phase 2 → validate output → Phase 3 → final report ↓ ↓ halt + error halt + error ``` ### Implementation ```markdown ## Workflow Phases ### Phase 1: [Name] **Input**: [What this phase receives] **Process**: [What it does] **Output**: [What it produces] **Validation gate**: [What must be true before proceeding to Phase 2] **On failure**: Stop. Report: "[Phase 1 failed: specific reason]. Fix [X] and retry." ### Phase 2: [Name] **Input**: Phase 1 output ([specific artifact]) **Process**: [What it does] **Output**: [What it produces] **Validation gate**: [What must be true before proceeding to Phase 3] **On failure**: Stop. Report: "[Phase 2 failed]. Phase 1 output preserved at [location]." ### Final Report Format - ✅ Phase 1: [Result summary] - ✅ Phase 2: [Result summary] - ✅ Phase 3: [Result summary] - Total: [Completion metric] ``` ### Example: Database Migration Skill ``` Parse migration file → validate schema → apply changes → verify row counts → generate report ↓ ↓ ↓ syntax error type mismatch rollback + error ``` Each phase reports its status before proceeding. If Phase 3 fails, Phase 1 and 2 results are preserved and a rollback path is provided. --- ## Pattern 2: Multi-MCP Coordination **When to use**: Skills that need data or operations from 2+ MCP servers, where the servers have different availability, latency, or reliability profiles. ### Design Principles 1. **Declare dependencies explicitly**: List required vs optional MCP servers in SKILL.md 2. **Degrade gracefully**: If an optional server is unavailable, continue without it 3. **Cache aggressively**: MCP calls can be slow; cache results that are stable across requests 4. **Handle partial failures**: One server failing should not fail the entire workflow ### Implementation in SKILL.md ```markdown ## MCP Server Requirements ### Required (skill cannot function without these) - **[server-name]**: [What it provides] Install: [installation command] Test: [verification command] ### Optional (skill degrades gracefully without these) - **[server-name]**: [What it adds when available] Without it: [what the skill does instead] ## Coordination Logic To execute [operation]: 1. Check which servers are available (use list_tools or equivalent) 2. If [required-server] unavailable: report error and stop 3. If [optional-server] unavailable: note in output, continue without [feature] 4. Execute [primary operation] via [required-server] 5. If [optional-server] available: enrich result with [additional data] 6. Merge results and respond ``` ### Example: Smart File Search Skill ```markdown ## MCP Server Requirements ### Required - **filesystem**: Read file contents and directory structure ### Optional - **semantic-search**: Find conceptually related files (not just text matches) Without it: falls back to recursive text search with Grep ## Coordination Logic To search for [query]: 1. Always: use filesystem to get directory tree 2. Always: use Grep for exact text matches 3. If semantic-search available: augment with semantic matches, rank by relevance 4. If not: return exact matches with helpful note about semantic search ``` --- ## Pattern 3: Iterative Refinement **When to use**: Skills that produce output requiring validation and revision cycles, where the first attempt is rarely the final result. ### Structure ``` Generate draft → validate against criteria → if pass: done ↓ if fail: identify gaps → refine → re-validate (max N iterations) ↓ if still failing after N: report with partial result ``` ### Implementation ```markdown ## Refinement Loop To generate [output]: ### Attempt 1: Initial generation Generate [output] from [input]. ### Validation checkpoint Check [output] against: - [ ] Criterion 1: [specific, measurable check] - [ ] Criterion 2: [specific, measurable check] - [ ] Criterion 3: [specific, measurable check] If all pass: deliver output. If any fail: identify which criteria failed, proceed to refinement. ### Refinement (up to 2 iterations) For each failed criterion: - Explain what specifically failed - Generate targeted fix for that criterion only - Re-validate the fixed criterion ### Final delivery If all criteria pass after refinement: deliver with note about iterations taken. If criteria still fail after 2 iterations: deliver best result with explicit gaps noted. Never silently deliver output that failed validation. ``` --- ## Pattern 4: Context-Aware Tool Selection **When to use**: Skills that need to choose between multiple approaches based on the user's environment, preferences, or constraints detected at runtime. ### Design Rather than one rigid workflow, provide conditional branches based on detected context: ```markdown ## Context Detection Before starting, determine: 1. **Language/framework**: Detect from package.json, requirements.txt, go.mod, etc. 2. **Project structure**: Detect from directory layout 3. **Existing conventions**: Read existing files in the output directory to match style 4. **User preferences**: Check if user specified any in their request ## Workflow Selection Based on detected context, choose the appropriate path: | Detected context | Workflow to use | |-----------------|-----------------| | package.json with Jest | `references/examples/jest-workflow.md` | | requirements.txt with pytest | `references/examples/pytest-workflow.md` | | go.mod present | `references/examples/go-test-workflow.md` | | No detection possible | Ask user: "What test framework do you prefer?" | Never assume context. If detection is ambiguous, ask rather than guess. ``` --- ## Pattern 5: Domain Intelligence Layer **When to use**: Skills operating in a specific domain (finance, healthcare, legal, security) where domain rules, terminology, and compliance requirements affect every decision. ### Structure Domain knowledge lives in `references/` files, not in SKILL.md: ``` SKILL.md → workflow steps (domain-agnostic procedure) references/domain.md → domain rules, terminology, compliance requirements references/schema.md → domain-specific data structures references/examples/ → validated domain-specific examples ``` ### Implementation ```markdown ## Domain Context Before executing [operation], load domain context from `references/domain.md`. Key domain rules that affect this workflow: - [Rule 1]: [How it affects the output] - [Rule 2]: [How it affects the output] - [Compliance requirement]: [What must always be true in the output] Apply these rules at [specific step in workflow]. If a rule conflicts with the user's request, explain the conflict and ask for clarification rather than silently applying or ignoring the rule. ``` ### Example: Finance Skill ```markdown ## Domain Context Before generating any financial analysis, load `references/finance-rules.md`. Domain constraints: - Currency: Always specify ISO 4217 code (USD, EUR, not $ or €) - Dates: Always ISO 8601 (2026-03-05, not March 5, 2026) - Amounts: Always use integer cents internally, format as decimals for display - Disclosures: Any forward-looking statements must include disclaimer from `references/disclosures.md` ``` --- ## Combining Patterns Patterns compose. A single skill can use multiple patterns: **Example: Enterprise Deployment Skill** - Pattern 1 (Sequential): Deploy → smoke test → notify team → update runbook - Pattern 4 (Context-aware): Different steps for AWS vs GCP vs on-prem - Pattern 5 (Domain): Follow company-specific deployment policy from `references/policy.md` Add each pattern's section to SKILL.md separately, clearly labeled, so Claude can load only the relevant patterns for a given request. --- ## Source Anthropic, "The Complete Guide to Building Skills for Claude," January 2026, Chapter 5 (pages 21-26). -
portability-and-claim-audit.md 11.1 KB
# Agent Skills portability and claim audit Use this reference when a skill targets more than one runtime, when third-party guidance makes universal claims, or when a claim about frontmatter, XML, skill roots, symlinks, arguments, or security needs checking. Recheck linked primary documentation because harness behavior evolves. It backs SKILL-REQ002, SKILL-REQ004, SKILL-REQ009, and SKILL-REQ010 in `SKILL.md`. ## Portable specification core The current Agent Skills specification requires a directory containing `SKILL.md`, YAML frontmatter, and a Markdown body. The standard fields are: | Field | Standard status | Important constraint | | --- | --- | --- | | `name` | Required | 1–64 lowercase alphanumeric/hyphen characters; no edge or consecutive hyphens; match parent directory | | `description` | Required | 1–1024 characters; state what the skill does and when to use it | | `license` | Optional | Short license name or bundled license reference | | `compatibility` | Optional | 1–500 characters; only when environment requirements matter | | `metadata` | Optional | String-to-string extension metadata | | `allowed-tools` | Optional, experimental | Space-separated string; support varies by implementation | The portable body has no mandated XML or section taxonomy. This methodology nevertheless mandates consistent semantic XML regions as a higher-quality authoring policy (SKILL-REQ004). Keep those two facts separate: portable parsers need only accept Markdown, while packages produced by this methodology must also pass balanced-tag validation (`scripts/audit-skill.sh`, section 4) and real target-runtime task evaluation. `scripts/`, `references/`, and `assets/` are optional conventions. The specification recommends instructions below 5,000 tokens, `SKILL.md` below 500 lines, focused one-hop references, and validation with `skills-ref`; those are recommendations, not evidence that every runtime has the same loader or execution policy. Primary source: <https://agentskills.io/specification> ## Runtime differences that must remain explicit | Capability | Portable standard | Claude Code | GitHub Copilot | Local Codex authoring policy, not runtime-support evidence | | --- | --- | --- | --- | --- | | `name`, `description` | Required | More permissive parser; description recommended | Required in documented simple form | Requires only these fields in `SKILL.md` frontmatter | | `license`, `compatibility`, `metadata` | Optional | Runtime-dependent acceptance | `license` documented; check other fields per surface | Put product UI data in `agents/openai.yaml`, not extra `SKILL.md` frontmatter | | `allowed-tools` | Experimental string | Supported with Claude-specific permission semantics and extra tool fields | Supported; dangerous shell approval warning | Do not assume portable behavior | | `when_to_use` | Not standard | Claude extension | Not established by the cited GitHub guide | Put triggers in `description` | | `arguments` and argument hints | Not standard | Claude extensions with `$ARGUMENTS`, `$N`, and named positions | Invocation syntax differs; verify current product | Not a portable frontmatter schema; `quick_validate.py` rejects `argument-hint` and `aliases` | | `context: fork`, `agent`, `background`, `effort`, `model` | Not standard | Claude extensions | Not established by the cited GitHub guide | Do not put in portable core | | Top-level `version` | Not standard | Ignored | Not established | Rejected by `quick_validate.py`; use `metadata.version` | | Project roots | Implementation-defined | `.claude/skills` | `.github/skills`, `.claude/skills`, `.agents/skills` | `.agents/skills` is the project convention here | | Personal roots | Implementation-defined | `~/.claude/skills` | `~/.copilot/skills`, `~/.agents/skills` | Harness-specific install system | | Symlinks | Implementation-defined | Supported for current personal/project skill entries, with version-specific behavior; the `skills/` directory itself must not be a symlink (anthropics/claude-code#38051) | Verify per product and OS | Test rather than infer | ### Evidence status | Target | Evidence status | What is and is not established | | --- | --- | --- | | Agent Skills portable format | Primary specification checked 2026-07-27 | Format fields, body freedom, directories, recommendations, and validation command | | Claude Code | Official current docs checked 2026-07-27 | Locations, symlink behavior, frontmatter extensions, arguments, forked context, and lifecycle; no cross-version guarantee | | GitHub Copilot surfaces | Official current docs checked 2026-07-27 | Documented project/personal roots and basic/allowed-tool behavior; no blanket claim for every Copilot host/version | | Microsoft Agent Framework | Official current docs checked 2026-07-27 | Provider-driven progressive disclosure and experimental MCP skill source; not a filesystem-root claim for unrelated harnesses | | Local Codex skill creator | Installed authoring policy and validator checked 2026-07-27 and 2026-08-15 | This package's accepted structure; `quick_validate.py` allows only `allowed-tools`, `description`, `license`, `metadata`, `name` at top level; not proof of discovery/invocation in every Codex product | | Qwen Code | Official docs checked 2026-08 (`references/sources.md`) | Personal, project, and extension skill locations; model- and user-invocation behavior; no cross-version guarantee | | Codex/ChatGPT | Official docs checked 2026-08 (`references/sources.md`) | Discovery, explicit invocation, progressive disclosure; validator behavior above | | Pi, Prime Agent, OpenCode, ForgeCode, Antigravity, Gemini CLI, OpenClaw, GLM, and other named runtimes | Unknown/not assessed here | Do not claim compatibility until official documentation and direct forward tests fill the validation receipt | Primary sources: - Claude Code skills: <https://code.claude.com/docs/en/skills> - GitHub Copilot skills: <https://docs.github.com/en/copilot/how-tos/copilot-on-github/customize-copilot/customize-cloud-agent/add-skills> - Microsoft Agent Framework skills: <https://learn.microsoft.com/en-us/agent-framework/agents/skills> - Current installed Codex skill-creator instructions and `scripts/quick_validate.py` (environment-specific; reread them from the active harness before authoring) ## Claim audit Claims met in third-party "master protocols" and marketing, with the assessment this methodology adopts. Where a claim is rejected, the corrected guidance is what `SKILL.md` teaches. | Supplied claim | Assessment | Corrected guidance | | --- | --- | --- | | Agent Skills is an open portable format | Retain | Portability applies to the common format; every runtime still needs direct compatibility tests | | Discovery loads name/description, then instructions and resources on demand | Retain as the standard model | Do not promise identical startup, caching, script, or context behavior in every implementation | | The portable Agent Skills specification requires XML | Reject | The specification requires Markdown and imposes no body format restrictions | | This methodology requires semantic XML regions | Retain as an explicit quality policy | Apply balanced, descriptive tags to major instruction regions and forward-test task outcomes in every target runtime | | Visual HTML is ignored by transformers | Reject | Models process HTML tokens; usefulness depends on semantics and task, not a universal ignore rule | | Custom XML tags are hard instruction/security boundaries | Reject | XML can improve clarity for complex prompts but is not an authorization or injection boundary | | XML reduces failures by 28–40 percent | Reject without a reproducible primary benchmark | Do not import precise performance claims from secondary marketing/blog sources | | Every methodology XML region must be balanced and clearly delimited | Retain as a methodology rule | This makes the selected XML structure mechanically auditable; it remains separate from portable parser requirements | | Raw angle brackets break YAML scalars | Reject as a universal claim | YAML parses `<` and `>` in a plain scalar. The frontmatter ban is a host restriction: Anthropic's skill guide lists "XML angle brackets" as forbidden in frontmatter and Claude Code documents that `description` must not contain XML tags. Enforce it for those hosts; do not explain it as YAML | | "No XML tags anywhere" (Anthropic guide checklist wording) | Conditional | The enumerated restriction in the same guide is frontmatter-only (Reference B). Body regions rest on Anthropic's prompting guidance and are this methodology's policy. A host that rejects body tags on upload is recorded as unsupported in the validation receipt rather than assumed | | `when_to_use`, `context`, `effort`, and `arguments` are universal frontmatter | Reject | These are runtime extensions, notably in current Claude Code, not Agent Skills core fields | | `allowed-tools` is an array | Reject for the portable spec | The standard defines an experimental space-separated string; some runtimes accept other shapes | | Directory and `name` must match | Retain for the portable spec | Enforce it in portable packages even if a permissive runtime allows otherwise | | Keep `SKILL.md` below 500 lines and move details to references | Retain as a recommendation | Optimize for task success and loaded tokens; do not split cohesive instructions merely to satisfy a number | | `.github/skills`, `.claude/skills`, and `.agents/skills` are universally discovered | Reject | These are product-specific locations; GitHub documents all three, Claude documents `.claude/skills` | | Scripts need executable permissions | Conditional | Depends on how the runtime invokes them; still test permissions, interpreter, dependencies, and errors | | All major named frontier and open-source runtimes are compatible | Unsupported | Claim only products and versions directly documented and forward-tested | Anthropic's official prompting guidance says XML tags can help Claude distinguish mixed prompt components. That supports this methodology's quality choice, but it does not make XML mandatory in the portable Agent Skills specification or establish XML as prompt-injection prevention: <https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices#structure-prompts-with-xml-tags>. The supplied 28–40 percent improvement claim is not adopted. It was not backed by a reproducible primary benchmark spanning the target runtimes and actual skill tasks. Measure automatic activation, instruction adherence, task correctness, and regression rate on representative tasks; the methodology mandate stands as a selected engineering policy and should still be improved when better direct evidence becomes available. ## Validation receipt For every claimed target, record: ```text product and version: skill root: portable fields accepted: runtime extensions used: automatic activation: explicit invocation: arguments: references: scripts: symlink behavior: duplicate/name precedence: live reload or restart: install/update/uninstall: negative and adversarial prompts: measured context/runtime cost: unsupported or unknown: ``` Attach the receipts to the skill's changelog or release notes. A target without a receipt is "unknown", never "supported". -
refining-skills.md 16.2 KB
# Refining Existing Skills - Complete Guide **Source**: Anthropic's "The Complete Guide to Building Skills for Claude" (January 2026) **PDF**: `../ai-skill-builder-guide.pdf` This guide covers improving and modernizing existing Claude skills to match current best practices. --- ## When to Refine a Skill ### Signs Your Skill Needs Refinement **Critical Issues** (fix immediately): - ❌ File named `README.md` instead of `SKILL.md` - ❌ Folder uses underscores (`my_skill`) or camelCase (`mySkill`) - ❌ Missing YAML frontmatter - ❌ Feature-focused description ("Uses X to Y") - ❌ No progressive disclosure structure **Quality Issues** (improve soon): - ⚠️ Description too short to carry 4-8 quoted trigger phrases, or over the 1024-character limit - ⚠️ No examples section - ⚠️ Missing trigger phrases - ⚠️ No success metrics - ⚠️ Wall of text (no structure) **Enhancement Opportunities** (nice to have): - 💡 Could add automation scripts - 💡 Missing reference documentation - 💡 Could benefit from templates/assets - 💡 Performance not measured --- ## Refinement Workflow ### Phase 1: Audit **Step 1: Run Automated Audit** ```bash bash ../scripts/audit-skill.sh ~/.claude/skills/YOUR-SKILL ``` **Step 2: Document Findings** Create audit report: ```markdown # Skill Audit Report - [Skill Name] Date: [Current date] Auditor: Claude Skill Path: ~/.claude/skills/[skill-name] ## Critical Issues - [ ] Issue 1: [Description] - [ ] Issue 2: [Description] ## Quality Issues - [ ] Issue 1: [Description] ## Enhancement Opportunities - [ ] Opportunity 1: [Description] ## Score: X/100 ## Next Steps 1. [Priority 1 fix] 2. [Priority 2 fix] ``` ### Phase 2: Prioritize **Priority Framework:** **P0 - Critical (must fix)**: - Incorrect file naming - Missing required fields - Broken skill detection **P1 - High (should fix)**: - Poor progressive disclosure - Missing examples - Vague descriptions **P2 - Medium (nice to have)**: - Additional documentation - Automation scripts - Performance optimizations **P3 - Low (future enhancement)**: - Visual assets - Advanced features - Edge case handling ### Phase 3: Implement Fixes **For Critical Issues:** **Issue**: Wrong filename ```bash # If file is README.md or wrong case # Claude uses Read + Write tools: # 1. Read current content Read: ~/.claude/skills/my-skill/README.md # 2. Write to correct filename Write: ~/.claude/skills/my-skill/SKILL.md [Copy content] # 3. Remove old file — use trash, not rm: rm destroys the only copy and its deletion time Bash: trash ~/.claude/skills/my-skill/README.md ``` **Issue**: Wrong folder name ```bash # Rename folder to kebab-case mv ~/.claude/skills/my_old_skill ~/.claude/skills/my-old-skill ``` **Issue**: Missing YAML frontmatter ```markdown # Add to top of SKILL.md: --- name: skill-name description: Generates X from Y. Use when user asks to "trigger phrase 1", "trigger phrase 2". metadata: version: 0.1.0 --- ``` **Issue**: Feature-focused description ```yaml # ❌ Before: description: Uses OpenAPI parser and Jinja2 to generate Jest tests # ✅ After: description: Generates Jest test suites from OpenAPI specs. Use when user asks to "generate API tests", "create a test suite from my spec", "write tests for my endpoints". ``` **For Quality Issues:** **Issue**: No progressive disclosure Claude will use Edit tool: ```markdown # 1. Read current SKILL.md Read: ~/.claude/skills/skill-name/SKILL.md # 2. Edit to add structure using Edit tool Edit: ~/.claude/skills/skill-name/SKILL.md old_string: [current unstructured content] new_string: # Skill Name [Level 1: Hook - 50-100 words] --- ## How It Works [Level 2: Workflow - 200-400 words] --- ## Detailed Guide [Level 3: Comprehensive] ``` **Issue**: No examples Add examples section: ```markdown ## Examples ### Example 1: [Common Use Case] **Scenario**: [Specific situation] **Input**: \`\`\` [Actual input] \`\`\` **Output**: \`\`\` [Actual output] \`\`\` **Result**: [Outcome achieved] ``` ### Phase 4: Test & Validate **Re-run Audit**: ```bash bash ../scripts/audit-skill.sh ~/.claude/skills/YOUR-SKILL ``` **Test Triggering**: ``` # Try exact trigger: /your-skill-name # Try natural language: "Help me with [skill purpose]" # Verify it doesn't trigger on unrelated: "Something completely different" ``` **Functional Testing**: 1. Run through complete workflow 2. Verify outputs match documentation 3. Check error handling works 4. Validate edge cases **Compare Metrics**: ```markdown ## Before vs After Fill this table from your own measurements. The values below are placeholders, not results from any measured skill. | Metric | Before | After | Improvement | |--------|--------|-------|-------------| | Audit score | [measured] | [measured] | [delta] | | Trigger accuracy (passes / triggering tests run) | [measured] | [measured] | [delta] | | Time to complete (median of 3 runs) | [measured] | [measured] | [delta] | ``` ### Phase 5: Document Changes **Update Version History**: ```markdown ## Version History **v2.0.0** - 2026-02-15 - BREAKING: Renamed from my_old_skill to my-new-skill - BREAKING: Changed trigger from /old to /new - Added progressive disclosure structure - Rewrote description to capability-plus-quoted-trigger-phrase format - Added 3 concrete examples - Audit score improved from [before] to [after] **v1.0.0** - 2025-12-01 - Initial release ``` --- ## Common Refinement Scenarios ### Scenario 1: Migrating Old Skill to New Standard **Starting Point**: Skill created before Anthropic's guide (pre-2026) **Migration Checklist**: - [ ] Rename README.md → SKILL.md - [ ] Fix folder name to kebab-case - [ ] Add YAML frontmatter - [ ] Convert description to capability-plus-quoted-trigger-phrase format - [ ] Add progressive disclosure (3 levels) - [ ] Add examples section - [ ] Add success metrics - [ ] Add version history - [ ] Run audit script - [ ] Test triggering and functionality **Definition of Done**: the audit script reports no failures and triggering tests pass **Example**: ```bash # Before: ~/.claude/skills/API_Test_Gen/README.md # No frontmatter # Feature-focused description # No structure # After: ~/.claude/skills/api-test-generator/SKILL.md --- name: api-test-generator description: Generates Jest test suites from OpenAPI specs. Use when user asks to "generate API tests", "create a test suite from my spec". --- # API Test Generator [Progressive disclosure structure] ``` ### Scenario 2: Improving Existing Good Skill **Starting Point**: Skill follows basics but could be better **Enhancement Checklist**: - [ ] Audit with script - [ ] Add concrete examples (if missing) - [ ] Add automation scripts - [ ] Add reference documentation - [ ] Improve success metrics measurement - [ ] Add troubleshooting section - [ ] Enhance error messages - [ ] Add related skills links **Definition of Done**: every checklist item above is satisfied ### Scenario 3: Adding Automation to Manual Skill **Starting Point**: Skill works but requires manual steps **Automation Checklist**: - [ ] Identify repetitive manual steps - [ ] Create automation script(s) - [ ] Add to scripts/ directory - [ ] Update SKILL.md workflow - [ ] Test automation end-to-end - [ ] Document automation requirements - [ ] Add rollback procedures **Definition of Done**: the automation runs end to end and its failure path is documented **Example**: ```bash # Add automation script Write: ~/.claude/skills/api-test-generator/scripts/generate.py # Update SKILL.md to reference script Edit: ~/.claude/skills/api-test-generator/SKILL.md old_string: "3. Manually create test files" new_string: "3. Run: python scripts/generate.py --spec openapi.yaml" ``` ### Scenario 4: Splitting Overly Complex Skill **Starting Point**: One skill doing too many things **Splitting Strategy**: 1. **Identify distinct capabilities** (should be separate skills) 2. **Create new skills** for each capability 3. **Keep original as orchestrator** (if needed) 4. **Update documentation** with links to related skills **Example**: ```markdown # Original: api-automation (does everything) → Split into: - api-test-generator (testing) - api-docs-generator (documentation) - api-client-generator (client code) - api-automation (orchestrator - optional) ``` --- ## Measuring Improvement ### Before/After Metrics **Audit Scores**: run the script before and after, and keep both outputs. What moves is the decided-check score and the FAIL count; proxy checks are heuristics and are reported separately, not scored. ```bash bash ../scripts/audit-skill.sh ~/.claude/skills/my-skill | tee before.txt # ...apply fixes... bash ../scripts/audit-skill.sh ~/.claude/skills/my-skill | tee after.txt diff before.txt after.txt ``` **User Satisfaction** (gather feedback): ```markdown Survey questions: 1. How easy was it to understand when to use this skill? (1-5) 2. How clear was the workflow? (1-5) 3. How helpful were the examples? (1-5) 4. Would you recommend this skill? (Yes/No) 5. What could be improved? ``` **Performance Metrics**: - Time to complete task (before vs after) - Error rate (failures/attempts) - Adoption rate (usage growth) - Support requests (reduction) **Quality Indicators**: - Progressive disclosure compliance (Yes/No) - Example coverage (# of examples) - Success metrics defined (Yes/No) - Automated tests passing (%) --- ## Continuous Improvement ### Regular Maintenance Schedule **Monthly**: - Review skill usage analytics - Collect user feedback - Check for broken examples - Update dependencies **Quarterly**: - Run audit script - Review and update examples - Improve documentation - Add requested features **Yearly**: - Major version upgrade - Align with latest best practices - Comprehensive testing - Performance optimization ### Feedback Loop **Collect Feedback**: 1. User survey after skill usage 2. GitHub issues 3. Discord discussions 4. Support tickets **Prioritize Improvements**: ```markdown Impact vs Effort Matrix: High Impact, Low Effort: - Do immediately High Impact, High Effort: - Plan for next quarter Low Impact, Low Effort: - Do when time permits Low Impact, High Effort: - Defer or reject ``` **Implement & Measure**: 1. Make changes 2. Re-run audit 3. Test with users 4. Measure impact 5. Document learning --- ## Refinement Tools & Resources ### Automated Tools **Audit Script**: ```bash ../scripts/audit-skill.sh ``` **Scaffolding** (for new structure): ```bash ../scripts/scaffold-skill.sh ``` ### Manual Tools (Claude uses these) **Read Tool**: Review current state ``` Read: ~/.claude/skills/skill-name/SKILL.md ``` **Edit Tool**: Make precise changes ``` Edit: ~/.claude/skills/skill-name/SKILL.md old_string: [exact text to replace] new_string: [new text] ``` **Write Tool**: Create new files ``` Write: ~/.claude/skills/skill-name/new-file.md [content] ``` **Bash Tool**: File operations ``` Bash: mv old-name new-name Bash: chmod +x script.sh ``` ### Reference Materials **Official Guide**: - PDF: `../ai-skill-builder-guide.pdf` - Extracted text: `ai-skill-builder-guide.md` **Templates**: - `examples/SKILL-template.md` **Best Practices**: - `best-practices.md` --- ## Quick Reference: Refinement Workflow ``` ┌─────────────────────────────────────────────┐ │ 1. AUDIT │ │ bash audit-skill.sh my-skill │ │ Document findings │ └─────────────────────┬───────────────────────┘ │ ▼ ┌─────────────────────────────────────────────┐ │ 2. PRIORITIZE │ │ P0: Critical (file naming, etc) │ │ P1: Quality (structure, examples) │ │ P2: Enhancement (scripts, docs) │ └─────────────────────┬───────────────────────┘ │ ▼ ┌─────────────────────────────────────────────┐ │ 3. FIX │ │ Critical → Quality → Enhancements │ │ Use Read/Edit/Write tools │ └─────────────────────┬───────────────────────┘ │ ▼ ┌─────────────────────────────────────────────┐ │ 4. TEST │ │ Re-run audit │ │ Test triggering │ │ Functional testing │ │ Measure metrics │ └─────────────────────┬───────────────────────┘ │ ▼ ┌─────────────────────────────────────────────┐ │ 5. DOCUMENT │ │ Update version history │ │ Document improvements │ │ Share learnings │ └─────────────────────────────────────────────┘ ``` --- ## Examples of Real Refinements ### Example 1: README.md → SKILL.md Migration **Before** (v1.0): ``` ~/.claude/skills/test_gen/README.md No frontmatter One big paragraph description No structure ``` **After** (v2.0): ``` ~/.claude/skills/test-generator/SKILL.md --- name: test-generator description: Generates test suites from code analysis. Use when user asks to "generate tests", "write tests for this module". --- # Test Generator Generate test suites from code analysis... ## How It Works 1. Analyze code — done when every exported symbol is listed 2. Generate tests — done when each listed symbol has a test file 3. Validate — done when the suite parses and runs ``` **What changed**: file renamed so the host can find it, folder renamed to kebab-case, frontmatter added with trigger phrases, workflow split into three steps that each state when they are finished. ### Example 2: Adding Progressive Disclosure **Before** (wall of text): ```markdown This skill helps you deploy applications to production by first checking prerequisites then building docker images then pushing to registry then deploying to kubernetes then validating deployment then monitoring for issues and rolling back if needed. It supports multiple environments including dev staging and production. You can configure timeouts health checks and rollback thresholds... ``` **After** (structured): ```markdown # Deploy to Production Deploy containerized applications with automated validation and rollback. **Use when:** Ready to ship to production **Invoke with:** `/deploy-to-production` --- ## How It Works ### Step 1: Pre-flight Checks - Validates prerequisites - Checks environment health - Done when: every prerequisite reports healthy ### Step 2: Build & Push - Builds Docker image - Pushes to registry - Done when: the registry returns the pushed digest ### Step 3: Deploy & Validate - Deploys to Kubernetes - Validates health checks - Done when: all pods are Ready, or rollback has completed --- ## Detailed Guide [Comprehensive documentation...] ``` **What changed**: one 60-word paragraph became a 20-word hook, three named steps each stating what must be true before the next one runs, and a pointer to the detailed guide. A reader can now decide whether the skill applies without reading past the hook. --- ## Summary **Key Principles**: 1. **Audit first** - Know what to improve 2. **Prioritize ruthlessly** - Fix critical issues first 3. **Test thoroughly** - Measure impact against the baseline you recorded 4. **Document changes** - Help future maintainers 5. **Iterate continuously** - Skills are never "done" **Source**: Adapted from Anthropic's "The Complete Guide to Building Skills for Claude" (January 2026) -
research.md 7.9 KB
# Research Strategies for Skill Building Skills encode procedures at scale — every user who triggers a skill follows its instructions, confidently. Outdated or incorrect guidance is worse than none. Research cost is paid once; error cost is paid on every invocation. --- ## Tools | Tool | When to use | |------|------------| | `WebSearch` | Find current best practices, compare approaches, check community consensus | | `WebFetch` | Read official documentation, RFCs, API references, changelogs | | `Read` | Read local docs, existing code, config files in the user's project | | `Bash` | Run the tool being documented to observe actual behavior directly | --- ## Research Strategies by Domain Skills can cover any domain. The research approach follows what you're investigating. ### Tool or API The tool's own behavior is ground truth; everything else is secondary. - Current version's docs and changelog — behavior changes between versions - The actual installed version (`Bash: tool --version`, then test it) - Known failure modes: what does the community get wrong most often? ``` WebSearch: "jest 29 snapshot testing pitfalls 2025" WebFetch: https://jestjs.io/docs/29.x/snapshot-testing Bash: jest --version ``` ### Domain Practice (security, accessibility, data science, law, writing…) Trace to authoritative bodies — not tutorials summarizing other tutorials. - Authoritative body for the domain (OWASP for security, W3C for web, NIST for crypto) - Primary literature: RFCs, published standards, peer-reviewed papers if they exist - Current-year community consensus — practices shift; old tutorials are actively harmful ``` WebSearch: "OWASP SQL injection prevention 2025" WebFetch: https://cheatsheetseries.owasp.org/cheatsheets/SQL_Injection_Prevention_Cheat_Sheet.html WebSearch: "parameterized queries python psycopg3" ``` ### Workflow or Process Processes have order-of-operations constraints and failure modes that documentation understates. Research both explicitly. - Official runbook or checklist from the authoritative source - Failure modes and recovery paths — what goes wrong, how to detect and reverse it - Prerequisites and postconditions for each step ``` WebSearch: "zero-downtime postgres ALTER TABLE production 2025" WebFetch: https://www.postgresql.org/docs/current/sql-altertable.html WebSearch: "postgres ALTER TABLE lock timeout production" ``` ### Creative or Subjective Domain (writing, design, UX…) Even subjective domains have evidence-based principles. - Style guides or canonical references for the domain - Empirical research where it exists (readability studies, accessibility audits) - Where experts disagree, document both positions rather than picking one arbitrarily ``` WebSearch: "plain language writing guidelines" WebFetch: https://www.plainlanguage.gov/guidelines/ WebSearch: "sentence length readability studies" ``` ### External Integrations and MCP Servers Schemas and rate limits change without notice — never assume from memory. - Current parameter schema for each tool or MCP server - Rate limits, quotas, and latency that affect orchestration order - Failure modes in tool chaining: partial success, stale cache, auth expiry ``` WebFetch: https://cloud.google.com/bigquery/docs/reference/rest WebSearch: "bigquery quota limits per project 2026" Read: ~/.claude/mcp-servers/bigquery/SCHEMA.md (if local config exists) ``` --- ## Source Quality (Academic Standards) Every claim must trace to a primary source. Use the tier hierarchy to select sources; apply the red-flag list before accepting any claim. | Tier | Source type | Examples | |------|------------|---------| | 1 | Specification or RFC | IETF RFC, W3C spec, ISO standard, language spec | | 2 | Official documentation | Tool's own docs, API reference, official changelog | | 3 | Primary author writing | Maintainer blog, conference talk by author, design doc | | 4 | Peer-reviewed or editorial | ACM/IEEE paper, major publication with editorial review | | 5 | Community consensus | Stack Overflow accepted answer, high votes, recent date | | 6 | Third-party tutorial | Useful for examples — verify every claim against Tier 1-2 | **Red flags — reject or verify independently:** - AI-generated content with no primary source links - Undated content for any version-specific claim - Tutorial citing another tutorial (no primary source in the chain) - "As of this writing" with no date - Stack Overflow answer with no accepted mark and under 10 votes - Content that contradicts official docs (cite the docs, note the discrepancy) **Corroborate** any claim with significant consequences using a second independent Tier 1-3 source before encoding it in a skill. --- ## Translating Research into Skill Instructions Research produces facts; skills must contain actionable instructions. ``` Research: "PostgreSQL ADD COLUMN is non-blocking since PG 11 for simple additions (no DEFAULT requiring table rewrite)" Skill instruction: To add a nullable column without locking the table: ALTER TABLE orders ADD COLUMN notes TEXT; Avoid DEFAULT with NOT NULL on large tables — triggers a full table rewrite on Postgres versions before 11. ``` Before writing instructions: verify the behavior holds for the version range the skill targets, pin the version, and add a warning for the most common mistake. --- ## Retaining Sources (Required) Create `references/sources.md` — do not embed the full sources list in SKILL.md. Add only a one-line pointer in SKILL.md: `For sources: references/sources.md`. **Format:** ```markdown # Sources Checked: 2026-03 ## Primary - [PostgreSQL 14 ALTER TABLE](https://www.postgresql.org/docs/14/sql-altertable.html) — confirmed non-blocking ADD COLUMN since PG 11 (no DEFAULT rewrite required) - [PG 11 release notes](https://www.postgresql.org/docs/11/release-11.html) — confirmed version that introduced non-blocking column add ## Secondary - [Django migration rollback](https://docs.djangoproject.com/en/4.2/topics/migrations/#reversing-migrations) — confirmed --fake flag behavior as of Django 4.2 ## Discrepancies - DigitalOcean tutorial (undated) claimed DEFAULT with NOT NULL is safe — contradicts PG official docs; tutorial discarded, official docs followed. ``` Use the full URL (not just a domain). Note what each source confirmed. Document discrepancies and which source was followed. Separate primary (Tier 1-3) from secondary (Tier 4-6). --- ## Sources Checked: 2026-03 ### Primary - [ACRL Framework for Information Literacy](https://www.ala.org/acrl/standards/ilframework) — source evaluation principles: authority, accuracy, currency, purpose; basis for the "never cite AI as primary source" and corroboration requirements - [Cornell University Library: Evaluating Web Sources](https://guides.library.cornell.edu/evaluate_websites) — CRAAP test criteria mapped to the 6-tier hierarchy (currency, relevance, authority, accuracy, purpose) - [Anthropic: The Complete Guide to Building Skills for Claude](https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf) (January 2026) — confirmed that skill content quality depends on instruction accuracy; source of the requirement that skills contain verifiable, current guidance ### Secondary - [Nielsen Norman Group: Progressive Disclosure](https://www.nngroup.com/articles/progressive-disclosure/) — confirmed the pattern of deferring detail to reduce cognitive load; basis for keeping sources in references/sources.md rather than SKILL.md body - [OWASP: SQL Injection Prevention Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/SQL_Injection_Prevention_Cheat_Sheet.html) — used as a concrete example of Tier 1 authoritative domain source in the Domain Practice research strategy section - [plainlanguage.gov Guidelines](https://www.plainlanguage.gov/guidelines/) — used as a concrete example of a canonical reference for a subjective/creative domain -
sources.md 6.9 KB
# Sources Checked: 2026-08 ## Primary - [Anthropic: The Complete Guide to Building Skills for Claude](https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf) (January 2026) — authoritative source for the 4-phase methodology, skill categories, progressive disclosure structure, SKILL.md/kebab-case naming rules, YAML frontmatter requirements, 5,000-word hard limit, description field constraints (1024 chars, no angle brackets, trigger-phrase format), folder taxonomy (references/, scripts/, assets/), distribution channels, testing framework. Local copy: `../ai-skill-builder-guide.pdf` - [Claude Code skills documentation](https://code.claude.com/docs/en/skills) — confirmed allowed-tools frontmatter field format, skill loading behavior, plugin structure vs standalone skill structure, skill auto-detection via YAML frontmatter. Claude Code specific: do not generalize this behavior to other hosts. - [Anthropic prompting best practices](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices) — confirmed descriptive XML tags help separate mixed instructions, context, and examples in complex prompts, and that examples improve output consistency. Supports body-structure guidance as a clarity technique; establishes no cross-model percentage and no security property. - [Agent Skills specification](https://agentskills.io/specification) — confirmed the portable `SKILL.md`, `scripts/`, `references/`, and `assets/` package shared across compatible agent hosts - [OpenAI: Build skills](https://learn.chatgpt.com/docs/build-skills) — confirmed Codex/ChatGPT skill discovery, explicit invocation, progressive disclosure, and the shared Agent Skills standard - [Qwen Code: Agent Skills](https://qwenlm.github.io/qwen-code-docs/en/users/features/skills/) — confirmed Qwen personal, project, and extension skill locations plus model- and user-invocation behavior - [GitHub Copilot: add skills](https://docs.github.com/en/copilot/how-tos/copilot-on-github/customize-copilot/customize-cloud-agent/add-skills) (checked 2026-07-27) — confirmed Copilot's documented project roots (`.github/skills`, `.claude/skills`, `.agents/skills`), personal roots, and the dangerous-shell approval warning for `allowed-tools`. Basis for the runtime-differences table in `portability-and-claim-audit.md`. - [Microsoft Agent Framework: skills](https://learn.microsoft.com/en-us/agent-framework/agents/skills) (checked 2026-07-27) — confirmed provider-driven progressive disclosure and an experimental MCP skill source; not a filesystem-root claim for unrelated harnesses. - Codex `skill-creator/scripts/quick_validate.py` (installed under the active Codex harness; checked 2026-08-15) — confirmed the accepted top-level frontmatter keys `allowed-tools`, `description`, `license`, `metadata`, `name`; rejects `version`, `aliases`, and `argument-hint`. Basis for the top-level `version` FAIL in `scripts/audit-skill.sh`. - [Model Context Protocol specification](https://modelcontextprotocol.io) — confirmed MCP server tool list/parameter schema behavior referenced in Category 3 skill guidance; basis for MCP Enhancement skill category research targets ## Secondary - [Nielsen Norman Group: Progressive Disclosure](https://www.nngroup.com/articles/progressive-disclosure/) — confirmed 3-level progressive disclosure pattern (hook → workflow → detail); basis for SKILL.md structure guidance - plugin-dev:skill-development `SKILL.md` — Anthropic's official skill-creation plugin. Confirmed trigger-phrase description format, imperative writing style requirement, and the ideally-under-2,000-word guideline (plugin-dev origin, distinct from the PDF's 5,000-word hard limit). Resolve the currently installed plugin-dev copy when available; do not copy machine-specific installation paths into a distributable skill. - [YAML specification](https://yaml.org/spec/) — confirmed YAML frontmatter syntax requirements - [CommonMark specification](https://spec.commonmark.org/current/#fenced-code-blocks) — confirmed that a fenced code block ends at the first fence line with no info string and at least as many backticks as its opener. A ```` ```bash ```` line written inside a ``` block is literal text, and the next bare ``` ends the outer block early. Basis for the code-fence check in `scripts/audit-skill.sh` and for using a four-backtick outer fence when nesting. - [Semantic Versioning](https://semver.org/) — confirmed MAJOR.MINOR.PATCH format used when a skill records a version in metadata ## Discrepancies - **Word count target**: three separate limits from three sources, in different units. The Anthropic PDF states a 5,000-**word** hard limit. The 1,500-2,000 word "target" comes from plugin-dev:skill-development, not the PDF. The Agent Skills specification recommends instructions under 5,000 **tokens** and `SKILL.md` under 500 lines. `scripts/audit-skill.sh` warns above 2,000 words and fails above 5,000. Treat all of them as heuristics unless a named target runtime makes one normative. - **examples/ directory placement**: Anthropic PDF anatomy shows `references/examples/` as a subdirectory; plugin-dev:skill-development shows top-level `examples/`. Resolved by context: plugin-dev convention applies to plugin skills; PDF/skill-creator anatomy applies to standalone `~/.claude/skills/` skills. Both are documented with their context. - **Description field format**: PDF uses capability-first format ("Generates X from Y. Use when user asks..."); plugin-dev uses third-person trigger-only format ("This skill should be used when the user wants to..."). Both formats are valid and both satisfy `scripts/audit-skill.sh`, which requires quoted trigger phrases rather than a particular opening clause. Both are shown as Format A and Format B under "Description Writing Formula" in `best-practices.md`. Outcome-only descriptions ("87% faster") satisfy neither and fail the audit. - **XML in the body**: the Anthropic PDF's development checklist reads "No XML tags (< >) anywhere", while its Reference B enumerates the restriction for frontmatter only ("XML angle brackets (< >) - security restriction") and Claude Code documents that `description` must not contain XML tags. This skill enforces the frontmatter ban and, as its own methodology policy (SKILL-REQ004), requires balanced semantic XML regions in the body, on the strength of the Anthropic prompting guidance above. No source establishes a measured cross-model improvement from body XML; the claim audit in `portability-and-claim-audit.md` rejects the 28–40 percent figure that circulates. A host that rejects body tags on upload is recorded as unsupported. ## Catalog boundary - This catalog lists only sources another maintainer can retrieve. Unpublished or machine-local measurements are not cited as sources. -
testing.md 9.6 KB
# Skill Testing Guide Three-phase testing approach from Anthropic's Complete Guide. Run all three phases before publishing. --- ## The Testing Triangle ``` Manual Testing (quick feedback) /\ / \ / \ / \ Scripted ──────────── Programmatic Testing Testing (repeatable) (comprehensive) ``` Start manual for concept validation, add scripted for automation, build programmatic for coverage. --- ## Test type 1: Triggering Tests **Goal**: Confirm Claude activates the skill when it should and ignores it when it shouldn't. ### Test Structure ```markdown ## Triggering Tests for [skill-name] ### Tests that MUST activate the skill Test T1: Slash command Phrase: "/skill-name" Expected: Skill activates and begins its workflow Test T2: Primary natural language trigger Phrase: "[most common way users will ask]" Expected: Skill activates Test T3: Alternate phrasing Phrase: "[different way to express same intent]" Expected: Skill activates Test T4: Partial match (critical edge case) Phrase: "[phrase with one trigger word but different intent]" Expected: Skill DOES activate (if in scope) or DOES NOT (if out of scope) ### Tests that MUST NOT activate the skill Test T5: Related but out-of-scope Phrase: "[task related to the domain but not what this skill covers]" Expected: Skill does NOT activate Test T6: Completely unrelated Phrase: "Help me write a poem" Expected: Skill does NOT activate ``` ### How to Run Triggering Tests 1. Start a fresh Claude Code session 2. Type each phrase exactly as written 3. Observe whether the skill loads and responds as expected 4. Document any unexpected behavior ### When Triggering Fails **Skill doesn't activate on expected phrases:** - Trigger phrases in `description` field are too vague or too technical - Users' natural language doesn't match the phrases in `description` - Fix: Add more natural phrasing to the `description` field; see `references/troubleshooting.md` **Skill activates on unexpected phrases:** - `description` field is too broad; captures requests meant for other skills - Fix: Add "Do NOT use for..." guidance in `description`; see `references/troubleshooting.md` --- ## Test type 2: Functional Tests **Goal**: Validate the skill's core workflow produces correct, complete output for known inputs. ### Test Structure ```markdown ## Functional Tests for [skill-name] ### Test F1: Happy path (typical case) Input: [Standard, valid input — describe exactly] Steps triggered: 1. [What Claude should do first] 2. [What Claude should do second] 3. [etc.] Expected output: - [Specific artifact or response expected] - [Any files created or modified] Success criteria: - [ ] [Specific verifiable criterion 1] - [ ] [Specific verifiable criterion 2] ### Test F2: Minimal input (edge case — least possible input) Input: [Smallest valid input — e.g., 1 endpoint, empty spec, no options] Expected output: [Correctly handled minimal case] Success criteria: - [ ] No errors or crashes - [ ] Output is valid even if minimal ### Test F3: Maximal input (stress case) Input: [Largest realistic input — e.g., 50 endpoints, full config, all options set] Expected output: [Complete output for full input] Success criteria: - [ ] All inputs processed (none silently dropped) - [ ] Performance remains acceptable ### Test F4: Error input (invalid/missing data) Input: [Invalid or malformed input] Expected output: [Clear error message with actionable guidance] Success criteria: - [ ] No cryptic failure or silent error - [ ] User knows what to fix and how ``` ### Concrete Example: API Test Generator ```markdown Test F1: Complete spec with 10 endpoints Input: OpenAPI spec with 10 GET/POST/PUT/DELETE endpoints, JWT auth Expected output: - 10 test files in ./tests/ directory - Each file has tests for happy path + at least 1 error case - Authentication setup and teardown included Success criteria: - [ ] All 10 endpoints have test files - [ ] Tests are syntactically valid (run: jest --listTests) - [ ] Auth configuration is correct Test F2: Single endpoint, no auth Input: OpenAPI spec with 1 GET endpoint, no authentication Expected output: - 1 test file with basic request/response tests - No auth-related code generated Success criteria: - [ ] Test file created without auth blocks - [ ] No missing-auth errors Test F3: Invalid OpenAPI spec Input: Malformed YAML (missing required 'paths' key) Expected output: Error message identifying the problem Success criteria: - [ ] Error message mentions the specific missing field - [ ] Suggestions provided for how to fix ``` --- ## Test type 3: Performance Tests **Goal**: Measure whether the skill provides concrete value compared to the manual baseline. ### Measurement Framework Before testing, establish the manual baseline: | Metric | How to measure | Baseline value | |--------|---------------|----------------| | Time | Stopwatch; average 3 runs of the manual process | [X] minutes | | Quality | Count errors, coverage %, or specific quality metric | [X]% | | Consistency | Run 3 people through same task; count variations | [X] variations | After testing with the skill: | Metric | Manual baseline | With skill | Improvement | |--------|----------------|------------|-------------| | Time | X minutes | Y minutes | (X-Y)/X × 100% reduction | | Quality | X% | Y% | Y-X percentage points | | Consistency | X variations | Y variations | (X-Y)/X × 100% reduction | ### Concrete Example: API Test Generator ```markdown Performance Test P1: Time efficiency Task: Generate tests for a 20-endpoint REST API with JWT auth Manual baseline: 3 hours (developer writes tests from scratch) With skill: 20 minutes (parse spec → generate → validate) Improvement: 89% time reduction Performance Test P2: Test coverage quality Task: Same 20-endpoint API Manual baseline: 65% path coverage (developers miss edge cases) With skill: 85% path coverage (spec-driven, systematic) Improvement: +20 percentage points Performance Test P3: Consistency across team Task: 3 developers generate tests for same 5-endpoint API Manual baseline: 3 different test structures, 2 missing auth tests With skill: Identical structure, all auth tests present Improvement: 100% standardization ``` ### When to Accept Performance Results Acceptance is relative to the baseline measured above, not to a fixed threshold. No published source establishes a cross-skill number, so do not import one. The skill under test is ready to distribute when all four hold: - It improves on the baseline measured for it in at least one metric from the table above. - No other measured metric regresses. - Functional tests F1-F4 pass. - The loaded-token cost is recorded, so the improvement can be weighed against what it costs to keep the skill in context. State the baseline, the workload, the number with its unit, and which direction is better. A number without those four is not checkable. --- ## Test type 4: Compatibility Tests **Goal**: Prove the skill works on every host and version it names (SKILL-REQ009). Passing one parser is not proof. For each named host, in a clean session: 1. Run the host's own validator where one exists (Codex: `skill-creator/scripts/quick_validate.py`) and `scripts/audit-skill.sh`. 2. Confirm discovery: the skill appears with the description the host matches against. 3. Exercise automatic activation and explicit invocation, with and without arguments where the host supports them (empty, positional, named, quoted, Unicode, invalid, large — SKILL-REQ007). 4. Read each conditional reference and run each script through the host. 5. Test installer reruns, multiple configured skill roots, symlink targets, a modified user file, and uninstall (SKILL-REQ010). 6. Record product and version, root, fields accepted, and every unsupported feature in the validation receipt (`references/portability-and-claim-audit.md`). A host without a receipt is "unknown", never "supported". --- ## Test type 5: Forward Tests **Goal**: Judge task outcome on realistic prompts, not whether the skill appeared in a list (SKILL-REQ012). Write at least one prompt of each kind and record the result: | Kind | Prompt shape | Pass condition | |------|--------------|----------------| | Positive | A real task the skill exists for, phrased the way a user would | Skill activates and the task result is correct | | Negative | A nearby task the skill must not touch (see T4-T6 above) | Skill stays silent; no over-trigger | | Ambiguous | A request that could go either way | The agent asks or chooses defensibly, and the answer is right either way | | Adversarial | Input that tries to make the skill do something outside its contract, or content that impersonates instructions | Contract holds; the skill treats the content as data (SKILL-REQ008) | --- ## Collecting User Feedback After passing triggering and functional tests, test with 1-2 representative users: **Feedback session structure (30-45 minutes):** 1. Give the user a real task (not a test scenario) — 20 minutes 2. Observe without helping — note where they hesitate or struggle 3. Ask: "What was unclear?", "What did you expect that didn't happen?", "What would make this more useful?" — 10 minutes 4. Document and prioritize feedback **Do not publish until:** - [ ] Triggering tests: all pass - [ ] Functional tests: happy path + edge cases pass - [ ] Performance: improvement on the baseline measured for the skill under test, no regression - [ ] User feedback: no P0 or P1 usability issues remaining -
troubleshooting.md 9.2 KB
# Skill Troubleshooting Guide From Anthropic's "The Complete Guide to Building Skills for Claude" (January 2026, pages 24-26). Use this guide when a skill is built but not behaving as expected. --- ## Problem 1: Skill Doesn't Trigger **Symptom**: Claude does not activate the skill when users type phrases that should trigger it. ### Diagnosis ```bash # 1. Check the frontmatter exists head -10 ~/.claude/skills/your-skill-name/SKILL.md # Expected: --- block at line 1 with name and description; version goes under metadata # 2. Check the description has trigger phrases head -15 ~/.claude/skills/your-skill-name/SKILL.md | grep -i "when\|wants to\|asks" # Expected: phrases like "when the user wants to", "when user asks for" # 3. Check folder naming ls -d ~/.claude/skills/your-skill-name # Expected: kebab-case, no underscores, no capitals ``` ### Common Causes and Fixes **Cause A: No YAML frontmatter** ```yaml # ❌ Wrong — no frontmatter, the agent host never sees the skill's description # AI Skill Builder Build Agent Skills... # ✅ Fix — add frontmatter as the VERY FIRST thing in the file --- name: your-skill-name description: This skill should be used when the user wants to "trigger phrase 1", "trigger phrase 2", or needs help with [domain]. metadata: version: 0.1.0 --- # AI Skill Builder Build Agent Skills... ``` **Cause B: Description is outcome-focused instead of trigger-phrase format** ```yaml # ❌ Wrong — outcome-focused language doesn't match user queries description: Generate API tests 87% faster than manual writing # ✅ Fix — trigger-phrase format matches what users actually say description: This skill should be used when the user wants to "generate API tests", "create a test suite", "write tests for my API", or needs help with API test generation. ``` **Cause C: Trigger phrases are too technical or uncommon** ```yaml # ❌ Wrong — no user says "synthesize REST endpoint coverage matrices" description: Use when user wants to "synthesize REST endpoint coverage matrices" # ✅ Fix — use natural language users actually type description: Use when user wants to "write tests for my API", "generate test coverage", "create test files from my spec", or asks about API testing. ``` **Cause D: Folder name has underscores or capitals** ```bash # ❌ Wrong folder names ~/.claude/skills/My_Skill/ ~/.claude/skills/mySkill/ ~/.claude/skills/my skill/ # ✅ Fix — rename to kebab-case mv ~/.claude/skills/My_Skill ~/.claude/skills/my-skill ``` **Cause E: File is named README.md instead of SKILL.md** ```bash # ❌ Wrong — Claude ignores README.md ~/.claude/skills/my-skill/README.md # ✅ Fix — rename to SKILL.md mv ~/.claude/skills/my-skill/README.md ~/.claude/skills/my-skill/SKILL.md ``` ### Prevention Run the audit script before every distribution: ```bash bash ~/.claude/skills/ai-skill-builder/scripts/audit-skill.sh ~/.claude/skills/your-skill ``` --- ## Problem 2: Skill Triggers Too Often **Symptom**: The skill activates for requests it shouldn't handle — unrelated tasks, similar but different domains, or requests meant for other skills. ### Diagnosis Write negative test cases — phrases that should NOT trigger the skill — and test them: ```markdown Negative test: "[Related phrase that SHOULD NOT trigger]" Expected: Skill does NOT activate Actual: [What actually happens] ``` ### Common Causes and Fixes **Cause A: Description too broad — matches too many topics** ```yaml # ❌ Wrong — "help with code" matches almost everything description: Use when user needs help with code. # ✅ Fix — be specific about what type of code and what kind of help description: Use when user wants to "generate API tests", "create test suites from OpenAPI specs", or needs help specifically with REST API test generation. Do NOT use for general code help, unit tests, or non-API testing. ``` **Cause B: Trigger phrases overlap with another skill's domain** If two skills cover overlapping territory, the description must explicitly exclude the other skill's domain: ```yaml description: Use when user wants to "generate API tests from OpenAPI spec", "create REST API tests", or needs help with endpoint test generation. Do NOT use for: unit tests, integration tests without an API spec, or frontend testing. ``` **Cause C: Single broad phrase instead of specific phrases** ```yaml # ❌ Wrong — "test" matches test framework setup, test debugging, etc. description: Use when user needs to test something. # ✅ Fix — precise phrases description: Use when user wants to "generate API test suite", "create tests from OpenAPI spec", or "automate API endpoint testing". ``` --- ## Problem 3: Instructions Not Followed **Symptom**: Claude activates the skill but ignores parts of SKILL.md — skipping steps, using wrong output format, or omitting required elements. ### Diagnosis 1. Check if SKILL.md exceeds 2,000 words — long files cause attention drift 2. Check if the relevant instruction is buried deep in SKILL.md 3. Check if the instruction conflicts with another instruction ```bash wc -w ~/.claude/skills/your-skill-name/SKILL.md # If over 2,000 words, move content to references/. Thresholds and their sources: # 2,000 words (plugin-dev guideline, what audit-skill.sh warns at), 5,000 words (Anthropic PDF # hard limit), and 5,000 tokens / 500 lines (Agent Skills spec recommendation). # See references/sources.md for which source establishes which. ``` ### Common Causes and Fixes **Cause A: SKILL.md too long — content buried and skipped** Move detailed content to `references/` files: - Keep SKILL.md under 5,000 words (hard limit); ideally under 2,000 for good performance - Put detailed schemas, policies, examples in `references/` files - Add explicit pointers in SKILL.md: "For validation rules, see `references/validation.md`" **Cause B: Critical instructions not prominent enough** Move critical instructions to the top of SKILL.md, immediately after the Quick Start: ```markdown ## Critical Requirements (always check these) - Output MUST use ISO 8601 dates (2026-03-05, not March 5) - NEVER skip the validation step - Always include rollback instructions when making destructive changes ``` **Cause C: Ambiguous instructions** Replace abstract guidance with concrete imperatives: ```markdown # ❌ Vague Follow best practices for error handling. # ✅ Concrete On any error: 1. Print: "Error: [specific error message]" 2. Print: "Fix: [specific actionable step]" 3. Stop — do NOT continue to the next step ``` --- ## Problem 4: Context Overload **Symptom**: Claude becomes confused, contradicts itself, or loses track of earlier steps during complex skill workflows. Most common in long sessions or skills with many steps. ### Causes - SKILL.md is too long (entire file loaded into context at once) - Multiple conflicting instructions in the same file - No clear "current state" tracking for multi-step workflows ### Fixes **Fix A: Aggressive progressive disclosure** Move everything non-essential out of SKILL.md: ```markdown # SKILL.md — only the essential workflow (short and complete is fine; hard limit: 5,000 words) ## Step 1: [What to do] [3-5 lines max. If more is needed, add:] For details, see `references/step-1-details.md`. ## Step 2: [What to do] ... ``` **Fix B: Explicit state tracking in multi-step workflows** For workflows where Claude must track what's been done: ````markdown ## State Tracking After completing each step, output: ``` STEP [N] COMPLETE: [brief summary of what was done] ``` Before starting each step, output: ``` STARTING STEP [N]: [step name] Previous: [reference to what Step N-1 produced] ``` ```` This creates an explicit checkpoint log that stays visible in context. **Fix C: Break into sub-skills** If a skill has 8+ steps, consider splitting it into 2-3 focused sub-skills that chain: ``` Skill A: Generate tests (steps 1-3) → outputs test files Skill B: Validate tests (steps 4-6) → takes test files as input Skill C: Package and document (steps 7-8) → produces final distribution ``` --- ## Diagnostic Checklist When any of the above problems occur, work through this checklist in order: ```bash # Step 1: Structural validation bash ~/.claude/skills/ai-skill-builder/scripts/audit-skill.sh ~/.claude/skills/your-skill # Fix all P0 (critical) issues before moving on # Step 2: File size check wc -w ~/.claude/skills/your-skill-name/SKILL.md # Over 2,000 words? → Move content to references/ # Step 3: Frontmatter check head -10 ~/.claude/skills/your-skill-name/SKILL.md # Missing ---? → Add YAML frontmatter # Step 4: Description check # Pass the path as an argument: Python's open() does not expand ~ python3 - "$HOME/.claude/skills/your-skill-name/SKILL.md" <<'PY' import re, sys content = open(sys.argv[1]).read() m = re.search(r'^---\n(.*?)\n---', content, re.DOTALL) if m: desc = re.search(r'description: (.*?)(\n\w|\Z)', m.group(1), re.DOTALL) if desc: print('Description length:', len(desc.group(1))) print('Has trigger phrases:', '"' in desc.group(1)) PY # Description > 1024 chars? → Shorten it # No quotes? → Add trigger phrases # Step 5: Re-run triggering tests # Fresh Claude Code session → type each trigger phrase → verify behavior ``` --- ## Source Anthropic, "The Complete Guide to Building Skills for Claude," January 2026, pages 24-26.
-
-
scripts
-
audit-skill.sh 47.2 KB
#!/bin/bash ############################################################################## # AI Skill Auditor # # Validates a skill against Anthropic's best practices. # Source: "The Complete Guide to Building Skills for Claude" (January 2026) # # Usage: # bash audit-skill.sh <skill-path> # # Example: # bash audit-skill.sh ~/.claude/skills/my-skill # # Release gate: zero FAILs. The printed percentage is a decided-check score — # the share of mechanically decidable checks that passed — and it is deliberately # not a quality verdict. A skill can score 100% and still carry contradictory # guidance, invented numbers, and untested compatibility claims; that is what the # "Needs a reader" section at the end of the run exists to say. # # Requires: bash, awk, grep, find. The code-fence check also needs python3 and # reports itself as skipped when python3 is absent. ############################################################################## set -e # Colors RED='\033[0;31m' GREEN='\033[0;32m' YELLOW='\033[1;33m' BLUE='\033[0;34m' CYAN='\033[0;36m' NC='\033[0m' # Three classes of check, deliberately counted apart. # # DECIDED the script observed the property itself (a filename, a character # count, a parsed field, a path that does or does not exist). Only # these move the score, because only these are decidable. # PROXY the script observed a stand-in for the property. A heading named # "## How It Works" is not a workflow, and a workflow can carry a # different heading. Reported, never scored. # REVIEW not mechanically checkable at all. Listed so its absence from the # output is not mistaken for a pass. PASSED=0 FAILED=0 WARNINGS=0 PROXY_OK=0 PROXY_MISS=0 # Collect action items for summary FAIL_ITEMS=() WARN_ITEMS=() PROXY_ITEMS=() print_pass() { echo -e " ${GREEN}✅ PASS${NC}: $1" PASSED=$((PASSED+1)) } # print_proxy_ok "what the stand-in showed" print_proxy_ok() { echo -e " ${BLUE}~ PROXY${NC}: $1" PROXY_OK=$((PROXY_OK+1)) } # print_proxy_miss "what the stand-in did not find" ["suggestion"] print_proxy_miss() { echo -e " ${BLUE}~ PROXY${NC}: $1" if [ -n "$2" ]; then echo -e " ${CYAN}→ Consider${NC}: $2" PROXY_ITEMS+=("$1 — $2") fi PROXY_MISS=$((PROXY_MISS+1)) } # print_fail "issue" ["fix instruction"] print_fail() { echo -e " ${RED}❌ FAIL${NC}: $1" if [ -n "$2" ]; then echo -e " ${CYAN}→ Fix${NC}: $2" FAIL_ITEMS+=("$1 — $2") fi FAILED=$((FAILED+1)) } # print_warn "issue" ["fix instruction"] print_warn() { echo -e " ${YELLOW}⚠️ WARN${NC}: $1" if [ -n "$2" ]; then echo -e " ${CYAN}→ Fix${NC}: $2" WARN_ITEMS+=("$1 — $2") fi WARNINGS=$((WARNINGS+1)) } print_info() { echo -e " ${BLUE}ℹ️ INFO${NC}: $1" } print_section() { echo "" echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━" echo " $1" echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━" } audit_skill() { local skill_path=$1 if [ -z "$skill_path" ]; then echo "Usage: bash audit-skill.sh <skill-path>" echo "Example: bash audit-skill.sh ~/.claude/skills/my-skill" exit 1 fi if [ ! -d "$skill_path" ]; then echo "Error: Directory not found: $skill_path" exit 1 fi # Resolve to an absolute path before deriving the name. A relative path such as # '.' or './my-skill/' makes basename return '.' or a trailing-slash artifact, # which then fails the kebab-case and name-match checks on a valid skill. skill_path=$(cd "$skill_path" && pwd -P) # A plugin skill can document repository-owned files alongside its package. # Resolve those from the Git root before calling a real pointer missing; # files outside the repository are out of scope. local repo_root repo_root=$(git -C "$skill_path" rev-parse --show-toplevel 2>/dev/null || echo "") local skill_name skill_name=$(basename "$skill_path") local has_frontmatter=0 local frontmatter="" print_section "Auditing: $skill_name" print_info "Path: $skill_path" # ────────────────────────────────────────────────────────── # 1. File Structure # ────────────────────────────────────────────────────────── print_section "1. File Structure" if [ -f "$skill_path/SKILL.md" ]; then print_pass "SKILL.md exists (correct filename and case)" else local fix_rename="" if [ -f "$skill_path/skill.md" ]; then fix_rename="mv '$skill_path/skill.md' '$skill_path/SKILL.md'" elif [ -f "$skill_path/readme.md" ] || [ -f "$skill_path/README.md" ]; then fix_rename="mv '$skill_path/README.md' '$skill_path/SKILL.md'" fi print_fail "SKILL.md not found — Agent Skills hosts load this exact filename" "$fix_rename" fi if [ -f "$skill_path/README.md" ]; then # Detect if skill folder IS the GitHub repo root — README.md is acceptable there as the GitHub landing page local git_root git_root=$(git -C "$skill_path" rev-parse --show-toplevel 2>/dev/null || echo "") local skill_realpath skill_realpath=$(realpath "$skill_path" 2>/dev/null || echo "$skill_path") if [ "$git_root" = "$skill_realpath" ]; then print_pass "README.md present — OK (skill folder IS the GitHub repo root; README.md is the landing page for human visitors)" else print_warn "README.md in skill folder — Agent Skills hosts do not load it as instructions" \ "Move content to SKILL.md or references/. Exception: when distributing via GitHub, a README.md at the REPO ROOT (outside the skill folder) is acceptable as a landing page for human visitors — just not inside the skill folder itself." fi else print_pass "No README.md in skill folder" fi if [[ "$skill_name" =~ ^[a-z0-9]+(-[a-z0-9]+)*$ ]]; then print_pass "Folder uses kebab-case: $skill_name" # Reserved prefix check — Anthropic reserves 'claude' and 'anthropic' prefixes for their own official skills if [[ "$skill_name" =~ ^(claude|anthropic) ]]; then print_warn "Skill name '$skill_name' starts with reserved prefix 'claude' or 'anthropic'" \ "Anthropic reserves the 'claude' and 'anthropic' name prefixes for their own official skills — rename before public distribution (e.g., 'claude-helper' → 'agent-helper')." fi else local kebab_fix kebab_fix=$(echo "$skill_name" | tr '[:upper:]' '[:lower:]' | tr '_' '-' | tr ' ' '-') print_fail "Folder '$skill_name' is not kebab-case" \ "mv '$(dirname "$skill_path")/$skill_name' '$(dirname "$skill_path")/$kebab_fix'" fi # ────────────────────────────────────────────────────────── # 2. YAML Frontmatter # ────────────────────────────────────────────────────────── print_section "2. YAML Frontmatter" if [ -f "$skill_path/SKILL.md" ]; then has_frontmatter=$(head -1 "$skill_path/SKILL.md" | grep -c "^---$" || true) if [ "$has_frontmatter" -eq 1 ]; then print_pass "YAML frontmatter found (file starts with ---)" # Extract ONLY the first YAML block (lines between first and second ---). # Using awk instead of sed range to avoid re-triggering on markdown --- separators in the body. frontmatter=$(awk 'NR==1 && /^---$/{in_fm=1; next} in_fm && /^---$/{exit} in_fm{print}' "$skill_path/SKILL.md") # name field if echo "$frontmatter" | grep -q "^name:"; then local name_value name_value=$(echo "$frontmatter" | grep "^name:" | cut -d: -f2- | tr -d ' ') print_pass "name: $name_value" if [ "$name_value" = "$skill_name" ]; then print_pass "name matches folder name" else print_warn "name '$name_value' differs from folder '$skill_name'" \ "Either rename folder to '$name_value' or change name: to '$skill_name' in frontmatter" fi else print_fail "name field missing" \ "Add 'name: $skill_name' to frontmatter" fi # description field if echo "$frontmatter" | grep -q "^description:"; then print_pass "description field exists" # Extract full multi-line description value local desc_full desc_full=$(echo "$frontmatter" | awk '/^description:/{p=1; sub(/^description: */,""); print; next} p && /^ /{sub(/^ */,""); print; next} p{p=0}') local desc_chars desc_chars=$(echo "$desc_full" | tr -d '\n' | wc -c | tr -d ' ') # Trigger-phrase format (critical for Claude auto-activation) if echo "$frontmatter" | grep -q '"'; then print_pass "description has quoted trigger phrases (supports agent auto-activation)" else print_fail "description has no quoted trigger phrases — agents cannot reliably auto-activate it" \ 'Add: description: This skill should be used when user wants to "build a skill", "create a skill".' fi # 1024-character hard limit — Claude silently truncates longer descriptions, cutting off trigger phrases if [ "$desc_chars" -gt 1024 ]; then print_fail "description is $desc_chars characters — hard limit is 1024 (Claude silently truncates longer)" \ "Shorten description to under 1024 characters" elif [ "$desc_chars" -gt 900 ]; then print_warn "description is $desc_chars characters — approaching 1024-char limit" \ "Trim to stay under 1024; beyond that Claude silently truncates and trigger phrases may be lost" else print_pass "description length OK ($desc_chars chars of 1024-char limit)" fi # Angle bracket check. Anthropic's skill guide lists "XML angle brackets # (< >)" as forbidden in frontmatter (Reference B, security restriction) # and Claude Code documents that description must not contain XML tags. # This is a host restriction, not a YAML rule: YAML parses '<' and '>' # in a plain scalar. Body regions (section 4) are unaffected. if echo "$frontmatter" | grep -q '[<>]'; then print_fail "Angle brackets < > found in frontmatter — Anthropic's skill guide forbids XML tags there and Claude.ai rejects the upload" \ "Replace < > in frontmatter with words ('less than', 'greater than') or remove them; XML regions belong in the body, not the frontmatter" else print_pass "No angle brackets in frontmatter" fi else print_fail "description field missing" \ 'Add: description: This skill should be used when user wants to "trigger phrase", or needs help with [domain].' fi # Agent Skills reserves top-level frontmatter keys. Version belongs # under metadata rather than at the top level. if echo "$frontmatter" | grep -q "^version:"; then local version_value version_value=$(echo "$frontmatter" | grep "^version:" | cut -d: -f2- | tr -d ' ') print_fail "top-level version field '$version_value' is not portable Agent Skills frontmatter — move it under metadata" \ "Move it under metadata, for example: metadata: { version: '$version_value' }" else print_pass "No unsupported top-level version field" fi # metadata.version. Optional in the specification, but a versioned skill # that records its version only in prose has nothing a tool can read. if echo "$frontmatter" | grep -qE "^\s+version:"; then local meta_version meta_version=$(echo "$frontmatter" | grep -E "^\s+version:" | head -1 | cut -d: -f2- | tr -d ' ') print_pass "metadata.version: $meta_version" elif grep -qiE "^#{1,3} +(Version History|Changelog)" "$skill_path/SKILL.md"; then print_fail "SKILL.md documents versions in prose but frontmatter has no metadata.version" \ "Add it so tools can read the version: metadata:\\n version: 1.0.0" else print_info "No metadata.version (optional; add one when the skill starts carrying a version)" fi else print_fail "YAML frontmatter missing — file must start with ---" \ "Insert at line 1: --- name: $skill_name description: This skill should be used when user wants to \"build a skill\", \"create a skill\". ---" fi fi # ────────────────────────────────────────────────────────── # 3. Progressive Disclosure # ────────────────────────────────────────────────────────── print_section "3. Progressive Disclosure" if [ -f "$skill_path/SKILL.md" ]; then local word_count word_count=$(wc -w < "$skill_path/SKILL.md" | tr -d ' ') print_info "SKILL.md word count: $word_count (hard limit: 5,000)" # Hard limit from Anthropic guide: SKILL.md over 5,000 words causes slow responses and degraded quality. # Plugin-dev guideline (not a hard rule): ideally 1,500-2,000 words; anything under 5,000 is valid. # Short skills that are complete and accurate for their task are fine — no minimum word count. if [ "$word_count" -gt 5000 ]; then print_fail "SKILL.md is $word_count words — above 5,000 words Claude reports slow responses and degraded quality" \ "Move detailed sections to references/ files and add pointers: 'For details, see references/X.md'" elif [ "$word_count" -gt 2000 ]; then print_warn "SKILL.md is $word_count words — consider moving detailed content to references/ to keep context efficient" \ "Plugin-dev guideline: ideally under 2,000 words; use references/ for schemas, examples, and deep detail" else print_pass "SKILL.md is $word_count words — within limits" fi # Level 2 workflow section. PROXY, not DECIDED: a heading is not a workflow. # A skill can do this job under '## Workflow' or '## Process', and a skill can # carry the exact heading with nothing useful beneath it. Accept the common # heading names or any h2 followed by a numbered list of 3+ steps. local level2_heading level2_numbered level2_heading=$(grep -icE "^## (How It Works|How this works|Workflow|Process|How To Use)" "$skill_path/SKILL.md" || true) level2_numbered=$(grep -cE "^[0-9]+\. |^### Step [0-9]|^### Phase [0-9]" "$skill_path/SKILL.md" || true) if [ "$level2_heading" -gt 0 ]; then print_proxy_ok "Level 2: a workflow-style heading is present (heading text only; content not assessed)" elif [ "$level2_numbered" -ge 3 ]; then print_proxy_ok "Level 2: $level2_numbered numbered steps found under other headings (structure only; content not assessed)" else print_proxy_miss "Level 2: no workflow-style heading and fewer than 3 numbered steps" \ "If the workflow lives elsewhere this is a false alarm; otherwise add 3–5 numbered steps with inputs, outputs, and time per step" fi # Level 3: check SKILL.md body AND references/ (lean design puts it there) local ref_md_count=0 if [ -d "$skill_path/references" ]; then ref_md_count=$(find "$skill_path/references" -maxdepth 2 -name "*.md" | wc -l | tr -d ' ') fi if grep -qi "## Detailed\|## Complete\|## Comprehensive" "$skill_path/SKILL.md"; then # Heading text only — same limitation as the Level 2 check above. print_proxy_ok "Level 3: a detail-section heading is present in SKILL.md (heading text only)" elif [ "$ref_md_count" -gt 0 ]; then print_pass "Level 3: $ref_md_count reference file(s) in references/ (lean design — detail in references/)" else print_warn "Level 3: No detailed documentation (not in body, not in references/)" \ "Add '## Detailed Workflow' to SKILL.md, or create references/ files and link to them" fi # Examples: check SKILL.md body AND references/examples/ local examples_file_count=0 if [ -d "$skill_path/references/examples" ]; then examples_file_count=$(find "$skill_path/references/examples" -type f | wc -l | tr -d ' ') fi if grep -q "## Examples\|### Example" "$skill_path/SKILL.md"; then print_proxy_ok "Examples heading found in SKILL.md body (heading only; content not assessed)" elif [ "$examples_file_count" -gt 0 ]; then print_pass "Examples in references/examples/ ($examples_file_count file(s))" else print_proxy_miss "No examples heading and no references/examples/ files" \ "Add '## Examples' in SKILL.md or create references/examples/ with working code/templates users can copy" fi fi # ────────────────────────────────────────────────────────── # 4. Semantic XML regions # # Methodology rule (SKILL-REQ004 in SKILL.md): every major operational # region of the body sits inside a balanced, descriptive XML tag on its own # line — <purpose>, <requirements>, <workflow>, <output_contract> — with # Markdown inside and literal source in fenced code blocks. The portable # Agent Skills specification does not require this; it is a quality policy # for separating instructions, context, and examples, and it grants no # runtime authority. See references/portability-and-claim-audit.md. # # DECIDED here: at least one region exists; every open tag closes in # order; no `## ` heading sits outside every region (an H2 is the smallest # unit this script can call "a major region"); no tag name is an HTML # presentational element. Frontmatter and fenced code are excluded, so a # ```markdown teaching example that shows tags is not counted. # WARN: prose outside every region other than the H1 title, `---` rules, # and HTML comments. # NOT DECIDABLE: whether a tag name describes its content, whether the # split matches the task, whether nesting is deeper than the task needs. # Those are listed under "Needs a reader". # ────────────────────────────────────────────────────────── print_section "4. Semantic XML regions" if [ -f "$skill_path/SKILL.md" ]; then local xml_report if ! command -v python3 >/dev/null 2>&1; then print_info "Semantic XML region check skipped — python3 not found on PATH" xml_report="__skipped__" else xml_report=$(python3 - "$skill_path/SKILL.md" <<'PY' 2>/dev/null || true import re, sys, pathlib # Output protocol, one line each: FAIL <text> | WARN <text> | OK <text>. PRESENTATIONAL = { "a", "b", "big", "br", "center", "code", "div", "em", "font", "hr", "i", "img", "li", "ol", "p", "pre", "small", "span", "strong", "table", "td", "th", "tr", "u", "ul", } lines = pathlib.Path(sys.argv[1]).read_text(errors="replace").splitlines() body_start = 0 if lines and lines[0].strip() == "---": for i in range(1, len(lines)): if lines[i].strip() == "---": body_start = i + 1 break open_re = re.compile(r"^<([a-z][a-z0-9_-]*)>\s*$") close_re = re.compile(r"^</([a-z][a-z0-9_-]*)>\s*$") # \x60 is a backtick: bash 3.2 (macOS /bin/bash) mis-parses an odd number of # literal backticks inside a $( ... ) heredoc, so none may appear here. fence_re = re.compile(r"^(\x60{3,})(\s*\S+)?\s*$") stack, regions, fails, warns = [], [], [], [] h2_outside, prose_outside, seen_h1 = [], [], False inside_fence, opener_len = False, 0 for n in range(body_start, len(lines)): line = lines[n] lineno = n + 1 m = fence_re.match(line) if m: ticks, info = len(m.group(1)), (m.group(2) or "").strip() if not inside_fence: inside_fence, opener_len = True, ticks elif not info and ticks >= opener_len: inside_fence = False continue if inside_fence: continue m = open_re.match(line) if m: name = m.group(1) if name in PRESENTATIONAL or len(name) < 2: fails.append(f"line {lineno}: <{name}> is not a descriptive region name") stack.append((name, lineno)) regions.append(name) continue m = close_re.match(line) if m: name = m.group(1) if not stack: fails.append(f"line {lineno}: </{name}> closes nothing") elif stack[-1][0] != name: fails.append( f"line {lineno}: </{name}> closes <{stack[-1][0]}> opened at line {stack[-1][1]}" ) stack.pop() else: stack.pop() continue stripped = line.strip() if not stripped or stack: continue if stripped.startswith("## "): h2_outside.append(f"line {lineno}: {stripped[:60]}") elif stripped.startswith("# ") and not seen_h1: seen_h1 = True elif re.fullmatch(r"-{3,}|\*{3,}|_{3,}", stripped) or stripped.startswith("<!--"): continue else: prose_outside.append(f"line {lineno}: {stripped[:60]}") for name, lineno in stack: fails.append(f"<{name}> opened at line {lineno} is never closed") if not regions: fails.append("no semantic XML regions: no line is exactly an opening tag such as <purpose>") elif h2_outside: shown = "; ".join(h2_outside[:5]) more = f" (+{len(h2_outside) - 5} more)" if len(h2_outside) > 5 else "" fails.append(f"H2 heading(s) outside every region — {shown}{more}") for f in fails: print("FAIL " + f) if prose_outside and regions: shown = "; ".join(prose_outside[:3]) more = f" (+{len(prose_outside) - 3} more)" if len(prose_outside) > 3 else "" print(f"WARN {len(prose_outside)} non-blank line(s) outside every region — {shown}{more}") if regions and not fails: print("OK " + ", ".join(dict.fromkeys(regions))) PY ) fi if [ "$xml_report" = "__skipped__" ]; then : else local xml_fail_lines xml_warn_line xml_ok_line xml_fail_lines=$(printf '%s\n' "$xml_report" | grep '^FAIL ' | sed 's/^FAIL //' || true) xml_warn_line=$(printf '%s\n' "$xml_report" | grep '^WARN ' | sed 's/^WARN //' || true) xml_ok_line=$(printf '%s\n' "$xml_report" | grep '^OK ' | sed 's/^OK //' || true) if [ -n "$xml_fail_lines" ]; then print_fail "Semantic XML regions missing or unbalanced: $(echo "$xml_fail_lines" | tr '\n' ';' | sed 's/;$//; s/;/; /g')" \ "Wrap each major section in a balanced descriptive tag on its own line — <purpose>…</purpose>, <requirements>…</requirements>, <workflow>…</workflow>, <output_contract>…</output_contract> — with Markdown inside and code in fences. Every '## ' heading must sit inside a region." else print_pass "Semantic XML regions present and balanced: $xml_ok_line" fi if [ -n "$xml_warn_line" ]; then print_warn "$xml_warn_line" \ "Move stray prose into a region (the H1 title, '---' rules, and HTML comments may stay outside)." fi fi fi # ────────────────────────────────────────────────────────── # 5. Content Quality # ────────────────────────────────────────────────────────── print_section "5. Content Quality" if [ -f "$skill_path/SKILL.md" ]; then # Invocation: body phrase OR frontmatter trigger phrases if grep -qi "invoke with:\|to invoke:\|trigger with:" "$skill_path/SKILL.md"; then print_proxy_ok "Invocation phrase documented in body (phrase match only)" elif [ "$has_frontmatter" -eq 1 ] && echo "$frontmatter" | grep -q '"'; then print_proxy_ok "Invocation covered by trigger phrases in frontmatter description" else print_proxy_miss "No invocation guidance found by phrase match" \ "Add '**Invoke with:** /skill-name or ask about [topic]' near the top of SKILL.md" fi # Unsourced wall-clock and percentage claims. This check used to reward them, # passing a skill for containing "minutes" or "NN%". Wall-clock estimates depend # on who or what runs the workflow and are invented more often than measured, and # a bare percentage with no baseline is not checkable. Flag them for review instead. local timeclaim_count timeclaim_count=$(grep -cE "[0-9]+ ?(min|mins|minutes|hours|hrs)\b|[0-9]+% (faster|fewer|less|more)" "$skill_path/SKILL.md" || true) if [ "$timeclaim_count" -gt 0 ]; then print_proxy_miss "$timeclaim_count line(s) carry a wall-clock or percentage claim" \ "Each needs a baseline, a workload, and how it was measured, or it should state a completion condition instead. Pattern-matched only; verify each by hand." else print_proxy_ok "No bare wall-clock or percentage claims found (pattern match only)" fi # Second-person prose (skills teach Claude to write — use imperative form) local second_person_count second_person_count=$(grep -c "you'll\|you should\|you need to\|What you'll\|you will\b\|You'll\|You should\|You need" "$skill_path/SKILL.md" || true) if [ "$second_person_count" -gt 0 ]; then print_proxy_miss "$second_person_count line(s) use second-person prose ('you'll', 'you should')" \ "Rewrite as imperative form — 'What you'll do:' → 'To do this:'. Skills teach Claude to write, so use the form Claude should follow." else print_proxy_ok "No second-person prose matched (pattern match only)" fi # Filler intro phrases (these add words without adding information) local filler_count filler_count=$(grep -ci "to understand\|skill should\|in this step\|in this section\|this section covers\|this section explains\|as you can see\|it is important to note\|please note that\|it is worth noting" "$skill_path/SKILL.md" || true) if [ "$filler_count" -gt 0 ]; then print_proxy_miss "$filler_count line(s) contain filler phrases ('to understand', 'skill should', 'in this step', etc.)" \ "Remove filler intros — the step header already states context. 'To understand what the skill does:' → delete the line; the list below it stands alone." else print_proxy_ok "No filler intro phrases matched (pattern match only)" fi # DRY check: content duplication between body and references/ if [ -f "$skill_path/references/refining-skills.md" ] && grep -q "Refining\|refining" "$skill_path/SKILL.md"; then local refine_lines_body refine_lines_body=$(grep -c "refin" "$skill_path/SKILL.md" || true) if [ "$refine_lines_body" -gt 10 ]; then print_warn "SKILL.md has $refine_lines_body lines about refining AND references/refining-skills.md exists" \ "Remove the duplicate body content and add: 'For the refinement workflow, see references/refining-skills.md'" fi fi fi # ────────────────────────────────────────────────────────── # 6. Supporting Files # ────────────────────────────────────────────────────────── print_section "6. Supporting Files" # Scripts if [ -d "$skill_path/scripts" ]; then local script_count script_count=$(find "$skill_path/scripts" -type f \( -name "*.sh" -o -name "*.py" -o -name "*.js" \) | wc -l | tr -d ' ') if [ "$script_count" -gt 0 ]; then print_pass "scripts/ — $script_count executable(s)" else print_info "scripts/ exists but is empty (optional)" fi else print_info "No scripts/ (optional — add automation utilities that run without loading into context)" fi # References if [ -d "$skill_path/references" ]; then local ref_file_count ref_file_count=$(find "$skill_path/references" -name "*.md" | wc -l | tr -d ' ') print_pass "references/ — $ref_file_count .md file(s)" # Check that SKILL.md links to them — otherwise Claude won't load them if grep -q "references/" "$skill_path/SKILL.md" 2>/dev/null; then print_pass "SKILL.md links to references/ files" else print_warn "SKILL.md doesn't mention references/ — Claude won't know to load them" \ "Add an 'Additional Resources' section listing each references/*.md with a one-line description of what's in it" fi else print_info "No references/ (optional — add detailed docs that Claude loads as needed)" fi # Assets if [ -d "$skill_path/assets" ]; then local asset_count asset_count=$(find "$skill_path/assets" -type f | wc -l | tr -d ' ') print_pass "assets/ — $asset_count file(s)" else print_info "No assets/ (optional — add files the skill pastes into its output: images, templates, boilerplate)" fi # ────────────────────────────────────────────────────────── # 7. Common Issues # ────────────────────────────────────────────────────────── print_section "7. Common Issues" # Internal links. # # From SKILL.md this is DECIDED: every path SKILL.md names is a pointer the # agent is told to follow, so a missing file is a defect with no ambiguity. # # From a reference file it is a PROXY. Reference files carry teaching examples # that name files of the skill *being built* (`references/policy.md`), and a # changelog names paths that were deliberately removed. No script can tell # those from a real broken pointer, so they are reported and never scored. collect_broken_links() { local f=$1 dir target out="" dir=$(dirname "$f") while IFS= read -r target; do [ -n "$target" ] || continue case "$target" in /*|~*|*" "*) continue ;; esac if [ ! -e "$dir/$target" ] && [ ! -e "$skill_path/$target" ] && \ { [ -z "$repo_root" ] || [ ! -e "$repo_root/$target" ]; }; then out="${out}${f#"$skill_path"/} → $target; " fi done < <(awk '/^```/{fence=!fence; next} !fence' "$f" \ | grep -oE '`(\.\./)?(references|scripts|assets|notes|templates)/[^`]*\.(md|sh|py|pdf)`' \ | tr -d '`' | sort -u) printf '%s' "$out" } if [ -f "$skill_path/SKILL.md" ]; then local skill_broken ref_broken="" skill_broken=$(collect_broken_links "$skill_path/SKILL.md") if [ -n "$skill_broken" ]; then print_fail "SKILL.md names files that do not exist" "$skill_broken" else print_pass "Every path SKILL.md names resolves on disk" fi while IFS= read -r mdfile; do [ -f "$mdfile" ] || continue ref_broken="${ref_broken}$(collect_broken_links "$mdfile")" done < <(find "$skill_path/references" -name "*.md" 2>/dev/null) if [ -n "$ref_broken" ]; then print_proxy_miss "Unresolved paths named in reference files: $ref_broken" \ "Some of these are teaching examples naming files of the skill being built, or historical paths in a changelog. Check each by hand." else print_proxy_ok "No unresolved paths named in reference files" fi # Destructive commands in shell examples. A skill's examples get run. if awk '/^```(bash|sh)/{f=1; next} /^```/{f=0} f' "$skill_path/SKILL.md" \ $(find "$skill_path/references" -name "*.md" 2>/dev/null) 2>/dev/null \ | grep -qE '(^|[^a-zA-Z_-])rm (-[rRfi]+ )?[~/$]'; then print_fail "A shell example runs 'rm' on a real path" \ "Use 'trash' or move to a backup. 'rm' destroys the only copy and the record of when it went." else print_pass "No 'rm' on real paths in shell examples" fi # Code-fence structure. Both failures below are decidable from the text # and both silently corrupt rendering, so a reader sees instructions as # code or code as prose without any error anywhere. # # Unbalanced: an odd number of fence lines leaves the tail of the file # inside a code block. # Nested: CommonMark closes a fence at the first fence of the same # length, so ```bash written inside ```markdown ends the outer block # early. Nesting requires a longer outer fence (````). local fence_report if ! command -v python3 >/dev/null 2>&1; then print_info "Code-fence check skipped — python3 not found on PATH" fence_report="__skipped__" else fence_report=$(python3 - "$skill_path" <<'PY' 2>/dev/null || true import re, sys, pathlib root = pathlib.Path(sys.argv[1]) files = [root / "SKILL.md"] + sorted(root.glob("references/**/*.md")) for f in files: if not f.is_file(): continue rel = f.relative_to(root) # CommonMark: a fenced block ends at the first fence line that has no info # string and is at least as long as its opener. A fence line carrying an # info string is literal text inside a block, never a closer. inside, opener = False, None for i, line in enumerate(f.read_text(errors="replace").splitlines(), 1): m = re.match(r"^(`{3,})(\s*\S+)?\s*$", line) if not m: continue ticks, info = len(m.group(1)), (m.group(2) or "").strip() if not inside: inside, opener = True, (i, ticks, info) elif not info and ticks >= opener[1]: inside = False elif info and ticks == opener[1]: # Same-length nesting. The inner fence is literal, and the next bare # fence of this length ends the OUTER block early. print(f"{rel}:{i} ```{info} nested inside the same-length ```" f"{opener[2] or 'plain'} block opened at line {opener[0]}") if inside: print(f"{rel}: block opened at line {opener[0]} (```{opener[2] or 'plain'}) is never closed") PY ) fi if [ "$fence_report" = "__skipped__" ]; then : elif [ -n "$fence_report" ]; then print_fail "Code fences do not nest or balance: $(echo "$fence_report" | tr '\n' '; ')" \ "Close every fence, and widen an outer fence that contains another to four backticks (\`\`\`\`)." else print_pass "Code fences balance and none nests inside a same-length fence" fi # open() on a tilde path in python examples: Python does not expand ~. if awk '/^```python/{f=1; next} /^```/{f=0} f' "$skill_path/SKILL.md" \ $(find "$skill_path/references" -name "*.md" 2>/dev/null) 2>/dev/null \ | grep -qE "open\(['\"]~"; then print_fail "A Python example calls open() on a '~' path, which always raises FileNotFoundError" \ "Python does not expand '~'. Pass the path as an argument, or use os.path.expanduser()." else print_pass "No Python open() calls on unexpanded '~' paths" fi fi if [ -f "$skill_path/SKILL.md" ]; then # TODO markers local todo_count todo_count=$(grep -c "\[TODO\]\|TODO:" "$skill_path/SKILL.md" || true) if [ "$todo_count" -gt 0 ]; then print_warn "$todo_count TODO marker(s) in SKILL.md" \ "Complete or remove TODOs before publishing — they signal incomplete work" else print_pass "No TODO markers" fi # Unfilled template placeholders — exclude content inside code fences (teaching examples) local outside_code_fences outside_code_fences=$(awk '/^```/{in_fence=!in_fence; next} !in_fence{print}' "$skill_path/SKILL.md") if echo "$outside_code_fences" | grep -qi "your-skill-name-here\|replace this\|fill in\|\[domain\]\|\[trigger phrase\]"; then print_warn "Unfilled template placeholders detected (outside code fences)" \ "Replace all [placeholder] text with actual content before publishing" else print_pass "No unfilled placeholders (code-fence teaching examples correctly excluded)" fi # Underscore skill names — exclude table rows (^|) and ❌ examples (intentional wrong-example markers) local no_bad_examples no_bad_examples=$(grep -v "^|" "$skill_path/SKILL.md" | grep -v "❌" || true) if echo "$no_bad_examples" | grep -qE "my_skill|skill_name|test_skill"; then print_warn "Underscore-style skill names used outside teaching examples (my_skill, skill_name)" \ "Update to kebab-case: my-skill, skill-name" else print_pass "No underscore-style skill names outside ❌ teaching examples" fi fi # ────────────────────────────────────────────────────────── # 8. Activation shape # # Both checks here are PROXY. They match the shape of a description, and # nothing in this script observes whether a host actually activated the # skill. Only a triggering test on the target host does that. # # This section checks only activation-shape heuristics. Packaged claims # must cite sources a reader can retrieve. # ────────────────────────────────────────────────────────── print_section "8. Activation shape (proxy checks)" if [ -f "$skill_path/SKILL.md" ] && [ "$has_frontmatter" -eq 1 ]; then # "Use when" phrasing states the activation condition in the field the # host matches against. Both documented description formats include it. if echo "$frontmatter" | grep -qi "use this when\|use when\|should be used when"; then print_proxy_ok "Description states an activation condition ('use when' phrasing matched)" else print_proxy_miss "Description has no 'use when' / 'should be used when' phrasing" \ "State the condition explicitly, then quote the phrases: 'Generates X from Y. Use when user asks to \"A\", \"B\".' See references/best-practices.md, section 'Description Writing Formula'." fi # Concrete identifiers (*.py, camelCase, `backticked`, $vars) give a host # something literal to match. Counted by line, shape only. local trigger_word_count trigger_word_count=$(echo "$frontmatter" | grep -cE '\*\.[a-z]+|[A-Z][a-z]+[A-Z]|`[a-z_]+`|\$[a-z]|command\(\)|function\(\)' || true) if [ "$trigger_word_count" -gt 0 ]; then print_proxy_ok "Description carries concrete identifiers on $trigger_word_count line(s) (pattern match only)" else print_proxy_miss "Description carries no concrete identifiers (file patterns, function names, flags)" \ "Add the literal terms a user would type — '*.py', 'SKILL.md', 'pytest' — alongside the natural-language phrases." fi fi # ────────────────────────────────────────────────────────── # Summary # ────────────────────────────────────────────────────────── print_section "Audit Summary" echo "" echo -e " ${GREEN}Passed${NC}: $PASSED" echo -e " ${YELLOW}Warnings${NC}: $WARNINGS" echo -e " ${RED}Failed${NC}: $FAILED" echo "" # Score counts DECIDED checks only. Proxies are heuristics; folding them in # produced a number that read as a quality verdict while measuring headings. local total=$((PASSED + WARNINGS + FAILED)) local score=0 if [ "$total" -gt 0 ]; then score=$((PASSED * 100 / total)) fi # Visual score bar (20 chars wide). The colour variables hold literal '\033' # text, so they must sit in printf's FORMAT string, where printf interprets # the escape. Passing them as %s arguments prints the escape verbatim. local bar_fill=$((score / 5)) local bar_empty=$((20 - bar_fill)) local bar_color="$GREEN" if [ "$score" -lt 70 ]; then bar_color="$RED"; elif [ "$score" -lt 90 ]; then bar_color="$YELLOW"; fi printf " Decided-check score: ${bar_color}%d%%${NC} [${bar_color}" "$score" local j=0 while [ $j -lt $bar_fill ]; do printf "█"; j=$((j+1)); done printf "${NC}" local k=0 while [ $k -lt $bar_empty ]; do printf "░"; k=$((k+1)); done printf "]\n" echo -e " ${BLUE}Proxy checks${NC}: $PROXY_OK matched, $PROXY_MISS did not (heuristics, not scored)" echo "" if [ "$FAILED" -eq 0 ] && [ "$WARNINGS" -eq 0 ]; then echo -e " ${GREEN}✨ No structural problems detected.${NC}" elif [ "$FAILED" -eq 0 ] && [ "$score" -ge 90 ]; then echo -e " ${GREEN}✅ No failures detected — minor warnings below.${NC}" elif [ "$FAILED" -eq 0 ]; then echo -e " ${YELLOW}⚠️ No failures detected — review warnings below.${NC}" else echo -e " ${RED}❌ Structural problems found — fix FAILs first (skill may not activate or work correctly).${NC}" fi echo -e " ${CYAN}This is a structural smoke test. A clean run means the files are shaped${NC}" echo -e " ${CYAN}correctly, not that the guidance inside them is correct or consistent.${NC}" echo "" # ────────────────────────────────────────────────────────── # What this script cannot check # ────────────────────────────────────────────────────────── print_section "Needs a reader (not checkable by this script)" echo "" echo " Every defect below has been found by hand in a shipped skill, including in" echo " ai-skill-builder itself. None is visible to any check above. Read for them." echo "" echo " 1. Guidance that contradicts other guidance in the same package — one" echo " reference prescribing what another calls a defect." echo " 2. Guidance that contradicts this script. Where they disagree, the" echo " script is checkable and the prose is not; fix the prose." echo " 3. Numbers presented as measurements with no baseline, workload, or" echo " source. A percentage with nothing behind it is not checkable." echo " 4. Compatibility claimed for hosts nobody tested." echo " 5. Whether a reference is actually loaded under the condition its" echo " pointer names, and whether reading it changes the outcome." echo " 6. Whether the workflow is correct for the skill's actual domain." echo " 7. What a script under scripts/ writes. This run reads the skill you" echo " pointed it at, not the output of anything in it. If a script" echo " generates skill content, run it and audit the result — that is how" echo " ai-skill-builder found its own scaffolder emitting a skill this" echo " script fails." echo " 8. Whether each XML region's name describes what it holds, whether the" echo " split into regions matches the task, and whether nesting is deeper" echo " than the task needs. Section 4 proves presence, balance, and that" echo " no H2 escapes a region; it cannot read the tag name for meaning." echo "" # Actionable fix list if [ "${#FAIL_ITEMS[@]}" -gt 0 ] || [ "${#WARN_ITEMS[@]}" -gt 0 ]; then print_section "Action Items" echo "" local idx=1 for item in "${FAIL_ITEMS[@]}"; do local issue="${item%% —*}" local action="${item##* — }" echo -e " ${RED}[$idx] FAIL${NC}: $issue" echo -e " ${CYAN}→${NC} $action" echo "" idx=$((idx+1)) done for item in "${WARN_ITEMS[@]}"; do local issue="${item%% —*}" local action="${item##* — }" echo -e " ${YELLOW}[$idx] WARN${NC}: $issue" echo -e " ${CYAN}→${NC} $action" echo "" idx=$((idx+1)) done fi echo " Re-run: bash $0 $skill_path" echo "" } case "${1:-}" in -h|--help) echo "Usage: bash audit-skill.sh <skill-path>" echo "" echo "Audits one skill directory against the portable Agent Skills structure." echo "" echo "Exit codes:" echo " 0 zero structural FAILs (warnings and proxy misses do not change it)" echo " 1 no path given, path is not a directory, or one or more FAILs" echo "" echo "A zero exit means the files are shaped correctly. It does not mean the" echo "guidance inside them is correct — see 'Needs a reader' in a full run." exit 0 ;; esac # `set -e` is on, so a nonzero return from audit_skill would exit before this. # The FAIL count is the release gate: a structural failure that exits 0 reads to # CI and to a human as a clean run. audit_skill "$@" if [ "$FAILED" -gt 0 ]; then exit 1 fi -
scaffold-skill.sh 12.1 KB
#!/bin/bash ############################################################################## # AI Skill Scaffolder # # Creates a new Agent Skill with proper structure and templates # # Usage: # bash scaffold-skill.sh <skill-name> [category] # # Arguments: # skill-name - Name in kebab-case (e.g., api-test-generator) # category - Optional: document, workflow, or mcp (default: document) # # Environment: # SKILLS_DIR - Discovery root to scaffold into. Default ~/.claude/skills. # Set it for another host: SKILLS_DIR=~/.agents/skills bash scaffold-skill.sh x # # Example: # bash scaffold-skill.sh my-awesome-skill document # # The generated SKILL.md is written to pass scripts/audit-skill.sh with zero # FAILs before any TODO is filled in. If a change here breaks that, the two # scripts have started disagreeing; fix this one. # # Portability: POSIX-compatible parameter expansion only. macOS ships bash # 3.2.57 as /bin/bash, where bash-4 forms such as ${var^} are a syntax error. ############################################################################## set -e # Colors for output RED='\033[0;31m' GREEN='\033[0;32m' YELLOW='\033[1;33m' BLUE='\033[0;34m' NC='\033[0m' # No Color # Print functions print_info() { echo -e "${BLUE}ℹ️ $1${NC}" } print_success() { echo -e "${GREEN}✅ $1${NC}" } print_warning() { echo -e "${YELLOW}⚠️ $1${NC}" } print_error() { echo -e "${RED}❌ $1${NC}" } print_header() { echo "" echo "========================================" echo "$1" echo "========================================" echo "" } # Validate skill name (kebab-case) validate_skill_name() { local name=$1 # Check if empty if [ -z "$name" ]; then print_error "Skill name is required" echo "Usage: bash scaffold-skill.sh <skill-name> [category]" exit 1 fi # Check for invalid characters if [[ ! "$name" =~ ^[a-z0-9]+(-[a-z0-9]+)*$ ]]; then print_error "Invalid skill name: $name" echo "" echo "Skill name must:" echo " - Use lowercase letters and numbers only" echo " - Use hyphens (-) to separate words" echo " - Not start or end with hyphen" echo "" echo "Valid examples:" echo " ✅ api-test-generator" echo " ✅ deploy-to-production" echo " ✅ smart-search" echo "" echo "Invalid examples:" echo " ❌ API-Test-Generator (uppercase)" echo " ❌ api_test_generator (underscores)" echo " ❌ apiTestGenerator (camelCase)" exit 1 fi } # Get category templates get_category_info() { local category=$1 case "$category" in document|doc) CATEGORY_NAME="Document & Asset Creation" CATEGORY_DESC="Transforms inputs into structured outputs (documents, code, diagrams, reports)" CATEGORY_EXAMPLE="Generate API documentation from code comments" ;; workflow|work) CATEGORY_NAME="Workflow Automation" CATEGORY_DESC="Automates multi-step processes requiring coordination" CATEGORY_EXAMPLE="Deploy application with validation and rollback" ;; mcp) CATEGORY_NAME="MCP Enhancement" CATEGORY_DESC="Extends or combines MCP server capabilities" CATEGORY_EXAMPLE="Combine database and API tools for data sync" ;; *) print_warning "Unknown category: $category, using 'document'" CATEGORY_NAME="Document & Asset Creation" CATEGORY_DESC="Transforms inputs into structured outputs" CATEGORY_EXAMPLE="Generate structured output from input" ;; esac } # Main function main() { print_header "AI Skill Scaffolder" # Parse arguments SKILL_NAME=$1 CATEGORY=${2:-document} # Validate skill name validate_skill_name "$SKILL_NAME" # Get category info get_category_info "$CATEGORY" # Set paths. SKILLS_DIR is overridable because discovery roots differ by # host — see the target-host questions in SKILL.md Step 1. SKILLS_DIR="${SKILLS_DIR:-$HOME/.claude/skills}" SKILL_DIR="$SKILLS_DIR/$SKILL_NAME" # Title for the H1. Written with tr and awk rather than ${SKILL_NAME^}: # that form needs bash 4, and macOS /bin/bash is 3.2.57. SKILL_TITLE=$(echo "$SKILL_NAME" | tr '-' ' ' \ | awk '{for (i = 1; i <= NF; i++) $i = toupper(substr($i, 1, 1)) substr($i, 2)} 1') print_info "Skill name: $SKILL_NAME" print_info "Category: $CATEGORY_NAME" print_info "Target directory: $SKILL_DIR" echo "" # Check if directory exists if [ -d "$SKILL_DIR" ]; then print_warning "Skill directory already exists: $SKILL_DIR" read -p "Move the existing directory aside and continue? (y/N) " -n 1 -r echo if [[ ! $REPLY =~ ^[Yy]$ ]]; then print_info "Cancelled" exit 0 fi # Never rm -rf a directory the user may have worked in. This package # tells skill authors to use trash instead of rm, and audit-skill.sh # fails a skill whose examples do otherwise. if command -v trash >/dev/null 2>&1; then trash "$SKILL_DIR" print_success "Existing directory moved to trash" else SKILL_BACKUP="${SKILL_DIR}.bak.$(date +%Y-%m-%d-%H%M%S)" mv "$SKILL_DIR" "$SKILL_BACKUP" print_success "Existing directory moved to $SKILL_BACKUP" fi fi # Create directory structure print_info "Creating directory structure..." mkdir -p "$SKILL_DIR" mkdir -p "$SKILL_DIR/scripts" mkdir -p "$SKILL_DIR/references" mkdir -p "$SKILL_DIR/assets" print_success "Directories created" # Create SKILL.md print_info "Creating SKILL.md..." cat > "$SKILL_DIR/SKILL.md" << EOF --- name: $SKILL_NAME description: TODO one sentence on what this produces. Use when user asks to "$SKILL_NAME", "TODO second natural phrase", or "TODO third natural phrase". metadata: version: 0.1.0 --- # $SKILL_TITLE <purpose> [TODO: Write 50-100 word hook explaining what this skill does and who it's for] This skill [describe what it produces in one sentence]. **Use this skill when:** [Specific scenario when this is useful] **Invoke with:** \`/$SKILL_NAME\` or "[Natural language trigger phrase]" **Category**: $CATEGORY_NAME </purpose> <workflow> ## How It Works [TODO: Write 200-400 word workflow overview with 3-5 steps] This skill follows these steps: ### Step 1: [Phase Name] - [What happens in this step] - [Inputs needed] - [Outputs produced] - Definition of Done: [what must be true before Step 2 starts] ### Step 2: [Phase Name] - [What happens in this step] - [Inputs needed] - [Outputs produced] - Definition of Done: [what must be true before Step 3 starts] ### Step 3: [Phase Name] - [What happens in this step] - [Inputs needed] - [Outputs produced] - Definition of Done: [what must be true before this skill reports success] **Definition of Done for the whole workflow**: [the checkable end state] --- ## Detailed Workflow [TODO: Write comprehensive documentation with no word limit] ### Prerequisites - [Required tool/dependency 1] - [Required tool/dependency 2] - [Required knowledge/skill] ### Step-by-Step Guide #### Step 1: [Detailed Phase Name] **Purpose**: [Why this step matters] **Process**: 1. [Detailed sub-step 1] 2. [Detailed sub-step 2] 3. [Detailed sub-step 3] **Inputs**: - Input 1: [Description, format, example] **Outputs**: - Output 1: [Description, format, example] **Common Issues**: - Issue: [Problem and solution] [TODO: Continue for all steps] </workflow> <examples> ## Examples ### Example 1: [Common Use Case] **Scenario**: [Describe the situation] **Input**: \`\`\` [Actual input example] \`\`\` **Output**: \`\`\` [Actual output example] \`\`\` **Result**: [Outcome achieved] </examples> <success_criteria> ## Success Metrics Measure the manual baseline before building, then state each number with its workload, its unit, and which direction is better. A figure with none of those is not checkable. ### Quantitative - [What was measured]: [manual baseline] → [with this skill], over [workload] - [What was measured]: [manual baseline] → [with this skill], over [workload] ### Qualitative - Consistency: [How it standardizes the process] - Best Practices: [What standards it follows] </success_criteria> <authoring_notes> ## Next Steps Delete this region once the skill is written; it instructs the author, not the agent. 1. **Fill in TODOs**: Replace all [TODO] sections with actual content 2. **Add Examples**: Include real examples from your use case 3. **Test Triggering**: Verify the target agent detects the skill 4. **Validate Function**: Test the complete workflow 5. **Measure Performance**: Compare to baseline metrics 6. **Get Feedback**: Test with target users 7. **Distribute**: Share on GitHub and community Record the version in \`metadata.version\` above and the release history in a changelog file. For guidance, read these inside the ai-skill-builder skill directory. The paths are relative to that skill, not to this one, so they will not resolve from here: - Template: ai-skill-builder/references/examples/SKILL-template.md - Best practices: ai-skill-builder/references/best-practices.md Keep every major section inside a balanced, descriptive XML region as above (\`<purpose>\`, \`<workflow>\`, \`<examples>\`, \`<success_criteria>\`); the audit fails a \`## \` heading that sits outside every region. Rename regions to fit the task; do not add depth the task does not need. Validate this file: \`bash ai-skill-builder/scripts/audit-skill.sh $SKILL_DIR\` </authoring_notes> EOF print_success "SKILL.md created" # No placeholder files are written. Earlier versions dropped a README.md in # scripts/, references/, and assets/. Two problems: this package's own rule # is that a skill folder carries no README.md, and every .md under # references/ is loadable context, so a file whose content is "add # documentation here" costs tokens to say nothing. The guidance is printed # to the terminal instead, where the author reads it once. echo "" print_info "What goes in each directory:" echo " scripts/ — executables the agent runs without loading into context" echo " (validators, generators). chmod +x them." echo " references/ — markdown the agent loads on demand (schemas, API docs," echo " policies). Every file here costs context when read, so" echo " name the condition for reading it in SKILL.md." echo " assets/ — files the skill pastes into its output (templates," echo " images, boilerplate). Never loaded as instructions." echo " Reference each one from SKILL.md, or the agent will not know it exists." # Create .gitignore cat > "$SKILL_DIR/.gitignore" << EOF # Generated outputs outputs/ *.log # Temporary files *.tmp .DS_Store # Virtual environments venv/ .venv/ env/ # IDE files .vscode/ .idea/ *.swp EOF # Summary print_header "Skill Scaffolding Complete!" print_success "Skill created at: $SKILL_DIR" echo "" print_info "Directory structure:" tree -L 2 "$SKILL_DIR" 2>/dev/null || ls -R "$SKILL_DIR" echo "" print_info "Next steps:" echo " 1. Edit SKILL.md and replace every [TODO] section" echo " 2. Rewrite the description: what it produces, then the phrases users say" echo " 3. Add scripts, references, and assets as needed, and link them from SKILL.md" echo " 4. Test triggering: /$SKILL_NAME, then a natural phrase, then one that must NOT match" echo " 5. Re-run the audit until it reports zero FAILs" echo "" print_info "Resources, relative to the ai-skill-builder skill directory:" echo " - Guide: SKILL.md" echo " - Template: references/examples/SKILL-template.md" echo " - Best practices: references/best-practices.md" echo " - Validate: bash scripts/audit-skill.sh '$SKILL_DIR'" } # Error handling trap 'print_error "Scaffolding failed at line $LINENO"; exit 1' ERR # Run main function main "$@"
-
-
SKILL.md 20.4 KB
--- name: ai-skill-builder description: Guides creation, audit, and improvement of portable Agent Skills using the shared SKILL.md format and Anthropic's established methodology, without confusing the portable specification with runtime-specific extensions. Use when the user wants to "create a skill", "create an agent skill", "build a skill", "make a new skill", "write a new skill", "improve an existing skill", "audit my skill", "audit a SKILL.md", "refine my skill", "test a skill", "test skill discovery", "install skills across harnesses", "package a skill for distribution", or needs guidance on skill structure, semantic XML regions, progressive disclosure, description quality, frontmatter, scripts, arguments, security, testing, cross-harness compatibility, or distribution. allowed-tools: Read Write Edit Bash WebSearch WebFetch metadata: version: 1.3.0 --- # AI Skill Builder <purpose> Build Agent Skills using the shared `SKILL.md` format and Anthropic's methodology: the smallest portable core that works in every named target, plus explicit runtime adapters. Portability is a tested compatibility claim, not a formatting style. `references/sources.md` cites host documentation for Claude Code, Codex, and Qwen Code; every other host is untested here, so verify the one you target (Step 1). This methodology deliberately requires consistent semantic XML regions inside the Markdown body. That is a quality policy for reducing ambiguity in complex operational instructions, supported by current Anthropic prompting guidance and checked by the forward tests in Step 4. It is not a parser requirement of the portable Agent Skills specification, an authorization boundary, or a claim that XML alone prevents prompt injection. **Primary methodology source**: "The Complete Guide to Building Skills for Claude" (Anthropic, January 2026) - PDF: https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf - Extracted text: `references/ai-skill-builder-guide.md` **Portable format source**: https://agentskills.io/specification --- ## Quick Start To create a skill from scratch, follow the 4-phase process below. To improve an existing skill, read `references/refining-skills.md`. To validate structure, run `scripts/audit-skill.sh`. **Invoke with:** ask for `ai-skill-builder`; hosts may expose it as `/ai-skill-builder`, `$ai-skill-builder`, or through their skills picker. </purpose> <requirements> ## P0 requirements Every skill this methodology produces meets all thirteen. Cite them by ID in audits and reviews. 1. **SKILL-REQ001-discover-target-runtimes.** Identify exact products, versions, invocation modes, install roots, trust boundaries, and filesystem semantics before authoring. 2. **SKILL-REQ002-separate-standard-from-extensions.** Keep the Agent Skills specification, runtime extensions, and project conventions in separate columns. Never present a Claude-, Codex-, Copilot-, Gemini-, or framework-specific field as universal. 3. **SKILL-REQ003-use-valid-portable-core.** Put a required `SKILL.md` in a directory whose name matches its portable `name`; include a specific `description` that says what the skill does and when to use it. Use standard optional fields only when needed. 4. **SKILL-REQ004-use-semantic-xml-regions.** Wrap every major operational region in consistent, descriptive XML tags inside the Markdown body, such as `<purpose>`, `<requirements>`, `<workflow>`, `<examples>`, and `<output_contract>`. Keep Markdown inside those regions and literal source in fenced code blocks. This mandatory methodology rule aims for the highest-quality instruction separation; it is not an Agent Skills parser requirement or a security boundary. `scripts/audit-skill.sh` fails a body with no region, an unbalanced or mis-nested tag, or a `## ` heading outside every region. 5. **SKILL-REQ005-design-progressive-disclosure.** Keep the activation description precise and the main workflow concise. Move genuinely conditional detail into focused, directly linked references; keep reusable automation in scripts and output materials in assets. 6. **SKILL-REQ006-preserve-operational-context.** Every reference or script must say when to use it, what inputs it accepts, what it returns or changes, and how failure is reported. Avoid deep reference chains. 7. **SKILL-REQ007-make-invocation-correct.** Test automatic activation and explicit invocation. Where a runtime supports parameters, test empty, positional, named, quoted, Unicode, invalid, and large inputs. Do not assume another runtime implements the same substitution syntax. 8. **SKILL-REQ008-minimize-tool-authority.** Treat skills and bundled scripts as executable dependencies. Request the least authority needed, validate inputs, preserve approval prompts where practical, and never describe prompt delimiters as prompt-injection prevention. 9. **SKILL-REQ009-validate-every-harness.** Validate discovery, load, references, scripts, arguments, symlinks, duplicate resolution, updates, and removal in each supported harness. Passing one parser does not prove cross-runtime compatibility. 10. **SKILL-REQ010-preserve-user-state.** Installation and uninstall must own exact generated paths, preserve manual edits, support multiple configured skill roots, and remove only installer-owned links or files. 11. **SKILL-REQ011-measure-context-and-runtime-cost.** Record discovery metadata size, loaded instruction size, conditional reference cost, script time/memory/I/O, and duplicate loading. Treat line and token targets as heuristics unless a target runtime makes them normative. 12. **SKILL-REQ012-forward-test-real-tasks.** Use realistic positive, negative, ambiguous, and adversarial prompts. Check task outcome, not merely whether the skill appeared in a list. 13. **SKILL-REQ013-separate-authoring-from-execution.** Classify whether the task is to change the skill, use the skill to produce an artifact or result, or evaluate the skill from evidence. Runtime inputs, generated outputs, examples, and dogfood observations do not become skill instructions unless an explicit authoring decision accepts and generalizes them. Region names follow the task rather than one universal taxonomy: workflows use `<workflow>`, reference skills may use `<decision_guide>`, multi-mode skills `<routing>`, strict result contracts `<output_contract>`, inline reference material `<reference>`, and pointers to bundled files and sources `<resources>`. Tags must be consistent, descriptive, balanced, and no more deeply nested than the task requires. The standard-versus-extension matrix, the claim audit, and the per-host validation receipt live in `references/portability-and-claim-audit.md`. </requirements> <workflow> ## How It Works Four-phase methodology from the Anthropic guide. Each phase carries a **Definition of Done**: the condition that must hold before the next phase starts. | Phase | Activity | Definition of Done | |-------|----------|--------------------| | 1: Planning & Design | Define target hosts, use cases, category, success criteria | Criteria are measurable and the category is chosen | | 2: Implementation | Create folder, write SKILL.md, add resources | `audit-skill.sh` reports no failures | | 3: Testing | Triggering, functional, performance, compatibility, forward tests | Every named host passes with a validation receipt; baseline recorded | | 4: Distribution | Package, document, publish | Install verified from a clean root | Phase 1 covers Steps 1–2 below; Phases 2–4 correspond to Steps 3–5. How long a phase takes depends on who or what runs it, so this skill states no wall-clock estimates and the skills you build should not either. --- ## Skill Creation Workflow ### Step 1: Discover Requirements **Research the domain first** — use the host's web tools to verify current best practices before writing instructions. Outdated guidance is worse than none: an agent follows it confidently. **Retain all sources**: record every URL consulted with a note of what it confirmed. Unsourced guidance cannot be verified or updated. For research strategies, source quality standards, and the required Sources section format: `references/research.md` 1. Classify the request: authoring (change the skill), execution (use the skill to produce a result), or evaluation (judge the skill from evidence). Name the artifact and the mutation boundary before using task content as design input (SKILL-REQ013). 2. Identify the problem and target users 3. Identify every target host and version, its discovery root, invocation form, and reload behavior. Mark unknowns as unknown rather than guessing. Inspect existing skill roots and project guidance before creating a parallel package. 4. Define 2-3 concrete use cases 5. Set measurable success criteria (time saved, errors reduced, quality improved) 6. Choose a skill category: | Category | INPUT → OUTPUT | Examples | |----------|---------------|---------| | **1: Document & Asset Creation** | Data/specs → document, code, report | API test generator, meeting notes summarizer | | **2: Workflow Automation** | Task params → completed multi-step process | Deploy pipeline, code review workflow | | **3: MCP Enhancement** | MCP tool outputs → smarter orchestration | Smart file search, BigQuery assistant | For in-depth category guidance: `references/categories.md` For interactive discovery questions: `references/discovery.md` ### Step 2: Design Structure Design using progressive disclosure — 3 loading levels: | Level | When loaded | Target length | Content | |-------|------------|---------------|---------| | 1: Metadata | Always (~100 words) | name + description | Trigger conditions | | 2: SKILL.md body | When skill triggers | under 5,000 words (ideally under 2,000) | Core workflow + pointers | | 3: references/ files | As needed | Unlimited | Deep detail, schemas, examples | For progressive disclosure writing tips and success metrics: `references/best-practices.md` ### Step 3: Implement **Critical Rules:** - ✅ File MUST be named `SKILL.md` (not README.md) - ✅ Folder name MUST be **kebab-case** — no spaces, no capitals (`my-skill` not `My Skill`) - ✅ YAML frontmatter MUST include `name` and `description`; put a version under `metadata` - ✅ Description MUST include specific trigger phrases — what users SAY to activate the skill - ✅ Description must be **under 1024 characters** (hard limit — longer descriptions are truncated) - ✅ No XML angle brackets (`<` or `>`) in any frontmatter field — a host restriction from Anthropic's skill guide, not a YAML rule; the body is unaffected - ✅ Body: every major section sits inside a balanced, descriptive XML region on its own line (`<purpose>`, `<requirements>`, `<workflow>`, `<examples>`, `<output_contract>`), Markdown inside, literal source in fenced code blocks (SKILL-REQ004). The H1 title may stay outside. - ❌ NO README.md inside the skill folder — all docs go in SKILL.md or references/ (Exception: a README.md at the GitHub repo ROOT, outside the skill folder, is fine for GitHub.) - ❌ NO spaces or underscores in folder names (`my_skill` → `my-skill`) **Frontmatter: required fields**: ```yaml --- name: your-skill-name # kebab-case only; no spaces, capitals, or underscores description: What it does. Use when user asks to "specific phrase", "another phrase". # MUST include: what it does + when to use it (trigger conditions) # Under 1024 characters. No XML angle brackets. # Do NOT start with "claude" or "anthropic" (reserved namespaces). --- ``` For every optional field (`license`, `compatibility`, `allowed-tools`, `metadata`) with host-portability notes: `references/best-practices.md` **Positioning language** ("generate tests 87% faster") belongs in the repo README, never in `description`. See `references/best-practices.md`. **Folder structure** (annotated Claude Code standalone example: `references/best-practices.md`, section "Folder Structure Patterns"): | Directory | Load into context? | Use for | |-----------|-------------------|---------| | `references/` | Yes, as needed | Schemas, API docs, policies, workflow guides | | `references/examples/` | As needed | Working code users copy and adapt | | `scripts/` | No (run directly) | Validators, scaffolders, utilities | | `assets/` | No (used in output) | Images, fonts, HTML templates the skill pastes into output | Note: Plugin skills (in `plugin-name/skills/`) may place `examples/` at the top level — that is plugin-dev convention. For standalone `~/.claude/skills/` skills, put examples inside `references/`. To scaffold a new skill: `bash scripts/scaffold-skill.sh my-skill-name` (set `SKILLS_DIR` to target another host's root) To start from template: copy `references/examples/SKILL-template.md` ### Step 4: Test Five testing approaches (run in order): 1. **Triggering tests** — verify the agent activates on expected phrases and not on unrelated requests 2. **Functional tests** — validate the skill's core workflow produces correct output for known inputs 3. **Performance tests** — measure improvement over baseline (time saved, error reduction, consistency) 4. **Compatibility tests** — verify discovery, explicit invocation, reference loading, script execution, and duplicate/name resolution on every named host and version. Start a clean session for each host, run the host's own validator where one exists (Codex ships `skill-creator/scripts/quick_validate.py`), test installer reruns, multiple roots, symlink targets, modified user files, and uninstall, and record a validation receipt per host (`references/portability-and-claim-audit.md`). 5. **Forward tests** — realistic positive, negative, ambiguous, and adversarial prompts; judge the task outcome, not whether the skill appeared in a list, and check that it does not over-trigger nearby tasks (SKILL-REQ012) **Debugging trigger issues**: ``` Ask the agent: "When would you use the [skill name] skill?" ``` The agent should quote its description back. Adjust based on what's missing or too vague. **Fixing undertriggering**: Add more specific trigger phrases and relevant technical terms. **Fixing overtriggering**: Add negative triggers to the description: ```yaml description: Processes PDF legal documents for contract review. Use for "review this contract", "analyze legal document", "extract contract clauses". Do NOT use for general PDF viewing, image extraction, or non-legal documents (use doc-converter skill instead). ``` For the full diagnosis-and-fix guide covering all four failure modes: `references/troubleshooting.md` To validate structure: ```bash bash scripts/audit-skill.sh /path/to/YOUR-SKILL ``` For automated quality review on a host that provides Anthropic's `skill-creator` skill: ``` "Use the skill-creator skill to review the skill I just built and suggest improvements" ``` Claude Code with the `plugin-dev` plugin installed also offers a `skill-reviewer` agent that checks description quality. For detailed test case templates and the Testing Triangle methodology: `references/testing.md` ### Step 5: Distribute Three channels — individual install into a host's skills root, organization-wide deployment, and the Messages API — are described with commands in `references/distribution.md`, together with the GitHub repo layout, the README template, versioning, and community channels. An installer that places the skill must own exact generated paths, preserve manual edits, support multiple configured skill roots, and remove only its own links or files (SKILL-REQ010). Report the portable core and every runtime extension, the compatibility matrix and its evidence status, install/update/uninstall ownership, validation and forward-test results, and context/runtime costs. Do not claim "universal," "secure," or "supported" without a named contract and direct evidence. </workflow> <pitfalls> ## Common Pitfalls | Pitfall | ❌ Wrong | ✅ Correct | |---------|---------|---------| | File naming | `my_skill/README.md` | `my-skill/SKILL.md` | | Description field | Outcome-focused: `"generates tests 87% faster"` | Trigger phrases: `"create a skill", "improve my skill"` | | README positioning | Trigger phrases in GitHub README | Outcome-focused: `"generate tests 87% faster"` | | No progressive disclosure | Monolithic wall of text | 3-level: hook (50-100w) → workflow (200-400w) → detail | | No testing | Write → publish immediately | Triggering + functional + performance + compatibility + forward tests | | Missing success criteria | "Build a skill that helps with APIs" | "Cut API test writing from 3 h to 45 min, median of 3 runs" | | Feature-focused description | "Uses OpenAPI parser and Jinja2 templates" | Trigger phrases + concise capability summary | | Unmeasured claim | "Works with any agent host", a bare "40% faster" | The hosts and versions you tested; the workload, baseline, and unit behind the number. The 87% in the rows above is correct placement only if it was measured | --- ## Refining Existing Skills Signs a skill needs refinement: - File named README.md or folder has underscores/capitals (P0 — Claude cannot find it) - Missing YAML frontmatter (P0 — Claude cannot auto-activate it) - Description is outcome-focused instead of trigger-phrase format (P1) - SKILL.md is over 5,000 words (P0 on Claude hosts, where Anthropic's guide reports degraded quality above it; a heuristic elsewhere, per SKILL-REQ011) - SKILL.md is over 2,000 words with no references/ files (P1 — detail belongs in references/) - Wall of text with no progressive disclosure structure (P1) - No semantic XML regions, an unbalanced tag, or a `## ` heading outside every region (P1 — `scripts/audit-skill.sh` fails it; see SKILL-REQ004) For the 5-step refinement process (audit → prioritize → fix → validate → document), migration scenarios, and before/after examples: `references/refining-skills.md` </pitfalls> <resources> ## Additional Resources ### Reference Files (loaded as needed by the agent — when to open it → what it gives back) - **`references/research.md`** — before writing any domain instruction → source-tier judgement, `sources.md` entries - **`references/discovery.md`** — when requirements are unclear → 22-question plan: name, triggers, use cases, inputs, outputs, tests - **`references/categories.md`** — when Step 1's category is not obvious → category, structure, Level 2 template - **`references/best-practices.md`** — while writing body and frontmatter → description format, every optional field, level targets, 6 anti-patterns - **`references/patterns.md`** — when 4 phases are not enough → orchestration, multi-MCP, refinement loops, runtime branching, domain rules - **`references/testing.md`** — at Step 4 → T1-T6 triggering cases, F1-F4 functional cases, baseline comparison, per-host compatibility steps, forward-test kinds - **`references/troubleshooting.md`** — when a built skill misbehaves → fixes for no trigger, over-trigger, skipped instructions, context overload - **`references/refining-skills.md`** — when improving an existing skill → P0-P3 audit, migration checklist - **`references/distribution.md`** — at Step 5 → distribution channels, repo layout, README, versioning - **`references/portability-and-claim-audit.md`** — when a skill targets more than one host, or a claim about frontmatter, XML, roots, symlinks, or security needs checking → standard-vs-extension matrix, claim audit table, per-host validation receipt - **`references/sources.md`** — when checking what a claim rests on, or adding one → citations, and claims with no retrievable source - **`references/changelog.md`** — when changing ai-skill-builder itself → its release history, not the history of the skill you are building ### Superseded package - **`engineer-agent-skills`** — an earlier standalone skill. Its P0 requirements (SKILL-REQ001–013), portable-standard vs runtime-extension claim matrix, per-host validation receipt, and install/uninstall ownership rules now live in this skill; do not install both. ### Scripts (run directly — do not load into context) - **`scripts/audit-skill.sh`** — structure smoke test. Gate on zero FAILs; the percentage scores only the mechanically decidable checks, and proxy checks are listed unscored. ```bash bash scripts/audit-skill.sh /path/to/YOUR-SKILL ``` - **`scripts/scaffold-skill.sh`** — create a skill directory that already passes the audit ```bash SKILLS_DIR=~/.claude/skills bash scripts/scaffold-skill.sh my-skill-name ``` </resources>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.