ai-patterns-tool-use-patterns
Provider-agnostic patterns for LLM function calling, tool loops, and agentic workflows
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-patterns-tool-use-patterns/skills/ai-patterns-tool-use-patterns
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Tool Use Patterns
Quick Guide: Tool use (function calling) lets LLMs invoke external functions. The universal pattern is: define tool schemas (JSON Schema for parameters) -> send tools + message to LLM -> detect tool_use in response -> execute locally -> return result to LLM -> repeat until the model responds with text. Guard every loop with a max-step limit, validate all tool inputs before execution, and return structured errors so the model can recover. Use tool choice control (
auto,required,none, specific tool) to steer model behavior.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST guard every tool loop with a maximum step limit -- unbounded loops risk infinite API calls and runaway costs)
(You MUST validate all tool input arguments before execution -- LLM-generated arguments are untrusted input)
(You MUST return structured error messages to the model when tool execution fails -- never silently swallow errors or return empty results)
(You MUST use JSON Schema for tool parameter definitions -- all major providers require this format)
(You MUST treat tool definitions as token cost -- every tool schema is sent on every API call, so keep descriptions concise but precise)
</critical_requirements>
Auto-detection: tool use, function calling, tool_calls, tool_use, tool call loop, agent loop, tool definition, tool schema, toolChoice, tool_choice, parallel tool calls, human-in-the-loop, tool approval, agentic workflow, multi-step agent, tool result, tool error
When to use:
- Implementing LLM tool calling / function calling in any provider
- Building agent loops that call tools iteratively until a task is complete
- Handling parallel tool calls (multiple tools in one response)
- Reporting tool errors back to the model for recovery
- Controlling tool selection (auto, required, none, force specific)
- Adding human approval gates before dangerous tool execution
- Streaming responses that include tool calls
Key patterns covered:
- Tool definition schemas (JSON Schema for parameters, descriptions)
- The core tool call loop (send -> detect -> execute -> return -> re-send)
- Parallel tool calls (handling multiple calls in one response)
- Error handling (reporting tool failures back to the model)
- Tool choice control (auto, required, none, specific tool)
- Multi-step agent workflows with conversation state
- Human-in-the-loop approval patterns
- Type-safe tool definitions in TypeScript
- Security (input validation, sandboxing, least privilege)
- Streaming with tool calls
When NOT to use:
- Simple text generation without tool calling -- no tools needed
- Structured output / JSON extraction -- use your provider's structured output feature instead
- Provider-specific SDK patterns -- use your provider's SDK skill for SDK-specific APIs
Detailed Resources:
- examples/core.md -- Tool definitions, the tool call loop, error handling, type-safe tools
- examples/advanced.md -- Parallel tool calls, multi-step agents, human-in-the-loop, streaming, security
- reference.md -- Decision frameworks, provider comparison, anti-pattern checklist
<decision_framework>
Decision Framework
Do You Need Tool Calling?
Does the task require information the model doesn't have?
+-- YES -> Tool calling (fetch data from APIs, databases, files)
+-- NO -> Does the task require side effects?
+-- YES -> Tool calling (send email, create record, execute code)
+-- NO -> Do you need structured JSON output?
+-- YES -> Use structured output features (NOT tool calling)
+-- NO -> Plain text generation, no tools needed
Which Loop Pattern?
How many tools might the model call?
+-- Single tool call per request
| +-- Simple request-response with one tool execution
| +-- No loop needed, just one round-trip
+-- Multiple sequential tool calls
| +-- Use the bounded tool call loop (Pattern 2)
| +-- Set MAX_TOOL_STEPS based on task complexity
+-- Multiple parallel tool calls in one response
| +-- Execute all tool calls concurrently (Promise.all)
| +-- Then return all results and loop
+-- Complex multi-step agent
+-- Use the bounded loop with conversation state
+-- Add human-in-the-loop for dangerous operations
+-- Consider per-step tool filtering
How to Handle Tool Errors?
Tool execution failed. What to do?
+-- Return structured error to the model
| +-- Include error message and context
| +-- Model can retry, choose alternative, or explain failure
+-- NEVER: silently return empty result
+-- NEVER: crash the loop
+-- NEVER: retry automatically without telling the model
(the model should decide whether to retry)
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Unbounded tool loop (
while (true)) without a step counter -- risks infinite API calls and runaway costs - Executing tool arguments without validation -- LLM-generated arguments are untrusted input, treat like user input
- Swallowing tool errors silently (returning
nullor{}) -- the model cannot recover from failures it doesn't know about - Allowing arbitrary code execution from tool arguments without sandboxing -- prompt injection can escalate to code execution
- Tool descriptions that say "Gets data" -- vague descriptions cause wrong tool selection and malformed arguments
Medium Priority Issues:
- Sending all tools on every API call when only a subset is relevant -- wastes tokens and confuses the model
- Not including the assistant message (with tool calls) in conversation history before tool results -- breaks the message sequence
- Using
tool_choice: "required"without a fallback for when no tool makes sense -- forces meaningless tool calls - Returning raw database rows or full API responses as tool results -- overwhelms context with irrelevant data; summarize or truncate
- Not logging tool calls and results -- impossible to debug agent behavior in production
Common Mistakes:
- Forgetting that tool call arguments arrive as a JSON string, not a parsed object -- always
JSON.parse()before use - Assuming the model will always call tools when tools are available -- with
automode, it may respond with text directly - Treating tool calling as structured output -- they solve different problems (actions vs data extraction)
- Putting business logic in tool descriptions instead of tool implementations -- descriptions guide selection, not execution
Gotchas & Edge Cases:
- Tool definitions consume tokens on every API call -- 10 tools with detailed schemas can use 1000+ tokens per request
- Parallel tool calls may arrive in any order -- never assume execution order matches definition order
- Some models hallucinate tool names or arguments that don't match any definition -- always validate the tool name exists in your registry
- Streaming responses with tool calls require accumulating partial JSON chunks before parsing -- the arguments arrive incrementally, not all at once
- Returning very large tool results (>4000 tokens) can push the conversation past context limits -- truncate or summarize large results
- The model maintains full conversation context including all tool calls and results -- long agent runs accumulate significant token usage
- Different providers use different message formats for tool results (
role: "tool"vs content blocks) -- abstract this in your callLLM wrapper
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST guard every tool loop with a maximum step limit -- unbounded loops risk infinite API calls and runaway costs)
(You MUST validate all tool input arguments before execution -- LLM-generated arguments are untrusted input)
(You MUST return structured error messages to the model when tool execution fails -- never silently swallow errors or return empty results)
(You MUST use JSON Schema for tool parameter definitions -- all major providers require this format)
(You MUST treat tool definitions as token cost -- every tool schema is sent on every API call, so keep descriptions concise but precise)
Failure to follow these rules will produce agents that run up API costs in infinite loops, execute unvalidated input, or silently fail without the model being able to recover.
</critical_reminders>
Files (skills)
-
examples
-
advanced.md 16.1 KB
# Tool Use Patterns -- Advanced Examples > Advanced patterns for parallel tool calls, multi-step agents, human-in-the-loop approval, streaming with tool calls, and security. See [core.md](core.md) for fundamentals. **Prerequisites**: Understand the tool call loop, tool execution with error handling, and the tool registry from core examples. --- ## Pattern 6: Parallel Tool Calls When a model requests multiple tool calls in a single response, execute them concurrently for better performance. Four 300ms calls complete in ~300ms total instead of ~1200ms sequentially. ### Good Example -- Concurrent Execution with Individual Error Handling ```typescript // agent/parallel.ts async function executeToolCallsInParallel( toolCalls: ToolCall[], ): Promise<ToolResultMessage[]> { const results = await Promise.allSettled( toolCalls.map(async (toolCall) => { const result = await executeTool(toolCall); return { role: "tool" as const, toolCallId: toolCall.id, content: JSON.stringify(result), }; }), ); return results.map((result, index) => { if (result.status === "fulfilled") { return result.value; } // Promise.allSettled catches individual failures return { role: "tool" as const, toolCallId: toolCalls[index].id, content: JSON.stringify({ success: false, error: `Tool execution failed: ${result.reason}`, }), }; }); } // Integration with the tool loop async function runToolLoopWithParallel( userMessage: string, tools: ToolDefinition[], systemPrompt: string, ): Promise<string> { const messages: Message[] = [ { role: "system", content: systemPrompt }, { role: "user", content: userMessage }, ]; const MAX_STEPS = 10; for (let step = 0; step < MAX_STEPS; step++) { const response = await callLLM({ messages, tools }); if (!response.toolCalls?.length) { return response.text; } messages.push({ role: "assistant", content: response.text ?? null, toolCalls: response.toolCalls, }); // Execute all tool calls concurrently const toolResults = await executeToolCallsInParallel(response.toolCalls); messages.push(...toolResults); } throw new Error(`Tool loop exceeded ${MAX_STEPS} steps`); } ``` **Why good:** `Promise.allSettled` executes all calls concurrently and handles individual failures without aborting others. Each tool result is matched to its `toolCallId` regardless of completion order. Failed calls return structured errors so the model can recover. ### Bad Example -- Sequential Execution of Parallel Calls ```typescript // BAD: Sequential execution wastes time for (const toolCall of response.toolCalls) { const result = await executeTool(toolCall); // Each waits for the previous messages.push({ role: "tool", toolCallId: toolCall.id, content: JSON.stringify(result), }); } ``` **Why bad:** Sequential execution negates the performance benefit of parallel tool calls. If each call takes 300ms, 4 calls take 1200ms instead of 300ms. --- ## Pattern 7: Multi-Step Agent with Conversation State A complete agent that manages conversation state across multiple tool call rounds, with per-step tool filtering and graceful termination. ### Good Example -- Agent with Step-Aware Tool Control ```typescript // agent/multi-step.ts const MAX_AGENT_STEPS = 15; const SUMMARIZE_AFTER_STEPS = 10; interface AgentConfig { systemPrompt: string; tools: ToolRegistry; /** Optional: restrict which tools are available at each step */ getActiveTools?: (step: number, messages: Message[]) => string[]; } async function runAgent( userMessage: string, config: AgentConfig, ): Promise<AgentResult> { const messages: Message[] = [ { role: "system", content: config.systemPrompt }, { role: "user", content: userMessage }, ]; const allToolDefs = config.tools.toDefinitions(); for (let step = 0; step < MAX_AGENT_STEPS; step++) { // Per-step tool filtering (optional) const activeToolNames = config.getActiveTools?.(step, messages); const tools = activeToolNames ? config.tools.subset(activeToolNames) : allToolDefs; // Context management: summarize conversation if getting long if (step === SUMMARIZE_AFTER_STEPS) { const summary = await summarizeConversation(messages); // Replace middle messages with summary, keep system + first user + recent messages.splice(2, messages.length - 4, { role: "system", content: `Previous conversation summary: ${summary}`, }); } const response = await callLLM({ messages, tools, toolChoice: "auto" }); if (!response.toolCalls?.length) { return { text: response.text, steps: step, messages, }; } messages.push({ role: "assistant", content: response.text ?? null, toolCalls: response.toolCalls, }); const results = await executeToolCallsInParallel(response.toolCalls); messages.push(...results); } // Graceful termination const finalResponse = await callLLM({ messages: [ ...messages, { role: "user", content: "Please provide your final answer based on all information gathered.", }, ], tools: [], toolChoice: "none", }); return { text: finalResponse.text, steps: MAX_AGENT_STEPS, messages }; } ``` **Why good:** Per-step tool filtering via `getActiveTools` reduces token cost and prevents irrelevant tool calls. Conversation summarization prevents context overflow in long runs. Graceful termination asks the model for a final answer instead of crashing. Parallel tool execution for each step. --- ## Pattern 8: Human-in-the-Loop Approval For dangerous or irreversible operations, pause the agent loop and wait for human approval before executing. ### Good Example -- Approval Gate Pattern ```typescript // agent/approval.ts type ApprovalDecision = "approve" | "reject" | "modify"; interface ApprovalRequest { toolName: string; arguments: unknown; reason: string; } interface ApprovalResult { decision: ApprovalDecision; modifiedArguments?: unknown; rejectionReason?: string; } // Tools that require human approval before execution const DANGEROUS_TOOLS = new Set([ "delete_record", "send_email", "execute_sql", "deploy_service", "transfer_funds", ]); async function executeWithApproval( toolCall: ToolCall, requestApproval: (req: ApprovalRequest) => Promise<ApprovalResult>, ): Promise<ToolResult> { // Non-dangerous tools execute immediately if (!DANGEROUS_TOOLS.has(toolCall.name)) { return executeTool(toolCall); } // Request human approval const approval = await requestApproval({ toolName: toolCall.name, arguments: toolCall.arguments, reason: `Agent wants to call "${toolCall.name}" with these arguments.`, }); switch (approval.decision) { case "approve": return executeTool(toolCall); case "modify": // Execute with human-modified arguments return executeTool({ ...toolCall, arguments: approval.modifiedArguments ?? toolCall.arguments, }); case "reject": // Return rejection to the model so it can adjust return { success: false, error: `Action "${toolCall.name}" was rejected by the user. ` + `Reason: ${approval.rejectionReason ?? "No reason given"}. ` + "Please suggest an alternative approach.", }; } } ``` **Why good:** Clear separation between safe tools (auto-execute) and dangerous tools (require approval). Three decision options: approve, modify (change arguments), reject. Rejection returns a structured error to the model with the reason, so it can adjust. Named constant set for dangerous tools. #### Integration with the Tool Loop ```typescript // In the tool loop, replace executeTool with executeWithApproval: for (const toolCall of response.toolCalls) { const result = await executeWithApproval(toolCall, promptUserForApproval); messages.push({ role: "tool", toolCallId: toolCall.id, content: JSON.stringify(result), }); } ``` --- ## Pattern 9: Streaming with Tool Calls When streaming LLM responses, tool call arguments arrive as partial JSON chunks that must be accumulated before parsing. ### Good Example -- Accumulating Streamed Tool Calls ```typescript // agent/streaming.ts interface StreamedToolCall { id: string; name: string; argumentChunks: string[]; } async function processToolCallStream( stream: AsyncIterable<StreamChunk>, ): Promise<{ text: string; toolCalls: ToolCall[] }> { let text = ""; const partialToolCalls = new Map<number, StreamedToolCall>(); for await (const chunk of stream) { // Text content if (chunk.type === "text-delta") { text += chunk.text; // Optionally emit to UI: onTextDelta(chunk.text) continue; } // Tool call start -- model declares it wants to call a tool if (chunk.type === "tool-call-start") { partialToolCalls.set(chunk.index, { id: chunk.toolCallId, name: chunk.toolName, argumentChunks: [], }); continue; } // Tool call argument delta -- partial JSON arrives incrementally if (chunk.type === "tool-call-delta") { const partial = partialToolCalls.get(chunk.index); if (partial) { partial.argumentChunks.push(chunk.argumentDelta); } continue; } } // Assemble complete tool calls from accumulated chunks const toolCalls: ToolCall[] = Array.from(partialToolCalls.values()).map( (partial) => ({ id: partial.id, name: partial.name, arguments: partial.argumentChunks.join(""), }), ); return { text, toolCalls }; } ``` **Why good:** Accumulates partial argument chunks per tool call index. Handles interleaved text and tool call deltas. Assembles complete tool calls only after the stream ends. Supports multiple concurrent tool calls (indexed by position). ### Bad Example -- Parsing Each Chunk Individually ```typescript // BAD: Treating each argument chunk as a complete tool call for await (const chunk of stream) { if (chunk.type === "tool-call-delta") { // BAD: Partial JSON is not valid -- JSON.parse will throw const args = JSON.parse(chunk.argumentDelta); await executeTool({ name: chunk.toolName, arguments: args }); } } ``` **Why bad:** Argument deltas are partial JSON fragments (e.g., `{"loc` then `ation":` then `"London"}`). Parsing each chunk individually will always fail. Must accumulate all chunks before parsing. --- ## Pattern 10: Security -- Input Validation and Sandboxing LLM-generated tool arguments are untrusted input. Apply defense-in-depth: validate schemas, enforce permissions, sandbox execution, and rate limit. ### Good Example -- Defense-in-Depth Tool Execution ```typescript // agent/security.ts import { z } from "zod"; const TOOL_TIMEOUT_MS = 30_000; const MAX_TOOL_CALLS_PER_MINUTE = 30; interface SecureToolConfig { schema: z.ZodType; execute: (args: unknown) => Promise<unknown>; /** Maximum execution time in ms */ timeoutMs?: number; /** Allowed operations (for tools that access external resources) */ allowedDomains?: string[]; /** Require human approval */ requiresApproval?: boolean; } async function executeSecurely( toolCall: ToolCall, config: SecureToolConfig, ): Promise<ToolResult> { // 1. Schema validation (reject malformed arguments) const validation = config.schema.safeParse(toolCall.arguments); if (!validation.success) { return { success: false, error: `Validation failed: ${validation.error.message}`, }; } // 2. Domain allowlist (prevent exfiltration via URL arguments) if (config.allowedDomains && "url" in validation.data) { const url = new URL(validation.data.url as string); if (!config.allowedDomains.includes(url.hostname)) { return { success: false, error: `Domain "${url.hostname}" is not in the allowlist.`, }; } } // 3. Timeout (prevent hung tools from blocking the agent) const timeoutMs = config.timeoutMs ?? TOOL_TIMEOUT_MS; try { const result = await Promise.race([ config.execute(validation.data), new Promise((_, reject) => setTimeout( () => reject(new Error("Tool execution timed out")), timeoutMs, ), ), ]); return { success: true, data: result }; } catch (error) { return { success: false, error: error instanceof Error ? error.message : "Execution failed", }; } } ``` **Why good:** Schema validation before execution catches malformed arguments. Domain allowlist prevents data exfiltration through URL parameters. Timeout prevents hung tools from blocking the agent indefinitely. Named constants for timeout and rate limits. ### Security Checklist ```typescript // Defense-in-depth layers for tool execution: // // 1. VALIDATE: Parse arguments with Zod schema (reject malformed input) // 2. AUTHORIZE: Check that the tool is in the allowed set for this context // 3. CONSTRAIN: Enforce allowlists for URLs, file paths, SQL tables // 4. TIMEOUT: Wrap execution in a timeout to prevent hangs // 5. SANDBOX: Run code execution tools in isolated environments // 6. RATE LIMIT: Cap tool calls per time window to prevent abuse // 7. AUDIT: Log every tool call with arguments and results // 8. APPROVE: Require human approval for destructive operations ``` --- ## Pattern 11: Per-Step Tool Filtering Reduce token cost and improve model focus by only sending relevant tools at each step. A research agent's first step needs `search`, not `send_email`. ### Good Example -- Step-Aware Tool Selection ```typescript // agent/tool-filter.ts type AgentPhase = "research" | "analysis" | "action" | "report"; const PHASE_TOOLS: Record<AgentPhase, string[]> = { research: ["search_web", "search_docs", "fetch_url"], analysis: ["analyze_data", "compare_options", "calculate"], action: ["create_record", "update_record", "send_notification"], report: ["format_report", "generate_chart"], }; function determinePhase(step: number, messages: Message[]): AgentPhase { // Simple heuristic: research first, then analyze, then act, then report const toolCallCount = messages.filter((m) => m.role === "tool").length; if (toolCallCount === 0) return "research"; if (step < 5) return "analysis"; if (step < 10) return "action"; return "report"; } // Used with getActiveTools in the agent config: const agentConfig: AgentConfig = { systemPrompt: "You are a research assistant...", tools: registry, getActiveTools: (step, messages) => { const phase = determinePhase(step, messages); return PHASE_TOOLS[phase]; }, }; ``` **Why good:** Reduces token cost by only sending relevant tools. Prevents the model from calling irrelevant tools (e.g., `send_email` during research). Phase-based filtering matches natural agent workflow. Named constant for phase-tool mapping. --- ## Pattern 12: Tool Result Caching For idempotent tools (same input always produces same output), cache results to avoid redundant API calls when the model retries. ### Good Example -- Cache by Tool Name + Arguments ```typescript // agent/cache.ts const CACHE_TTL_MS = 60_000; const toolResultCache = new Map< string, { result: ToolResult; timestamp: number } >(); function getCacheKey(toolCall: ToolCall): string { return `${toolCall.name}:${JSON.stringify(toolCall.arguments)}`; } async function executeWithCache( toolCall: ToolCall, idempotentTools: Set<string>, ): Promise<ToolResult> { // Only cache idempotent tools (read-only operations) if (!idempotentTools.has(toolCall.name)) { return executeTool(toolCall); } const key = getCacheKey(toolCall); const cached = toolResultCache.get(key); if (cached && Date.now() - cached.timestamp < CACHE_TTL_MS) { return cached.result; } const result = await executeTool(toolCall); if (result.success) { toolResultCache.set(key, { result, timestamp: Date.now() }); } return result; } const IDEMPOTENT_TOOLS = new Set([ "search_database", "get_user", "fetch_weather", "calculate", ]); ``` **Why good:** Only caches idempotent (read-only) tools -- write operations always execute. TTL prevents stale results. Cache key includes arguments for correct invalidation. Failed results are not cached (retry might succeed). Named constant for TTL. -
core.md 13.9 KB
# Tool Use Patterns -- Core Examples > Core patterns for tool definitions, the tool call loop, error handling, and type-safe tools. See [advanced.md](advanced.md) for parallel calls, multi-step agents, human-in-the-loop, and security. **Prerequisites**: Understand the tool calling flow: define tools -> send to LLM -> detect tool calls -> execute -> return results -> repeat. --- ## Pattern 1: Complete Tool Definition with JSON Schema Tool definitions tell the model what functions are available, what they do, and what arguments they accept. The schema uses JSON Schema format, which all major providers require. ### Good Example -- Precise, Constrained Definition ```typescript // tools/definitions.ts const MAX_SEARCH_RESULTS = 50; const VALID_SORT_OPTIONS = ["relevance", "date", "popularity"] as const; const searchDocsTool: ToolDefinition = { name: "search_documentation", description: "Search the project documentation for pages matching a query. " + "Returns titles, URLs, and relevance scores. " + "Use when the user asks about project features, configuration, or troubleshooting. " + "Do NOT use for general programming questions unrelated to this project.", parameters: { type: "object", properties: { query: { type: "string", description: "Natural language search query. Be specific -- " + "'authentication setup' works better than 'auth'.", minLength: 1, maxLength: 200, }, section: { type: "string", description: "Limit search to a specific docs section. " + "One of: 'guides', 'api-reference', 'troubleshooting', 'changelog'.", enum: ["guides", "api-reference", "troubleshooting", "changelog"], }, limit: { type: "integer", description: `Number of results to return. Max ${MAX_SEARCH_RESULTS}.`, minimum: 1, maximum: MAX_SEARCH_RESULTS, default: 10, }, sortBy: { type: "string", description: "How to order results.", enum: VALID_SORT_OPTIONS, default: "relevance", }, }, required: ["query"], additionalProperties: false, }, }; ``` **Why good:** Description explains what the tool returns, when to use it, and when NOT to use it. Each parameter has a description with examples. Numeric parameters have min/max constraints. Enum parameters restrict values to valid options. `additionalProperties: false` prevents hallucinated extra fields. Named constants for limits. ### Bad Example -- Vague, Unconstrained Definition ```typescript // BAD: Everything wrong with tool definitions const badSearchTool: ToolDefinition = { name: "search", // Too generic, collides with other tools description: "Searches for stuff", // Useless description parameters: { type: "object", properties: { q: { type: "string" }, // Cryptic name, no description n: { type: "number" }, // No constraints, could be negative or 99999 sort: { type: "string" }, // No enum, model will hallucinate values }, // No required fields -- model might omit the query entirely }, }; ``` **Why bad:** Generic name risks collision. No description guidance means wrong tool selection. Cryptic parameter names produce worse argument quality. No constraints means invalid values (negative limit, unknown sort). No `required` means the model might omit essential fields. --- ## Pattern 2: The Tool Call Loop -- Complete Implementation The full implementation of the request-tool-respond cycle with proper error handling and conversation state management. ### Good Example -- Bounded Loop with Full State ```typescript // agent/tool-loop.ts import type { Message, ToolCall, LLMResponse } from "./types"; const MAX_TOOL_STEPS = 10; interface ToolLoopResult { text: string; toolCallsExecuted: number; messages: Message[]; } async function runToolLoop( userMessage: string, tools: ToolDefinition[], systemPrompt: string, ): Promise<ToolLoopResult> { const messages: Message[] = [ { role: "system", content: systemPrompt }, { role: "user", content: userMessage }, ]; let toolCallsExecuted = 0; for (let step = 0; step < MAX_TOOL_STEPS; step++) { const response: LLMResponse = await callLLM({ messages, tools, toolChoice: "auto", }); // Model responded with text (no tool calls) -- we're done if (!response.toolCalls || response.toolCalls.length === 0) { return { text: response.text, toolCallsExecuted, messages, }; } // CRITICAL: Append the assistant message WITH tool calls to history // before appending tool results. The message sequence must be: // assistant (with tool_calls) -> tool result(s) -> next assistant messages.push({ role: "assistant", content: response.text ?? null, toolCalls: response.toolCalls, }); // Execute each tool call and append results for (const toolCall of response.toolCalls) { const result = await executeTool(toolCall); messages.push({ role: "tool", toolCallId: toolCall.id, content: JSON.stringify(result), }); toolCallsExecuted++; } } // Exhausted step limit -- ask the model to wrap up without tools messages.push({ role: "user", content: "You have reached the maximum number of tool calls. " + "Please provide your best answer based on the information gathered so far.", }); const finalResponse = await callLLM({ messages, tools, toolChoice: "none", // Force text-only response }); return { text: finalResponse.text, toolCallsExecuted, messages, }; } ``` **Why good:** Bounded loop with named constant `MAX_TOOL_STEPS`. Assistant message (with tool calls) is appended before tool results, maintaining correct message sequence. Graceful degradation when step limit is reached (asks model to summarize). Returns metadata (calls executed, full message history) for logging. Uses `toolChoice: "none"` to force final text response. ### Bad Example -- Broken Message Sequence ```typescript // BAD: Missing assistant message in conversation history async function brokenLoop(message: string): Promise<string> { const messages = [{ role: "user", content: message }]; for (let i = 0; i < 5; i++) { const response = await callLLM({ messages, tools }); if (!response.toolCalls?.length) return response.text; // BAD: Tool results added without the preceding assistant message // Most providers will reject this message sequence for (const tc of response.toolCalls) { const result = await executeTool(tc); messages.push({ role: "tool", toolCallId: tc.id, content: JSON.stringify(result), }); } } return "Failed"; // BAD: generic error, no graceful degradation } ``` **Why bad:** Missing assistant message before tool results breaks the message sequence (most providers reject this). Generic "Failed" string instead of asking the model to summarize. No metadata returned for debugging. --- ## Pattern 3: Tool Execution with Validation and Error Reporting The tool executor validates arguments, catches errors, and returns structured results the model can reason about. ### Good Example -- Full Validation Pipeline ```typescript // agent/executor.ts import { z } from "zod"; interface ToolHandler { schema: z.ZodType; execute: (args: unknown) => Promise<unknown>; } interface ToolResult { success: boolean; data?: unknown; error?: string; } const toolRegistry: Record<string, ToolHandler> = {}; function registerTool( name: string, schema: z.ZodType, execute: (args: unknown) => Promise<unknown>, ): void { toolRegistry[name] = { schema, execute }; } async function executeTool(toolCall: ToolCall): Promise<ToolResult> { // 1. Validate tool exists const handler = toolRegistry[toolCall.name]; if (!handler) { return { success: false, error: `Unknown tool: "${toolCall.name}". ` + `Available tools: ${Object.keys(toolRegistry).join(", ")}`, }; } // 2. Parse arguments (they arrive as a JSON string) let parsedArgs: unknown; try { parsedArgs = typeof toolCall.arguments === "string" ? JSON.parse(toolCall.arguments) : toolCall.arguments; } catch { return { success: false, error: `Invalid JSON in arguments for "${toolCall.name}": ${toolCall.arguments}`, }; } // 3. Validate arguments against schema const validation = handler.schema.safeParse(parsedArgs); if (!validation.success) { const issues = validation.error.issues .map((i) => `${i.path.join(".")}: ${i.message}`) .join("; "); return { success: false, error: `Invalid arguments for "${toolCall.name}": ${issues}`, }; } // 4. Execute with error boundary try { const data = await handler.execute(validation.data); return { success: true, data }; } catch (error) { const message = error instanceof Error ? error.message : "Unknown execution error"; return { success: false, error: `Tool "${toolCall.name}" execution failed: ${message}`, }; } } ``` **Why good:** Four-step validation pipeline: tool exists -> JSON parses -> schema validates -> execution succeeds. Each failure returns a descriptive error the model can act on. Arguments are parsed from JSON string (how they arrive from most providers). Zod `safeParse` gives field-level error messages. Error boundary catches execution failures without crashing the loop. ### Bad Example -- No Validation, Crashes on Error ```typescript // BAD: Everything that can go wrong async function unsafeExecute(toolCall: ToolCall): Promise<unknown> { const handler = toolRegistry[toolCall.name]!; // Crashes if missing const args = JSON.parse(toolCall.arguments); // Crashes on invalid JSON return handler.execute(args); // No schema validation, throws on error } ``` **Why bad:** Non-null assertion crashes on unknown tool. Unhandled JSON.parse throws on malformed arguments. No schema validation lets invalid data reach execution. Unhandled execution errors crash the entire agent loop. --- ## Pattern 4: Tool Registry Pattern A typed registry that maps tool names to their handlers, supporting dynamic tool registration and serialization to API format. ### Good Example -- Typed Registry with Serialization ```typescript // tools/registry.ts import { z } from "zod"; import { zodToJsonSchema } from "zod-to-json-schema"; interface Tool<T extends z.ZodType = z.ZodType> { name: string; description: string; schema: T; execute: (args: z.infer<T>) => Promise<unknown>; } class ToolRegistry { private tools = new Map<string, Tool>(); register<T extends z.ZodType>(config: { name: string; description: string; schema: T; execute: (args: z.infer<T>) => Promise<unknown>; }): void { this.tools.set(config.name, config); } get(name: string): Tool | undefined { return this.tools.get(name); } /** Serialize all tools to the JSON Schema format LLM providers expect */ toDefinitions(): ToolDefinition[] { return Array.from(this.tools.values()).map((tool) => ({ name: tool.name, description: tool.description, parameters: zodToJsonSchema(tool.schema, { target: "openApi3", $refStrategy: "none", }), })); } /** Get a subset of tools by name (for per-step filtering) */ subset(names: string[]): ToolDefinition[] { return this.toDefinitions().filter((t) => names.includes(t.name)); } listNames(): string[] { return Array.from(this.tools.keys()); } } // Usage const registry = new ToolRegistry(); registry.register({ name: "get_user", description: "Look up a user by their ID. Returns name, email, and role.", schema: z.object({ userId: z.string().uuid().describe("The user's unique identifier"), }), execute: async ({ userId }) => { return db.users.findById(userId); }, }); // Pass to LLM const response = await callLLM({ messages, tools: registry.toDefinitions(), }); ``` **Why good:** Type-safe registration with Zod schema + inferred execute arguments. `toDefinitions()` serializes to the format LLM providers expect. `subset()` enables per-step tool filtering (reduce token cost). Single source of truth for tool metadata and execution logic. --- ## Pattern 5: Formatting Tool Results Tool results should be concise, structured, and include only information the model needs to formulate its response. ### Good Example -- Summarized, Structured Results ```typescript // tools/result-formatter.ts const MAX_RESULT_LENGTH = 2000; function formatToolResult(data: unknown): string { const json = JSON.stringify(data, null, 2); // Truncate oversized results to prevent context overflow if (json.length > MAX_RESULT_LENGTH) { const truncated = json.slice(0, MAX_RESULT_LENGTH); return ( truncated + `\n... [truncated, ${json.length - MAX_RESULT_LENGTH} chars omitted]` ); } return json; } // For list results, return summary + items function formatListResult( items: unknown[], totalCount: number, limit: number, ): string { return JSON.stringify({ items, showing: items.length, totalAvailable: totalCount, note: items.length < totalCount ? `Showing ${items.length} of ${totalCount}. Ask the user if they want more.` : undefined, }); } ``` **Why good:** Truncation prevents context overflow from oversized results. List results include count and pagination hint. Structured JSON is easier for the model to parse than free text. Named constant for max length. ### Bad Example -- Raw Database Dump ```typescript // BAD: Returning entire database rows async function rawResult(userId: string): Promise<unknown> { // Returns ALL columns including internal IDs, timestamps, soft-delete flags return db.query("SELECT * FROM users WHERE id = $1", [userId]); } ``` **Why bad:** `SELECT *` returns irrelevant internal fields, wasting context tokens. No truncation means a large row set could overflow context. No structure or summary for the model to work with.
-
-
reference.md 8.3 KB
# Tool Use Patterns -- Quick Reference > Decision frameworks, provider comparison, and anti-pattern checklist. See [examples/core.md](examples/core.md) for full implementations. --- ## Provider Tool Format Comparison All providers use JSON Schema for parameters. The wrapping structure differs: | Provider | Tool Wrapper | Parameter Key | Tool Call Response | Result Message | | --------- | ------------------------------------------------------------------- | -------------- | --------------------------------------------------------------------- | ------------------------------------------------------------- | | OpenAI | `{ type: "function", function: { name, description, parameters } }` | `parameters` | `message.tool_calls[]` with `function.arguments` (JSON string) | `{ role: "tool", tool_call_id, content }` | | Anthropic | `{ name, description, input_schema }` | `input_schema` | Content block `{ type: "tool_use", id, name, input }` (parsed object) | Content block `{ type: "tool_result", tool_use_id, content }` | | Google | `FunctionDeclaration { name, description, parameters }` | `parameters` | `function_call { name, id, args }` (parsed object) | `function_response { name, id, response }` | **Key differences:** - OpenAI returns arguments as a **JSON string** -- you must `JSON.parse()` - Anthropic returns arguments as a **parsed object** -- no parsing needed - OpenAI uses `role: "tool"` messages; Anthropic uses `tool_result` content blocks - Tool choice syntax differs: OpenAI `"required"` vs Anthropic `{ type: "any" }` - Google Gemini 3+ models generate a unique `id` for each function call -- include the matching `id` in your `functionResponse` **Recommendation:** Abstract the provider format in a `callLLM()` wrapper that normalizes tool calls into a common shape. All examples in this skill use this normalized format. --- ## Tool Choice Quick Reference | Mode | OpenAI | Anthropic | Google | Effect | | ----------------------- | ------------------------------------------ | ------------------------------- | -------------------------------- | ----------------------------------------------- | | Auto (default) | `"auto"` | `{ type: "auto" }` | `AUTO` | Model decides | | Must call a tool | `"required"` | `{ type: "any" }` | `ANY` | Forces at least one tool call | | No tools | `"none"` | `{ type: "auto" }` + omit tools | `NONE` | Text-only response | | Specific tool | `{ type: "function", function: { name } }` | `{ type: "tool", name }` | `ANY` + `allowed_function_names` | Forces specific tool | | Validated (Google only) | N/A | N/A | `VALIDATED` (preview) | Schema-validated, allows text or function calls | --- ## Tool Definition Checklist For each tool definition, verify: - [ ] **Name** is descriptive and unique (`search_documentation`, not `search`) - [ ] **Description** explains what the tool returns, when to use it, and when NOT to use it - [ ] **Parameters** each have a `description` field - [ ] **Required fields** are listed in `required` array - [ ] **Enums** constrain string parameters to valid values - [ ] **Numeric constraints** use `minimum`, `maximum`, `minLength`, `maxLength` - [ ] **`additionalProperties: false`** prevents hallucinated extra fields - [ ] **Token cost** is reasonable (description is concise but precise) --- ## Common Anti-Patterns | Anti-Pattern | Problem | Fix | | ---------------------------------------------- | ------------------------------------------ | ----------------------------------------------------------- | | `while (true)` loop | Infinite API calls, runaway costs | Use `for` loop with `MAX_STEPS` constant | | No argument validation | Malformed input reaches execution | Validate with Zod before executing | | Silent error swallowing | Model cannot recover from unknown failures | Return structured error with message | | `SELECT *` in tool results | Context overflow with irrelevant data | Select specific fields, truncate large results | | All tools on every call | Token waste, model confusion | Filter tools per step based on context | | No assistant message before tool results | Provider rejects message sequence | Always append assistant message with tool calls first | | Parsing streaming argument chunks individually | JSON.parse fails on partial JSON | Accumulate all chunks, parse once after stream ends | | Direct code execution from arguments | Prompt injection becomes code execution | Sandbox execution, validate arguments, allowlist operations | | Caching write-operation results | Stale data, missed side effects | Only cache idempotent (read-only) tool results | | Generic tool descriptions | Wrong tool selection, malformed arguments | Include what it returns, when to use, when NOT to use | --- ## Token Cost Estimation Tool definitions are sent on **every** API call. Estimate token impact: | Component | Approximate Tokens | | -------------------------------------- | -------------------------- | | Tool name + description (50 words) | ~75 tokens | | Parameter with description (per param) | ~30 tokens | | JSON Schema overhead | ~20 tokens | | **Typical tool (3 params)** | **~185 tokens** | | **10 tools** | **~1,850 tokens per call** | **Optimization strategies:** - Filter tools per step (only send relevant tools) - Keep descriptions concise but precise - Use enums instead of verbose description constraints - Consider combining related tools into one with a `mode` parameter --- ## Message Sequence Diagram The correct message sequence for a tool call round-trip: ``` 1. User message { role: "user", content: "What's the weather?" } 2. Assistant response WITH tool calls { role: "assistant", toolCalls: [{ id: "tc_1", name: "get_weather", arguments: '{"location":"London"}' }] } 3. Tool result(s) -- one per tool call { role: "tool", toolCallId: "tc_1", content: '{"temp": 18, "conditions": "cloudy"}' } 4. Assistant final response (or more tool calls -- loop continues) { role: "assistant", content: "The weather in London is 18C and cloudy." } ``` **Critical:** Step 2 (assistant with tool calls) MUST be in the message history before step 3 (tool results). Omitting it breaks the message sequence for most providers. --- ## Step Limit Guidelines | Agent Complexity | Suggested MAX_STEPS | Rationale | | ---------------------------------- | ------------------- | -------------------------------------- | | Simple lookup (1-2 tools) | 3-5 | Quick data fetch, minimal iteration | | Research agent (search + analyze) | 10-15 | Multiple search rounds, analysis | | Complex workflow (multi-tool) | 15-25 | Multiple phases: research, act, verify | | Code assistant (edit + test + fix) | 20-30 | Edit-test-fix cycles can be lengthy | Always include graceful termination: when the step limit is reached, ask the model to provide its best answer with `toolChoice: "none"`. -
SKILL.md 17.5 KB
--- name: ai-patterns-tool-use-patterns description: Provider-agnostic patterns for LLM function calling, tool loops, and agentic workflows --- # Tool Use Patterns > **Quick Guide:** Tool use (function calling) lets LLMs invoke external functions. The universal pattern is: define tool schemas (JSON Schema for parameters) -> send tools + message to LLM -> detect tool_use in response -> execute locally -> return result to LLM -> repeat until the model responds with text. Guard every loop with a max-step limit, validate all tool inputs before execution, and return structured errors so the model can recover. Use tool choice control (`auto`, `required`, `none`, specific tool) to steer model behavior. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST guard every tool loop with a maximum step limit -- unbounded loops risk infinite API calls and runaway costs)** **(You MUST validate all tool input arguments before execution -- LLM-generated arguments are untrusted input)** **(You MUST return structured error messages to the model when tool execution fails -- never silently swallow errors or return empty results)** **(You MUST use JSON Schema for tool parameter definitions -- all major providers require this format)** **(You MUST treat tool definitions as token cost -- every tool schema is sent on every API call, so keep descriptions concise but precise)** </critical_requirements> --- **Auto-detection:** tool use, function calling, tool_calls, tool_use, tool call loop, agent loop, tool definition, tool schema, toolChoice, tool_choice, parallel tool calls, human-in-the-loop, tool approval, agentic workflow, multi-step agent, tool result, tool error **When to use:** - Implementing LLM tool calling / function calling in any provider - Building agent loops that call tools iteratively until a task is complete - Handling parallel tool calls (multiple tools in one response) - Reporting tool errors back to the model for recovery - Controlling tool selection (auto, required, none, force specific) - Adding human approval gates before dangerous tool execution - Streaming responses that include tool calls **Key patterns covered:** - Tool definition schemas (JSON Schema for parameters, descriptions) - The core tool call loop (send -> detect -> execute -> return -> re-send) - Parallel tool calls (handling multiple calls in one response) - Error handling (reporting tool failures back to the model) - Tool choice control (auto, required, none, specific tool) - Multi-step agent workflows with conversation state - Human-in-the-loop approval patterns - Type-safe tool definitions in TypeScript - Security (input validation, sandboxing, least privilege) - Streaming with tool calls **When NOT to use:** - Simple text generation without tool calling -- no tools needed - Structured output / JSON extraction -- use your provider's structured output feature instead - Provider-specific SDK patterns -- use your provider's SDK skill for SDK-specific APIs **Detailed Resources:** - [examples/core.md](examples/core.md) -- Tool definitions, the tool call loop, error handling, type-safe tools - [examples/advanced.md](examples/advanced.md) -- Parallel tool calls, multi-step agents, human-in-the-loop, streaming, security - [reference.md](reference.md) -- Decision frameworks, provider comparison, anti-pattern checklist --- <philosophy> ## Philosophy Tool use is the mechanism that turns LLMs from text generators into agents. The model cannot execute code, query databases, or call APIs -- it can only _request_ that your code does so by emitting structured tool calls. Your code is the executor; the model is the planner. **Core principles:** 1. **The model plans, you execute** -- The LLM emits tool call requests with structured arguments. Your code validates, executes, and returns results. Never let the model execute arbitrary code directly. 2. **Agents are loops** -- Every agent, from a simple weather bot to a complex coding assistant, follows the same loop: LLM decides -> system executes -> results feed back -> repeat. Complexity comes from the tools and state, not the loop itself. 3. **Tools are schemas** -- A tool definition is a JSON Schema that tells the model what function exists, what parameters it takes, and when to use it. Better descriptions produce better tool selection and argument quality. 4. **Errors are information** -- When a tool fails, return a structured error message to the model. The model can often recover by retrying with different arguments, choosing a different tool, or explaining the failure to the user. 5. **Defense in depth** -- LLM-generated arguments are untrusted input. Validate schemas, enforce types, limit argument ranges, sandbox execution, and require approval for dangerous operations. **When to use tool calling:** - The task requires real-world data the model doesn't have (weather, database, APIs) - The task requires side effects (sending email, creating records, file operations) - The task requires multi-step reasoning with intermediate data lookups - The task requires computation the model can't do reliably (math, code execution) **When NOT to use tool calling:** - The model can answer from its training data alone - You only need structured JSON output (use structured output features instead) - The "tool" is just prompt engineering disguised as a function - You want deterministic behavior (tool calling adds non-determinism from model decisions) </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Tool Definition Schema Every tool definition has three parts: a name, a description, and a parameter schema. The description is the most important part -- it guides the model's decision to call the tool and how it constructs arguments. ```typescript // Provider-agnostic tool definition shape interface ToolDefinition { name: string; // ^[a-zA-Z0-9_-]{1,64}$ description: string; // When and why to use this tool parameters: JsonSchema; // JSON Schema for input arguments } ``` #### Good Tool Definition ```typescript // Good: precise description, constrained parameters, .describe() on each property const VALID_UNITS = ["celsius", "fahrenheit"] as const; const getWeatherTool: ToolDefinition = { name: "get_current_weather", description: "Get the current weather conditions for a specific city. " + "Returns temperature, humidity, and conditions. " + "Use this when the user asks about current weather, NOT forecasts.", parameters: { type: "object", properties: { location: { type: "string", description: "City name and optional country code, e.g. 'London, UK'", }, unit: { type: "string", enum: VALID_UNITS, description: "Temperature unit. Defaults to celsius if not specified.", }, }, required: ["location"], }, }; ``` **Why good:** Description explains what the tool returns, when to use it, and when NOT to use it. Parameters have descriptions and constraints (enum). Named constant for valid values. #### Bad Tool Definition ```typescript // Bad: vague description, no parameter descriptions, no constraints const weatherTool: ToolDefinition = { name: "weather", description: "Gets weather", // BAD: too vague parameters: { type: "object", properties: { loc: { type: "string" }, // BAD: no description, cryptic name u: { type: "string" }, // BAD: no enum constraint }, }, }; ``` **Why bad:** Vague description causes incorrect tool selection, no parameter descriptions cause malformed arguments, no enum constraint causes invalid values, cryptic parameter names confuse the model --- ### Pattern 2: The Core Tool Call Loop The fundamental pattern for tool use. Send a message with tool definitions, check if the response contains tool calls, execute them, return results, and let the model generate a final response. See [examples/core.md](examples/core.md) for the complete implementation. ```typescript const MAX_TOOL_STEPS = 10; for (let step = 0; step < MAX_TOOL_STEPS; step++) { const response = await callLLM({ messages, tools }); if (!response.toolCalls?.length) return response.text; messages.push({ role: "assistant", toolCalls: response.toolCalls }); for (const tc of response.toolCalls) { messages.push({ role: "tool", toolCallId: tc.id, content: JSON.stringify(await executeTool(tc)), }); } } ``` **Key points:** Always use a bounded `for` loop with a named constant (`MAX_TOOL_STEPS`) -- never `while (true)`. Always append the assistant message (with tool calls) before tool results. Include graceful termination when the step limit is reached. --- ### Pattern 3: Tool Execution with Error Handling When a tool fails, return a structured error to the model instead of crashing. The model can often recover by retrying, choosing a different approach, or explaining the failure. See [examples/core.md](examples/core.md) for the complete validation pipeline. ```typescript // Four-step validation: tool exists -> JSON parses -> schema validates -> execution succeeds async function executeTool(toolCall: ToolCall): Promise<ToolResult> { const handler = toolRegistry[toolCall.name]; if (!handler) return { success: false, error: `Unknown tool: "${toolCall.name}"` }; const validated = handler.schema.safeParse(toolCall.arguments); if (!validated.success) return { success: false, error: `Invalid args: ${validated.error}` }; try { return { success: true, data: await handler.execute(validated.data) }; } catch (error) { return { success: false, error: `Tool failed: ${error instanceof Error ? error.message : "Unknown"}`, }; } } ``` **Key points:** Always validate tool existence, parse JSON arguments, validate against schema, and catch execution errors. Return structured `{ success, data?, error? }` results so the model can recover. Never crash the loop on tool errors. --- ### Pattern 4: Tool Choice Control Control whether and how the model uses tools. All major providers support four modes. ```typescript type ToolChoice = | "auto" // Model decides whether to call tools (default) | "required" // Model MUST call at least one tool | "none" // Model MUST NOT call any tools | { name: string }; // Model MUST call this specific tool ``` #### When to Use Each Mode ```typescript // auto (default) -- let the model decide const response = await callLLM({ messages, tools, toolChoice: "auto", }); // required -- force tool use (e.g., first step of an agent must act) const response = await callLLM({ messages, tools, toolChoice: "required", }); // none -- disable tools (e.g., final response must be text only) const response = await callLLM({ messages, tools, toolChoice: "none", }); // specific tool -- force a particular tool (e.g., classification step) const response = await callLLM({ messages, tools, toolChoice: { name: "classify_intent" }, }); ``` **Decision guidance:** | Mode | Use When | | ---------- | --------------------------------------------------------------------- | | `auto` | General-purpose agent -- model decides based on context | | `required` | Agent's first step must always call a tool (e.g., data lookup) | | `none` | Final response generation -- no more tool calls allowed | | `{ name }` | Pipeline step that must use a specific tool (e.g., extract, classify) | **When to use:** When you need to constrain or guarantee tool behavior at specific points in a workflow --- ### Pattern 5: Type-Safe Tool Definitions in TypeScript Use a typed registry pattern with Zod schemas to get compile-time safety for tool definitions and runtime validation for execution. See [examples/core.md](examples/core.md) for the complete `ToolRegistry` class implementation. ```typescript import { z } from "zod"; // Define tool with schema + execute function -- TypeScript infers argument types function defineTool<T extends z.ZodType>(config: { name: string; description: string; schema: T; execute: (args: z.infer<T>) => Promise<unknown>; }) { return config; } ``` **When to use:** Any TypeScript project implementing tool calling -- Zod validates at runtime, TypeScript catches schema/handler mismatches at compile time </patterns> --- <decision_framework> ## Decision Framework ### Do You Need Tool Calling? ``` Does the task require information the model doesn't have? +-- YES -> Tool calling (fetch data from APIs, databases, files) +-- NO -> Does the task require side effects? +-- YES -> Tool calling (send email, create record, execute code) +-- NO -> Do you need structured JSON output? +-- YES -> Use structured output features (NOT tool calling) +-- NO -> Plain text generation, no tools needed ``` ### Which Loop Pattern? ``` How many tools might the model call? +-- Single tool call per request | +-- Simple request-response with one tool execution | +-- No loop needed, just one round-trip +-- Multiple sequential tool calls | +-- Use the bounded tool call loop (Pattern 2) | +-- Set MAX_TOOL_STEPS based on task complexity +-- Multiple parallel tool calls in one response | +-- Execute all tool calls concurrently (Promise.all) | +-- Then return all results and loop +-- Complex multi-step agent +-- Use the bounded loop with conversation state +-- Add human-in-the-loop for dangerous operations +-- Consider per-step tool filtering ``` ### How to Handle Tool Errors? ``` Tool execution failed. What to do? +-- Return structured error to the model | +-- Include error message and context | +-- Model can retry, choose alternative, or explain failure +-- NEVER: silently return empty result +-- NEVER: crash the loop +-- NEVER: retry automatically without telling the model (the model should decide whether to retry) ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Unbounded tool loop (`while (true)`) without a step counter -- risks infinite API calls and runaway costs - Executing tool arguments without validation -- LLM-generated arguments are untrusted input, treat like user input - Swallowing tool errors silently (returning `null` or `{}`) -- the model cannot recover from failures it doesn't know about - Allowing arbitrary code execution from tool arguments without sandboxing -- prompt injection can escalate to code execution - Tool descriptions that say "Gets data" -- vague descriptions cause wrong tool selection and malformed arguments **Medium Priority Issues:** - Sending all tools on every API call when only a subset is relevant -- wastes tokens and confuses the model - Not including the assistant message (with tool calls) in conversation history before tool results -- breaks the message sequence - Using `tool_choice: "required"` without a fallback for when no tool makes sense -- forces meaningless tool calls - Returning raw database rows or full API responses as tool results -- overwhelms context with irrelevant data; summarize or truncate - Not logging tool calls and results -- impossible to debug agent behavior in production **Common Mistakes:** - Forgetting that tool call arguments arrive as a JSON string, not a parsed object -- always `JSON.parse()` before use - Assuming the model will always call tools when tools are available -- with `auto` mode, it may respond with text directly - Treating tool calling as structured output -- they solve different problems (actions vs data extraction) - Putting business logic in tool descriptions instead of tool implementations -- descriptions guide selection, not execution **Gotchas & Edge Cases:** - Tool definitions consume tokens on every API call -- 10 tools with detailed schemas can use 1000+ tokens per request - Parallel tool calls may arrive in any order -- never assume execution order matches definition order - Some models hallucinate tool names or arguments that don't match any definition -- always validate the tool name exists in your registry - Streaming responses with tool calls require accumulating partial JSON chunks before parsing -- the arguments arrive incrementally, not all at once - Returning very large tool results (>4000 tokens) can push the conversation past context limits -- truncate or summarize large results - The model maintains full conversation context including all tool calls and results -- long agent runs accumulate significant token usage - Different providers use different message formats for tool results (`role: "tool"` vs content blocks) -- abstract this in your callLLM wrapper </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST guard every tool loop with a maximum step limit -- unbounded loops risk infinite API calls and runaway costs)** **(You MUST validate all tool input arguments before execution -- LLM-generated arguments are untrusted input)** **(You MUST return structured error messages to the model when tool execution fails -- never silently swallow errors or return empty results)** **(You MUST use JSON Schema for tool parameter definitions -- all major providers require this format)** **(You MUST treat tool definitions as token cost -- every tool schema is sent on every API call, so keep descriptions concise but precise)** **Failure to follow these rules will produce agents that run up API costs in infinite loops, execute unvalidated input, or silently fail without the model being able to recover.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.