ai-provider-cohere-sdk
Official Cohere TypeScript SDK patterns -- CohereClientV2, chat, embeddings, rerank, RAG with citations, tool use, streaming, and model selection
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-provider-cohere-sdk/skills/ai-provider-cohere-sdk
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Cohere SDK Patterns
Quick Guide: Use the
cohere-ainpm package withCohereClientV2for all new Cohere integrations. V2 API requiresmodelon every call. UsechatStreamfor streaming withcontent-deltaevents. Embeddings requireinputTypematching your use case (search_documentfor indexing,search_queryfor querying). Rerank scores documents by relevance. RAG works by passingdocumentstochat()-- the model returns inline citations automatically. Tool use follows a 4-step loop: user message, model returnstool_calls, you execute and return results, model generates cited response.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST use CohereClientV2 (not CohereClient) for all new code -- V2 is the current API with required model parameter)
(You MUST specify inputType on every embed call -- search_document for indexing, search_query for querying -- mismatched types produce garbage similarity scores)
(You MUST handle the tool use loop correctly: append the full assistant message (with tool_calls) to messages, then append tool role results with matching tool_call_id)
(You MUST check finish_reason in responses -- MAX_TOKENS means the output was truncated)
(You MUST never hardcode API keys -- pass via token constructor parameter sourced from environment variables)
</critical_requirements>
Auto-detection: Cohere, cohere-ai, CohereClientV2, CohereClient, command-a, command-r, command-r-plus, embed-v4, rerank-v4, chatStream, content-delta, inputType, search_document, search_query, embeddingTypes, topN, CO_API_KEY, COHERE_API_KEY
When to use:
- Building applications with Cohere Command models (chat, generation, summarization)
- Creating semantic search pipelines with Cohere embeddings
- Adding relevance scoring to search results with Cohere Rerank
- Implementing RAG with inline document grounding and automatic citations
- Building agentic workflows with Cohere tool use / function calling
- Streaming chat responses for real-time user interfaces
Key patterns covered:
- Client setup with
CohereClientV2(token, timeout, platform configs) - Chat and streaming (
chat,chatStream, event types) - Embeddings with
inputTypefor search/classification/clustering - Rerank for relevance scoring and search result ordering
- RAG with documents and automatic citation handling
- Tool use / function calling with multi-step loops
- Model selection (Command-A, Command-R, Embed v4, Rerank v4)
When NOT to use:
- Multi-provider applications needing OpenAI/Anthropic/Google switching -- use a unified provider SDK
- React-specific chat UI hooks -- use a framework-integrated AI SDK
- Simple text completion without Cohere-specific features (rerank, citations)
Examples Index
- Core: Setup, Chat & Error Handling -- CohereClientV2 init, basic chat, streaming, error handling
- Embeddings & Rerank -- Semantic search, input types, rerank scoring, RAG pipeline
- Tool Use & RAG -- Function calling, document grounding, citation handling
- Quick API Reference -- Model IDs, method signatures, event types, error classes
<decision_framework>
Decision Framework
Which Client Class to Use
New project?
+-- YES -> CohereClientV2 (always)
+-- Existing V1 code?
+-- Working fine? -> Keep CohereClient but plan migration
+-- Need V2 features? -> Migrate to CohereClientV2
Which Model to Choose
What is your task?
+-- General chat/generation -> command-a-03-2025 (most capable)
+-- Reasoning / multi-step -> command-a-reasoning-08-2025
+-- Image/document analysis -> command-a-vision-07-2025
+-- Translation -> command-a-translate-08-2025
+-- Lightweight / low latency -> command-r7b-12-2024
+-- Embeddings -> embed-v4.0 (or embed-english-v3.0 for English-only)
+-- Rerank quality -> rerank-v4.0-pro
+-- Rerank speed -> rerank-v4.0-fast
Embed inputType Selection
What are you embedding?
+-- Documents for a search index -> "search_document"
+-- Search queries against an index -> "search_query"
+-- Text for a classifier -> "classification"
+-- Text for clustering -> "clustering"
+-- Images -> "image" (embed-v4+ only)
When to Use Rerank
Do you have search results to re-order?
+-- YES -> Use rerank as a second-stage ranker
| +-- Quality matters most? -> rerank-v4.0-pro
| +-- Latency matters most? -> rerank-v4.0-fast
+-- NO -> Not applicable (rerank needs existing results to score)
RAG Approach
Do you need grounded answers with citations?
+-- YES -> Pass documents to chat()
| +-- Have pre-retrieved documents? -> Pass directly via documents param
| +-- Need retrieval first? -> Use embed + vector search + rerank pipeline, then pass top results to chat()
+-- NO -> Use plain chat without documents
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Using
CohereClientinstead ofCohereClientV2for new code (V1 is legacy) - Missing
modelparameter in V2 API calls (required on every call, unlike V1) - Using wrong
inputTypefor embeddings (search_queryfor documents or vice versa -- silently degrades results) - Hardcoding API keys instead of using environment variables
- Not appending the full assistant message (with
tool_calls) before appending tool results in the tool use loop
Medium Priority Issues:
- Not specifying
embeddingTypes(defaults may not match your storage format) - Ignoring
finish_reason: "MAX_TOKENS"(output was silently truncated) - Not handling
CohereTimeoutErrorseparately fromCohereError - Processing all stream events without checking
type(onlycontent-deltahas text) - Using V1 parameter names (
preamble,connectors,conversation_id) with V2 client
Common Mistakes:
- Accessing
response.textinstead ofresponse.message.content[0].text(V2 response shape changed) - Forgetting that
embeddingTypesis required in V2 Embed API - Not matching
tool_call_idwhen submitting tool results (model cannot correlate results) - Using
documentswith string values instead of{ data: { text: "..." } }objects in V2 - Expecting
response.message.citationsto exist when no documents were provided (citations only appear with grounded responses)
Gotchas & Edge Cases:
- The SDK is in beta -- pin your
cohere-aiversion in package.json to avoid breaking changes - V2 API is NOT yet supported for cloud deployments (Bedrock, SageMaker, Azure, OCI) -- use V1 client for cloud platforms
inputTypeis camelCase in TypeScript SDK (inputType) but snake_case in the REST API (input_type)- Embed API accepts max 96 texts per call -- batch larger sets yourself
embed-v4.0supportsoutputDimensionfor flexible sizing (256, 512, 1024, 1536) but v3 models have fixed dimensions- Rerank
relevanceScoreis normalized 0-1 but not calibrated across queries -- compare scores within a single query only - Stream events include
tool-plan-deltabeforetool-call-start-- the model's reasoning about which tool to call - V2 uses
systemrole for instructions (V1 usedpreambleparameter) - Citation
sourcesin tool use responses referencetool_call_idvalues, not document indices - The
clientNameconstructor parameter is for logging/analytics, not authentication responseFormat: { type: "json_object" }is NOT supported in RAG mode (withdocuments,tools, ortoolResults)toolChoiceis only supported oncommand-r7b-12-2024and newer models- First requests with
strictTools: trueand a new tool set take longer (schema compilation) thinking(reasoning mode) is only available on reasoning-capable models likecommand-a-reasoning-08-2025
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST use CohereClientV2 (not CohereClient) for all new code -- V2 is the current API with required model parameter)
(You MUST specify inputType on every embed call -- search_document for indexing, search_query for querying -- mismatched types produce garbage similarity scores)
(You MUST handle the tool use loop correctly: append the full assistant message (with tool_calls) to messages, then append tool role results with matching tool_call_id)
(You MUST check finish_reason in responses -- MAX_TOKENS means the output was truncated)
(You MUST never hardcode API keys -- pass via token constructor parameter sourced from environment variables)
Failure to follow these rules will produce broken embeddings, missing citations, or insecure AI integrations.
</critical_reminders>
Files (skills)
-
examples
-
core.md 5.9 KB
# Cohere SDK -- Setup, Chat & Streaming Examples > Client initialization, chat completions, streaming, multi-turn conversations, and error handling. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [embeddings-rerank.md](embeddings-rerank.md) -- Embeddings, rerank, semantic search pipeline - [tools-rag.md](tools-rag.md) -- Tool use, RAG with documents, citation handling --- ## Basic Client Setup ```typescript // lib/cohere.ts import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY, }); export { client }; ``` --- ## Production Configuration ```typescript // lib/cohere.ts import { CohereClientV2 } from "cohere-ai"; const TIMEOUT_MS = 30_000; const client = new CohereClientV2({ token: process.env.CO_API_KEY, timeout: TIMEOUT_MS, clientName: "my-app", }); export { client }; ``` --- ## Basic Chat Completion ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY }); async function chat(userMessage: string): Promise<string> { const response = await client.chat({ model: "command-a-03-2025", messages: [ { role: "system", content: "You are a helpful assistant. Be concise." }, { role: "user", content: userMessage }, ], }); const content = response.message.content[0].text; console.log( `Tokens: input=${response.usage?.tokens?.inputTokens}, output=${response.usage?.tokens?.outputTokens}`, ); if (response.finishReason === "MAX_TOKENS") { console.warn("Response was truncated"); } return content; } const answer = await chat("What is TypeScript in one sentence?"); console.log(answer); ``` --- ## Multi-Turn Conversation ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY }); const messages: Array<{ role: string; content: string }> = [ { role: "system", content: "You are a TypeScript expert." }, { role: "user", content: "What is a union type?" }, ]; const response = await client.chat({ model: "command-a-03-2025", messages, }); // Append assistant response for next turn const assistantText = response.message.content[0].text; messages.push({ role: "assistant", content: assistantText }); messages.push({ role: "user", content: "Give me a real-world example." }); const followUp = await client.chat({ model: "command-a-03-2025", messages, }); console.log(followUp.message.content[0].text); ``` --- ## Streaming with Content Deltas ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY }); const stream = await client.chatStream({ model: "command-a-03-2025", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: "Explain async/await in TypeScript." }, ], }); for await (const event of stream) { if (event.type === "content-delta") { process.stdout.write(event.delta?.message?.content?.text ?? ""); } if (event.type === "message-end") { console.log(`\nFinish reason: ${event.delta?.finishReason}`); } } ``` --- ## Streaming with All Event Types ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY }); const stream = await client.chatStream({ model: "command-a-03-2025", messages: [{ role: "user", content: "Tell me about TypeScript." }], }); for await (const event of stream) { switch (event.type) { case "message-start": // Stream has begun break; case "content-delta": process.stdout.write(event.delta?.message?.content?.text ?? ""); break; case "citation-start": // Citation generated (only with documents/tools) break; case "tool-plan-delta": // Model reasoning about which tool to call break; case "tool-call-start": // Tool call initiated break; case "message-end": console.log(`\nDone. Reason: ${event.delta?.finishReason}`); break; } } ``` --- ## Production Error Handling ```typescript import { CohereClientV2, CohereError, CohereTimeoutError } from "cohere-ai"; const TIMEOUT_MS = 30_000; const client = new CohereClientV2({ token: process.env.CO_API_KEY, timeout: TIMEOUT_MS, }); async function safeChat(prompt: string): Promise<string | null> { try { const response = await client.chat({ model: "command-a-03-2025", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: prompt }, ], }); if (response.finishReason === "MAX_TOKENS") { console.warn("Response was truncated"); } return response.message.content[0].text; } catch (error) { if (error instanceof CohereTimeoutError) { console.error("Request timed out"); return null; } if (error instanceof CohereError) { console.error(`Cohere API Error [${error.statusCode}]: ${error.message}`); if (error.statusCode === 429) { console.error("Rate limited -- back off and retry"); } if (error.statusCode === 401) { throw new Error( "Invalid API key. Check CO_API_KEY environment variable.", ); } return null; } // Unknown errors should be re-thrown throw error; } } const result = await safeChat("Hello!"); if (result) { console.log(result); } ``` --- ## Temperature and Output Control ```typescript const MAX_TOKENS = 500; const response = await client.chat({ model: "command-a-03-2025", messages: [{ role: "user", content: "Summarize this article." }], temperature: 0, // Deterministic output maxTokens: MAX_TOKENS, }); if (response.finishReason === "MAX_TOKENS") { console.warn("Output was truncated -- increase maxTokens"); } ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
embeddings-rerank.md 6.6 KB
# Cohere SDK -- Embeddings & Rerank Examples > Embedding generation, input type pairing, cosine similarity, rerank scoring, and semantic search pipeline. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, chat, streaming, error handling - [tools-rag.md](tools-rag.md) -- Tool use, RAG with documents, citation handling --- ## Basic Embedding Generation ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY }); const EMBEDDING_MODEL = "embed-v4.0"; // Embed documents for storage in a vector database const response = await client.embed({ model: EMBEDDING_MODEL, inputType: "search_document", texts: [ "TypeScript is a typed superset of JavaScript.", "React is a library for building user interfaces.", "Node.js is a JavaScript runtime.", ], embeddingTypes: ["float"], }); // response.embeddings.float is number[][] -- one vector per input text const vectors = response.embeddings.float; console.log( `Generated ${vectors.length} vectors of dimension ${vectors[0].length}`, ); ``` --- ## Correct Input Type Pairing The most critical gotcha with Cohere embeddings: `inputType` must match between indexing and querying. ```typescript const EMBEDDING_MODEL = "embed-v4.0"; // INDEXING: Use "search_document" for documents going into your vector store const docEmbeddings = await client.embed({ model: EMBEDDING_MODEL, inputType: "search_document", // Documents being indexed texts: documents, embeddingTypes: ["float"], }); // QUERYING: Use "search_query" for the user's search query const queryEmbedding = await client.embed({ model: EMBEDDING_MODEL, inputType: "search_query", // Query being searched texts: [userQuery], embeddingTypes: ["float"], }); ``` ```typescript // BAD: Using the same inputType for both const docs = await client.embed({ model: EMBEDDING_MODEL, inputType: "search_query", // WRONG: documents should use "search_document" texts: documents, embeddingTypes: ["float"], }); ``` **Why bad:** Cohere trains separate embedding spaces for documents vs queries. Using the wrong type silently produces vectors that don't align well, resulting in degraded search quality. There is no error -- results just get worse. --- ## Cosine Similarity Search ```typescript function cosineSimilarity(a: number[], b: number[]): number { let dotProduct = 0; let normA = 0; let normB = 0; for (let i = 0; i < a.length; i++) { dotProduct += a[i] * b[i]; normA += a[i] * a[i]; normB += b[i] * b[i]; } return dotProduct / (Math.sqrt(normA) * Math.sqrt(normB)); } // Rank documents by similarity to query const queryVector = queryEmbedding.embeddings.float[0]; const scored = docVectors.map((vec, index) => ({ index, score: cosineSimilarity(queryVector, vec), })); scored.sort((a, b) => b.score - a.score); console.log("Most similar:", scored[0]); ``` --- ## Reduced Dimensions with embed-v4 `embed-v4.0` supports `outputDimension` for faster similarity search at minimal quality loss. ```typescript const REDUCED_DIMENSION = 512; const response = await client.embed({ model: "embed-v4.0", inputType: "search_document", texts: documents, embeddingTypes: ["float"], outputDimension: REDUCED_DIMENSION, // 256 | 512 | 1024 | 1536 }); // Vectors are now 512-dimensional instead of default 1536 console.log(`Dimension: ${response.embeddings.float[0].length}`); ``` --- ## Compressed Embedding Types Use `int8` or `binary` for storage-efficient embeddings with minimal quality loss. ```typescript const response = await client.embed({ model: "embed-v4.0", inputType: "search_document", texts: documents, embeddingTypes: ["float", "int8"], }); // Access both formats const floatVectors = response.embeddings.float; // Full precision const int8Vectors = response.embeddings.int8; // 4x compressed ``` --- ## Batch Embedding (96 texts per call) ```typescript const BATCH_SIZE = 96; // Cohere max per embed() call async function embedAllDocuments( texts: string[], model: string, ): Promise<number[][]> { const allVectors: number[][] = []; for (let i = 0; i < texts.length; i += BATCH_SIZE) { const batch = texts.slice(i, i + BATCH_SIZE); const response = await client.embed({ model, inputType: "search_document", texts: batch, embeddingTypes: ["float"], }); allVectors.push(...response.embeddings.float); } return allVectors; } ``` --- ## Basic Rerank Score and reorder documents by relevance to a query. ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY }); const RERANK_MODEL = "rerank-v4.0-pro"; const TOP_N = 5; const result = await client.rerank({ model: RERANK_MODEL, query: "What is TypeScript?", documents: [ "TypeScript is a typed superset of JavaScript developed by Microsoft.", "Python is a general-purpose programming language.", "TypeScript adds static typing to JavaScript.", "Java is a class-based, object-oriented language.", "TypeScript compiles to plain JavaScript.", ], topN: TOP_N, }); for (const item of result.results) { console.log(`Doc[${item.index}] score: ${item.relevanceScore.toFixed(4)}`); } ``` --- ## Embed + Rerank Pipeline Two-stage retrieval: embed for initial recall, rerank for precision. ```typescript const EMBEDDING_MODEL = "embed-v4.0"; const RERANK_MODEL = "rerank-v4.0-pro"; const INITIAL_TOP_K = 20; const FINAL_TOP_N = 5; // Stage 1: Embed query and retrieve initial candidates via vector search const queryEmbedding = await client.embed({ model: EMBEDDING_MODEL, inputType: "search_query", texts: [userQuery], embeddingTypes: ["float"], }); // ... perform vector search to get INITIAL_TOP_K candidates ... const candidates: string[] = vectorSearchResults; // Stage 2: Rerank candidates for precision const reranked = await client.rerank({ model: RERANK_MODEL, query: userQuery, documents: candidates, topN: FINAL_TOP_N, }); // Top results by relevance const topDocuments = reranked.results.map((r) => ({ text: candidates[r.index], score: r.relevanceScore, })); ``` --- ## Classification Embeddings Use `inputType: "classification"` for text classifier inputs. ```typescript const response = await client.embed({ model: "embed-v4.0", inputType: "classification", texts: ["This product is amazing!", "Terrible experience."], embeddingTypes: ["float"], }); // Use these vectors as features for your classifier const featureVectors = response.embeddings.float; ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
tools-rag.md 11.2 KB
# Cohere SDK -- Tool Use & RAG Examples > Function calling, document grounding, citation handling, and full RAG pipelines. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, chat, streaming, error handling - [embeddings-rerank.md](embeddings-rerank.md) -- Embeddings, rerank, semantic search pipeline --- ## RAG with Inline Documents Pass documents directly to `chat()` for grounded answers with automatic citations. ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY }); const response = await client.chat({ model: "command-a-03-2025", messages: [ { role: "user", content: "What is TypeScript and who created it?" }, ], documents: [ { data: { text: "TypeScript is a typed superset of JavaScript that compiles to plain JavaScript.", title: "TypeScript Overview", url: "https://typescriptlang.org", }, }, { data: { text: "TypeScript was developed by Microsoft and first released in 2012.", title: "TypeScript History", }, }, ], }); // Grounded response text console.log(response.message.content[0].text); // Citations show which documents support each claim if (response.message.citations) { for (const citation of response.message.citations) { console.log(`"${citation.text}" [${citation.start}-${citation.end}]`); for (const source of citation.sources) { console.log(` Source: ${JSON.stringify(source)}`); } } } ``` **Why good:** Documents include metadata (title, url) for richer citations, response is grounded in provided facts --- ## V2 Document Format V2 uses `{ data: { ... } }` format for documents. The `data` object can contain any fields -- the model uses them for grounding. ```typescript // Good: V2 document format with data wrapper const documents = [ { data: { text: "Content here", title: "Title", id: "doc-1" } }, { data: { text: "More content", source: "internal-wiki" } }, ]; // BAD: V1 format (string or flat object) -- does not work with V2 const documents = [ "Content here", // V1 string format { text: "Content", title: "Title" }, // V1 flat object ]; ``` **Why bad:** V2 requires the `data` wrapper object. V1 formats cause errors or produce no citations. --- ## Full RAG Pipeline: Embed + Rerank + Chat ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY }); const EMBEDDING_MODEL = "embed-v4.0"; const RERANK_MODEL = "rerank-v4.0-pro"; const CHAT_MODEL = "command-a-03-2025"; const TOP_N = 3; async function ragQuery( query: string, corpus: Array<{ text: string; title: string }>, ): Promise<{ answer: string; citations: unknown[] }> { // Step 1: Embed the query const queryEmbed = await client.embed({ model: EMBEDDING_MODEL, inputType: "search_query", texts: [query], embeddingTypes: ["float"], }); // Step 2: Retrieve candidates (simplified -- use vector DB in production) // In production, use the query embedding against your vector store const candidateTexts = corpus.map((doc) => doc.text); // Step 3: Rerank for precision const reranked = await client.rerank({ model: RERANK_MODEL, query, documents: candidateTexts, topN: TOP_N, }); // Step 4: Build documents for chat const topDocs = reranked.results.map((r) => ({ data: { text: corpus[r.index].text, title: corpus[r.index].title, }, })); // Step 5: Chat with grounded documents const response = await client.chat({ model: CHAT_MODEL, messages: [{ role: "user", content: query }], documents: topDocs, }); return { answer: response.message.content[0].text, citations: response.message.citations ?? [], }; } ``` --- ## Tool Use: Basic Function Calling 4-step loop: user message -> model returns tool_calls -> execute tools -> return results. ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY }); // Step 1: Define tools with JSON Schema const tools = [ { type: "function" as const, function: { name: "get_weather", description: "Get the current weather for a location", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, unit: { type: "string", enum: ["celsius", "fahrenheit"] }, }, required: ["location"], }, }, }, ]; // Step 2: Send user message with tools const messages: Array<Record<string, unknown>> = [ { role: "user", content: "What is the weather in Tokyo?" }, ]; const response = await client.chat({ model: "command-a-03-2025", messages, tools, }); // Step 3: Check if model wants to call tools if (response.message.toolCalls && response.message.toolCalls.length > 0) { // CRITICAL: Append the FULL assistant message (including tool_calls and tool_plan) messages.push({ role: "assistant", toolPlan: response.message.toolPlan, toolCalls: response.message.toolCalls, }); // Step 4: Execute each tool and append results for (const toolCall of response.message.toolCalls) { const args = JSON.parse(toolCall.function.arguments); // Execute the actual function (simplified) const result = await getWeather(args.location, args.unit); // Append tool result with matching tool_call_id messages.push({ role: "tool", toolCallId: toolCall.id, content: [ { type: "document", document: { data: JSON.stringify(result) }, }, ], }); } // Step 5: Get final response with grounded citations const finalResponse = await client.chat({ model: "command-a-03-2025", messages, tools, }); console.log(finalResponse.message.content[0].text); // Citations reference the tool outputs if (finalResponse.message.citations) { for (const citation of finalResponse.message.citations) { console.log(`"${citation.text}": ${JSON.stringify(citation.sources)}`); } } } ``` --- ## Tool Use: Multiple Tools ```typescript const tools = [ { type: "function" as const, function: { name: "search_docs", description: "Search documentation and return relevant snippets", parameters: { type: "object", properties: { query: { type: "string", description: "Search query" }, topK: { type: "integer", description: "Max results" }, }, required: ["query"], }, }, }, { type: "function" as const, function: { name: "get_user_info", description: "Get information about a user by ID", parameters: { type: "object", properties: { userId: { type: "string", description: "User ID" }, }, required: ["userId"], }, }, }, ]; // The model may call multiple tools in a single response // Process ALL tool calls before sending results back const response = await client.chat({ model: "command-a-03-2025", messages: [ { role: "user", content: "Find docs about auth and look up user 123" }, ], tools, }); if (response.message.toolCalls) { messages.push({ role: "assistant", toolPlan: response.message.toolPlan, toolCalls: response.message.toolCalls, }); // Execute ALL tool calls for (const toolCall of response.message.toolCalls) { const args = JSON.parse(toolCall.function.arguments); const fn = functionMap[toolCall.function.name]; const result = await fn(args); messages.push({ role: "tool", toolCallId: toolCall.id, content: [ { type: "document", document: { data: JSON.stringify(result) } }, ], }); } // Model generates final response using all tool results const finalResponse = await client.chat({ model: "command-a-03-2025", messages, tools, }); } ``` --- ## Tool Result Document Format Tool results use the document format with `type: "document"` and `data`. ```typescript // Good: Document format with data messages.push({ role: "tool", toolCallId: toolCall.id, content: [ { type: "document", document: { data: JSON.stringify({ temperature: 22, condition: "sunny", location: "Tokyo", }), }, }, ], }); // BAD: Plain string content messages.push({ role: "tool", toolCallId: toolCall.id, content: "Temperature: 22, Condition: sunny", // May work but no citation grounding }); ``` **Why bad:** Using plain strings instead of document format means the model cannot generate fine-grained citations referencing specific fields from the tool output. --- ## Streaming with Tool Use ```typescript const stream = await client.chatStream({ model: "command-a-03-2025", messages: [{ role: "user", content: "What is the weather in Paris?" }], tools, }); let toolPlan = ""; const toolCalls: Array<{ id: string; name: string; arguments: string }> = []; let currentToolCall: { id: string; name: string; arguments: string } | null = null; for await (const event of stream) { switch (event.type) { case "tool-plan-delta": // Model's reasoning about which tools to call toolPlan += event.delta?.message?.toolPlan ?? ""; break; case "tool-call-start": currentToolCall = { id: event.delta?.message?.toolCalls?.id ?? "", name: event.delta?.message?.toolCalls?.function?.name ?? "", arguments: "", }; break; case "tool-call-delta": if (currentToolCall) { currentToolCall.arguments += event.delta?.message?.toolCalls?.function?.arguments ?? ""; } break; case "tool-call-end": if (currentToolCall) { toolCalls.push(currentToolCall); currentToolCall = null; } break; case "content-delta": process.stdout.write(event.delta?.message?.content?.text ?? ""); break; case "message-end": console.log(`\nFinish reason: ${event.delta?.finishReason}`); break; } } // If tool calls were collected, execute them and continue the conversation if (toolCalls.length > 0) { // ... execute tools and submit results as shown in previous examples } ``` --- ## Citation Processing Helper ```typescript interface ProcessedCitation { text: string; start: number; end: number; sourceIds: string[]; } function processCitations( responseText: string, citations: Array<{ start: number; end: number; text: string; sources: unknown[]; }>, ): ProcessedCitation[] { return citations.map((c) => ({ text: c.text, start: c.start, end: c.end, sourceIds: c.sources.map((s: { id?: string }) => s.id ?? "unknown"), })); } // Use with any response that has citations const response = await client.chat({ model: "command-a-03-2025", messages: [{ role: "user", content: "Tell me about TypeScript." }], documents: [{ data: { text: "TypeScript was created by Microsoft." } }], }); if (response.message.citations) { const processed = processCitations( response.message.content[0].text, response.message.citations, ); console.log("Citations:", processed); } ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
-
-
reference.md 9.2 KB
# Cohere SDK Quick Reference > Client configuration, model IDs, API methods, error types, and streaming events. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## Package Installation ```bash npm install cohere-ai ``` --- ## Client Configuration ```typescript import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY, // Required: API key timeout: 30_000, // Request timeout in ms (default: 120_000) clientName: "my-app", // Optional: for logging/analytics }); ``` ### Constructor Parameters | Parameter | Type | Default | Purpose | | ------------ | -------- | ---------------------------- | ------------------------------- | | `token` | `string` | (required) | API key for authentication | | `timeout` | `number` | `120_000` (2 min) | Request timeout in ms | | `clientName` | `string` | `undefined` | Client identifier for logging | | `baseUrl` | `string` | `"https://api.cohere.ai/v1"` | Override for custom deployments | --- ## Model IDs ### Command Models (Chat / Text Generation) | Model | Context | Max Output | Use Case | | ----------------------------- | ------- | ---------- | ---------------------------------------- | | `command-a-03-2025` | 256K | 8K | General purpose, strongest Command model | | `command-a-reasoning-08-2025` | 256K | 32K | Multi-step reasoning tasks | | `command-a-vision-07-2025` | 128K | 8K | Image/chart/document analysis | | `command-a-translate-08-2025` | 8K | -- | Translation (23 languages) | | `command-r7b-12-2024` | 128K | 4K | Lightweight, fast, RAG-optimized | | `command-r-08-2024` | 128K | 4K | General purpose (legacy) | | `command-r-plus-08-2024` | 128K | 4K | Enhanced quality (legacy) | ### Embed Models | Model | Context | Dimensions | Notes | | ------------------------------- | ------- | --------------- | ----------------------------- | | `embed-v4.0` | 128K | 256-1536 (flex) | Multimodal (text/images/PDFs) | | `embed-english-v3.0` | 512 | 1024 | English-focused | | `embed-english-light-v3.0` | 512 | 384 | Lightweight English | | `embed-multilingual-v3.0` | 512 | 1024 | 23-language support | | `embed-multilingual-light-v3.0` | 512 | 384 | Lightweight multilingual | ### Rerank Models | Model | Context | Notes | | -------------------------- | ------- | ----------------------------- | | `rerank-v4.0-pro` | 32K | Quality-focused, multilingual | | `rerank-v4.0-fast` | 32K | Latency-optimized | | `rerank-v3.5` | 4K | English, JSON support | | `rerank-english-v3.0` | 4K | English-only | | `rerank-multilingual-v3.0` | 4K | Non-English documents | --- ## API Methods Reference ### Chat (V2) ```typescript const response = await client.chat({ model: "command-a-03-2025", // Required messages: [], // Required: { role, content }[] temperature: 0.3, // 0-1 (default: 0.3) maxTokens: 4096, // Max output tokens tools: [], // Tool definitions documents: [], // Documents for RAG grounding citationOptions: { mode: "fast" }, // Citation generation mode safetyMode: "CONTEXTUAL", // "CONTEXTUAL" | "STRICT" | "OFF" stopSequences: [], // Stop generation at these strings toolChoice: "REQUIRED", // "REQUIRED" | "NONE" (optional, command-r7b+ only) strictTools: true, // Force tool calls to match schema exactly responseFormat: { type: "json_object" }, // Force JSON output thinking: { type: "enabled", tokenBudget: 4096 }, // Reasoning mode }); // Response access response.message.content[0].text; // Generated text response.message.citations; // Citation objects (if documents provided) response.message.toolCalls; // Tool calls (if tools provided) response.message.toolPlan; // Tool reasoning (if tools provided) response.finishReason; // "COMPLETE" | "MAX_TOKENS" response.usage; // { billedUnits, tokens } ``` ### Chat Stream (V2) ```typescript const stream = await client.chatStream({ model: "command-a-03-2025", // Required messages: [], // Required // Same optional params as chat() }); for await (const event of stream) { // Handle events by type } ``` ### Embed (V2) ```typescript const response = await client.embed({ model: "embed-v4.0", // Required inputType: "search_document", // Required for v3+ texts: [], // Up to 96 strings embeddingTypes: ["float"], // "float" | "int8" | "uint8" | "binary" | "ubinary" outputDimension: 1024, // embed-v4+ only: 256 | 512 | 1024 | 1536 truncate: "END", // "NONE" | "START" | "END" }); // Response access response.embeddings.float; // number[][] (one per input text) ``` ### Rerank (V2) ```typescript const response = await client.rerank({ model: "rerank-v4.0-pro", // Required query: "search query", // Required documents: ["doc1", "doc2"], // Required: string[] (max 1000) topN: 3, // Limit results maxTokensPerDoc: 4096, // Truncate long documents }); // Response access response.results; // { index, relevanceScore }[] ``` --- ## Message Roles (V2) | Role | Purpose | | ----------- | ------------------------------------------------ | | `system` | Instructions the model prioritizes | | `user` | User input | | `assistant` | Model response (including tool_calls, tool_plan) | | `tool` | Tool execution results (with tool_call_id) | --- ## Streaming Event Types | Event Type | Key Fields | Description | | ----------------- | --------------------------------------------- | ------------------------ | | `message-start` | `id`, `delta.message.role` | Stream begins | | `content-start` | `index` | Content block begins | | `content-delta` | `delta.message.content.text` | Text token received | | `content-end` | `index` | Content block ends | | `tool-plan-delta` | `delta.message.tool_plan` | Tool reasoning token | | `tool-call-start` | `delta.message.tool_calls` | Tool call begins | | `tool-call-delta` | `delta.message.tool_calls.function.arguments` | Tool arguments streaming | | `tool-call-end` | `index` | Tool call complete | | `citation-start` | `delta.message.citations` | Citation generated | | `citation-end` | `index` | Citation complete | | `message-end` | `delta.finish_reason`, `delta.usage` | Stream complete | --- ## Error Types | Error Class | When Thrown | | -------------------- | --------------------------------------- | | `CohereError` | API errors (4xx, 5xx) with `statusCode` | | `CohereTimeoutError` | Request exceeds configured timeout | ```typescript import { CohereError, CohereTimeoutError } from "cohere-ai"; // CohereError properties error.statusCode; // HTTP status code error.message; // Error message error.body; // Full error response body ``` --- ## Input Type Reference (Embeddings) | Value | Use Case | | ----------------- | ----------------------------------------------- | | `search_document` | Embeddings for documents stored in vector DB | | `search_query` | Embeddings for search queries against vector DB | | `classification` | Embeddings for text classification inputs | | `clustering` | Embeddings for clustering algorithms | | `image` | Embeddings for image inputs (embed-v4+ only) | **Critical:** Always pair `search_document` (indexing) with `search_query` (querying). Using the same type for both silently degrades similarity scores. --- ## Embedding Types | Type | Format | Models | Storage Size | | --------- | ---------- | ------ | -------------- | | `float` | `number[]` | All | Full precision | | `int8` | `number[]` | v3.0+ | 4x compressed | | `uint8` | `number[]` | v3.0+ | 4x compressed | | `binary` | `number[]` | v3.0+ | 32x compressed | | `ubinary` | `number[]` | v3.0+ | 32x compressed | | `base64` | `string` | v3.0+ | Base64-encoded | --- ## Citation Structure ```typescript interface Citation { start: number; // Start position in response text end: number; // End position in response text text: string; // Cited text span sources: Source[]; // Source references } // In RAG mode, sources reference document indices // In tool use mode, sources reference tool_call_id values ``` -
SKILL.md 20.1 KB
--- name: ai-provider-cohere-sdk description: Official Cohere TypeScript SDK patterns -- CohereClientV2, chat, embeddings, rerank, RAG with citations, tool use, streaming, and model selection --- # Cohere SDK Patterns > **Quick Guide:** Use the `cohere-ai` npm package with `CohereClientV2` for all new Cohere integrations. V2 API requires `model` on every call. Use `chatStream` for streaming with `content-delta` events. Embeddings require `inputType` matching your use case (`search_document` for indexing, `search_query` for querying). Rerank scores documents by relevance. RAG works by passing `documents` to `chat()` -- the model returns inline citations automatically. Tool use follows a 4-step loop: user message, model returns `tool_calls`, you execute and return results, model generates cited response. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST use `CohereClientV2` (not `CohereClient`) for all new code -- V2 is the current API with required `model` parameter)** **(You MUST specify `inputType` on every embed call -- `search_document` for indexing, `search_query` for querying -- mismatched types produce garbage similarity scores)** **(You MUST handle the tool use loop correctly: append the full assistant message (with `tool_calls`) to messages, then append `tool` role results with matching `tool_call_id`)** **(You MUST check `finish_reason` in responses -- `MAX_TOKENS` means the output was truncated)** **(You MUST never hardcode API keys -- pass via `token` constructor parameter sourced from environment variables)** </critical_requirements> --- **Auto-detection:** Cohere, cohere-ai, CohereClientV2, CohereClient, command-a, command-r, command-r-plus, embed-v4, rerank-v4, chatStream, content-delta, inputType, search_document, search_query, embeddingTypes, topN, CO_API_KEY, COHERE_API_KEY **When to use:** - Building applications with Cohere Command models (chat, generation, summarization) - Creating semantic search pipelines with Cohere embeddings - Adding relevance scoring to search results with Cohere Rerank - Implementing RAG with inline document grounding and automatic citations - Building agentic workflows with Cohere tool use / function calling - Streaming chat responses for real-time user interfaces **Key patterns covered:** - Client setup with `CohereClientV2` (token, timeout, platform configs) - Chat and streaming (`chat`, `chatStream`, event types) - Embeddings with `inputType` for search/classification/clustering - Rerank for relevance scoring and search result ordering - RAG with documents and automatic citation handling - Tool use / function calling with multi-step loops - Model selection (Command-A, Command-R, Embed v4, Rerank v4) **When NOT to use:** - Multi-provider applications needing OpenAI/Anthropic/Google switching -- use a unified provider SDK - React-specific chat UI hooks -- use a framework-integrated AI SDK - Simple text completion without Cohere-specific features (rerank, citations) --- ## Examples Index - [Core: Setup, Chat & Error Handling](examples/core.md) -- CohereClientV2 init, basic chat, streaming, error handling - [Embeddings & Rerank](examples/embeddings-rerank.md) -- Semantic search, input types, rerank scoring, RAG pipeline - [Tool Use & RAG](examples/tools-rag.md) -- Function calling, document grounding, citation handling - [Quick API Reference](reference.md) -- Model IDs, method signatures, event types, error classes --- <philosophy> ## Philosophy The Cohere TypeScript SDK (`cohere-ai`) provides **direct access to Cohere's API surface** -- chat, embeddings, rerank, and RAG with citations. The SDK is auto-generated from Cohere's API spec using Fern. **Core principles:** 1. **V2 API is current** -- `CohereClientV2` provides the modern API. `model` is required on every call. V1 methods on `CohereClient` are legacy. 2. **Embeddings are typed** -- The `inputType` parameter (`search_document`, `search_query`, `classification`, `clustering`) is mandatory for v3+ models. Mismatching input types between indexing and querying silently degrades results. 3. **RAG is first-class** -- Pass `documents` directly to `chat()` and the model returns grounded answers with inline citations. No external retrieval framework required for the grounding step. 4. **Rerank is a standalone primitive** -- Score and reorder search results without building a full RAG pipeline. Feed any list of documents and a query, get relevance scores back. 5. **Citations are automatic** -- When documents are provided (via RAG or tool results), the model generates fine-grained citations with start/end positions and source references. **When to use the Cohere SDK directly:** - You want Cohere-specific features: rerank, citation grounding, multilingual embeddings - You need semantic search with embed + rerank pipeline - You want RAG with automatic inline citations - You are building on Cohere's platform (or Bedrock/Azure/OCI with Cohere models) **When NOT to use:** - You need to switch between multiple LLM providers -- use a unified provider SDK - You want React-specific chat UI hooks -- use a framework-integrated AI SDK - You only need basic chat completion without Cohere differentiators </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Client Setup Initialize `CohereClientV2`. The `token` parameter is required (pass from environment). ```typescript // lib/cohere.ts -- basic setup import { CohereClientV2 } from "cohere-ai"; const client = new CohereClientV2({ token: process.env.CO_API_KEY, }); export { client }; ``` ```typescript // lib/cohere.ts -- production configuration const TIMEOUT_MS = 30_000; const client = new CohereClientV2({ token: process.env.CO_API_KEY, timeout: TIMEOUT_MS, }); ``` **Why good:** Explicit token from env var, named timeout constant, named export ```typescript // BAD: Hardcoded key, default CohereClient (V1) import { CohereClient } from "cohere-ai"; const client = new CohereClient({ token: "sk-abc123" }); ``` **Why bad:** Hardcoded API key is a security breach risk, `CohereClient` is the legacy V1 client **See:** [examples/core.md](examples/core.md) for error handling, platform configs (Bedrock, Azure) --- ### Pattern 2: Chat Completion V2 chat uses `messages` array with `system`, `user`, `assistant`, and `tool` roles. ```typescript const response = await client.chat({ model: "command-a-03-2025", messages: [ { role: "system", content: "You are a helpful coding assistant." }, { role: "user", content: "Explain TypeScript generics." }, ], }); console.log(response.message.content[0].text); ``` **Why good:** System message for instruction, `model` explicitly specified, correct V2 content access path ```typescript // BAD: Missing model (required in V2), wrong response access const response = await client.chat({ messages: [{ role: "user", content: "Hello" }], }); console.log(response.text); // WRONG: V2 uses response.message.content[0].text ``` **Why bad:** V2 requires `model`, response shape is `response.message.content[0].text` not `response.text` **See:** [examples/core.md](examples/core.md) for multi-turn, token tracking, temperature control --- ### Pattern 3: Streaming Use `chatStream` with `for await` and check event `type` for `content-delta`. ```typescript const stream = await client.chatStream({ model: "command-a-03-2025", messages: [{ role: "user", content: "Explain async/await." }], }); for await (const event of stream) { if (event.type === "content-delta") { process.stdout.write(event.delta?.message?.content?.text ?? ""); } } ``` **Why good:** Checks event type before accessing delta, handles nullable content safely ```typescript // BAD: Not checking event type for await (const event of stream) { console.log(event.delta?.message); // Many events don't have message delta } ``` **Why bad:** Only `content-delta` events have text content -- other events (`message-start`, `citation-start`, `tool-plan-delta`) have different shapes **See:** [examples/core.md](examples/core.md) for full streaming with all event types --- ### Pattern 4: Embeddings `inputType` is required for v3+ models. Mismatching types between indexing and querying silently degrades results. ```typescript const EMBEDDING_MODEL = "embed-v4.0"; // Index documents with search_document const docEmbeddings = await client.embed({ model: EMBEDDING_MODEL, inputType: "search_document", texts: ["TypeScript is a typed superset of JavaScript."], embeddingTypes: ["float"], }); // Query with search_query const queryEmbedding = await client.embed({ model: EMBEDDING_MODEL, inputType: "search_query", texts: ["What is TypeScript?"], embeddingTypes: ["float"], }); ``` **Why good:** Correct `inputType` pairing, `embeddingTypes` explicitly specified, named model constant ```typescript // BAD: Same inputType for both indexing and querying const docs = await client.embed({ model: "embed-v4.0", inputType: "search_query", // WRONG for documents texts: documents, embeddingTypes: ["float"], }); ``` **Why bad:** Using `search_query` for document indexing silently produces worse similarity scores -- documents must use `search_document` **See:** [examples/embeddings-rerank.md](examples/embeddings-rerank.md) for cosine similarity, dimension control, batch embedding --- ### Pattern 5: Rerank Score documents by relevance to a query. Returns ordered results with relevance scores. ```typescript const RERANK_MODEL = "rerank-v4.0-pro"; const TOP_N = 3; const result = await client.rerank({ model: RERANK_MODEL, query: "What is TypeScript?", documents: [ "TypeScript is a typed superset of JavaScript.", "Python is a general-purpose language.", "TypeScript compiles to JavaScript.", ], topN: TOP_N, }); for (const item of result.results) { console.log(`Doc ${item.index}: score ${item.relevanceScore}`); } ``` **Why good:** Named constants, `topN` limits results, accesses `index` and `relevanceScore` **See:** [examples/embeddings-rerank.md](examples/embeddings-rerank.md) for embed + rerank pipeline, rank fields --- ### Pattern 6: RAG with Documents Pass `documents` to `chat()` and the model returns grounded answers with inline citations. ```typescript const response = await client.chat({ model: "command-a-03-2025", messages: [{ role: "user", content: "What is TypeScript?" }], documents: [ { data: { text: "TypeScript is a typed superset of JavaScript.", title: "TS Docs", }, }, { data: { text: "TypeScript was developed by Microsoft.", title: "History", }, }, ], }); console.log(response.message.content[0].text); // Citations reference which documents support each claim if (response.message.citations) { for (const citation of response.message.citations) { console.log(`"${citation.text}" from doc ${citation.sources}`); } } ``` **Why good:** Documents passed inline with metadata, citations accessed from response, no external retrieval framework needed **See:** [examples/tools-rag.md](examples/tools-rag.md) for full RAG pipeline with embed + rerank + chat --- ### Pattern 7: Tool Use / Function Calling 4-step loop: user message -> model returns `tool_calls` -> execute tools -> return results with `tool_call_id`. ```typescript const tools = [ { type: "function" as const, function: { name: "get_weather", description: "Get weather for a city", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, }, required: ["location"], }, }, }, ]; const response = await client.chat({ model: "command-a-03-2025", messages: [{ role: "user", content: "Weather in Paris?" }], tools, }); // Check if model wants to call tools if (response.message.toolCalls) { // See examples/tools-rag.md for the complete tool execution loop } ``` **Why good:** Standard JSON Schema tool definition, checks for toolCalls before executing **See:** [examples/tools-rag.md](examples/tools-rag.md) for complete multi-step tool loop with tool result submission --- ### Pattern 8: Error Handling Catch `CohereError` for API errors, `CohereTimeoutError` for timeouts. ```typescript import { CohereError, CohereTimeoutError } from "cohere-ai"; try { const response = await client.chat({ model: "command-a-03-2025", messages: [{ role: "user", content: "Hello" }], }); } catch (error) { if (error instanceof CohereTimeoutError) { console.error("Request timed out"); } else if (error instanceof CohereError) { console.error(`API Error [${error.statusCode}]: ${error.message}`); console.error("Body:", error.body); } else { throw error; // Re-throw unknown errors } } ``` **Why good:** Specific error types with status codes, re-throws unexpected errors, timeout handled separately **See:** [examples/core.md](examples/core.md) for production error handling patterns </patterns> --- <performance> ## Performance Optimization ### Model Selection for Cost/Speed ``` General purpose (best) -> command-a-03-2025 (256K context, strongest) Reasoning tasks -> command-a-reasoning-08-2025 (multi-step reasoning) Vision/document analysis -> command-a-vision-07-2025 (images, charts, OCR) Translation -> command-a-translate-08-2025 (23 languages) Lightweight / edge -> command-r7b-12-2024 (7B, fast, 128K context) Legacy (still supported) -> command-r-08-2024, command-r-plus-08-2024 Embeddings (best) -> embed-v4.0 (multimodal, 128K context, flexible dims) Embeddings (English) -> embed-english-v3.0 (1024 dims) Embeddings (multilingual) -> embed-multilingual-v3.0 (23 languages) Rerank (quality) -> rerank-v4.0-pro (32K context, multilingual) Rerank (speed) -> rerank-v4.0-fast (32K context, latency-optimized) ``` ### Key Optimization Patterns - **Batch embeddings** -- pass up to 96 texts per `embed()` call instead of calling per-document - **Use `topN` in rerank** -- limit results to reduce response size and cost - **Use `outputDimension` with embed-v4** -- reduce dimensions (256/512/1024) for faster similarity search at minimal quality loss - **Check `finish_reason === "MAX_TOKENS"`** -- detect truncated output - **Use `temperature: 0`** for deterministic output (enables caching) - **Use embed-v4 `int8`/`binary` types** for compressed storage with minimal quality loss - **Use `strictTools: true`** to force tool calls to follow the schema exactly (structured outputs) - **Use `thinking: { type: "enabled" }`** with reasoning models for complex multi-step tasks - **Use `toolChoice: "REQUIRED"`** when you always want the model to call a tool (command-r7b+ only) </performance> --- <decision_framework> ## Decision Framework ### Which Client Class to Use ``` New project? +-- YES -> CohereClientV2 (always) +-- Existing V1 code? +-- Working fine? -> Keep CohereClient but plan migration +-- Need V2 features? -> Migrate to CohereClientV2 ``` ### Which Model to Choose ``` What is your task? +-- General chat/generation -> command-a-03-2025 (most capable) +-- Reasoning / multi-step -> command-a-reasoning-08-2025 +-- Image/document analysis -> command-a-vision-07-2025 +-- Translation -> command-a-translate-08-2025 +-- Lightweight / low latency -> command-r7b-12-2024 +-- Embeddings -> embed-v4.0 (or embed-english-v3.0 for English-only) +-- Rerank quality -> rerank-v4.0-pro +-- Rerank speed -> rerank-v4.0-fast ``` ### Embed `inputType` Selection ``` What are you embedding? +-- Documents for a search index -> "search_document" +-- Search queries against an index -> "search_query" +-- Text for a classifier -> "classification" +-- Text for clustering -> "clustering" +-- Images -> "image" (embed-v4+ only) ``` ### When to Use Rerank ``` Do you have search results to re-order? +-- YES -> Use rerank as a second-stage ranker | +-- Quality matters most? -> rerank-v4.0-pro | +-- Latency matters most? -> rerank-v4.0-fast +-- NO -> Not applicable (rerank needs existing results to score) ``` ### RAG Approach ``` Do you need grounded answers with citations? +-- YES -> Pass documents to chat() | +-- Have pre-retrieved documents? -> Pass directly via documents param | +-- Need retrieval first? -> Use embed + vector search + rerank pipeline, then pass top results to chat() +-- NO -> Use plain chat without documents ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Using `CohereClient` instead of `CohereClientV2` for new code (V1 is legacy) - Missing `model` parameter in V2 API calls (required on every call, unlike V1) - Using wrong `inputType` for embeddings (`search_query` for documents or vice versa -- silently degrades results) - Hardcoding API keys instead of using environment variables - Not appending the full assistant message (with `tool_calls`) before appending tool results in the tool use loop **Medium Priority Issues:** - Not specifying `embeddingTypes` (defaults may not match your storage format) - Ignoring `finish_reason: "MAX_TOKENS"` (output was silently truncated) - Not handling `CohereTimeoutError` separately from `CohereError` - Processing all stream events without checking `type` (only `content-delta` has text) - Using V1 parameter names (`preamble`, `connectors`, `conversation_id`) with V2 client **Common Mistakes:** - Accessing `response.text` instead of `response.message.content[0].text` (V2 response shape changed) - Forgetting that `embeddingTypes` is required in V2 Embed API - Not matching `tool_call_id` when submitting tool results (model cannot correlate results) - Using `documents` with string values instead of `{ data: { text: "..." } }` objects in V2 - Expecting `response.message.citations` to exist when no documents were provided (citations only appear with grounded responses) **Gotchas & Edge Cases:** - The SDK is in beta -- pin your `cohere-ai` version in package.json to avoid breaking changes - V2 API is NOT yet supported for cloud deployments (Bedrock, SageMaker, Azure, OCI) -- use V1 client for cloud platforms - `inputType` is camelCase in TypeScript SDK (`inputType`) but snake_case in the REST API (`input_type`) - Embed API accepts max 96 texts per call -- batch larger sets yourself - `embed-v4.0` supports `outputDimension` for flexible sizing (256, 512, 1024, 1536) but v3 models have fixed dimensions - Rerank `relevanceScore` is normalized 0-1 but not calibrated across queries -- compare scores within a single query only - Stream events include `tool-plan-delta` before `tool-call-start` -- the model's reasoning about which tool to call - V2 uses `system` role for instructions (V1 used `preamble` parameter) - Citation `sources` in tool use responses reference `tool_call_id` values, not document indices - The `clientName` constructor parameter is for logging/analytics, not authentication - `responseFormat: { type: "json_object" }` is NOT supported in RAG mode (with `documents`, `tools`, or `toolResults`) - `toolChoice` is only supported on `command-r7b-12-2024` and newer models - First requests with `strictTools: true` and a new tool set take longer (schema compilation) - `thinking` (reasoning mode) is only available on reasoning-capable models like `command-a-reasoning-08-2025` </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST use `CohereClientV2` (not `CohereClient`) for all new code -- V2 is the current API with required `model` parameter)** **(You MUST specify `inputType` on every embed call -- `search_document` for indexing, `search_query` for querying -- mismatched types produce garbage similarity scores)** **(You MUST handle the tool use loop correctly: append the full assistant message (with `tool_calls`) to messages, then append `tool` role results with matching `tool_call_id`)** **(You MUST check `finish_reason` in responses -- `MAX_TOKENS` means the output was truncated)** **(You MUST never hardcode API keys -- pass via `token` constructor parameter sourced from environment variables)** **Failure to follow these rules will produce broken embeddings, missing citations, or insecure AI integrations.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.