Claude Skill

ai-provider-cohere-sdk

Official Cohere TypeScript SDK patterns -- CohereClientV2, chat, embeddings, rerank, RAG with citations, tool use, streaming, and model selection

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_ai-provider-cohere-sdk_skills_ai-provider-cohere-sdk-3a51ef5.zip · 17 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-provider-cohere-sdk/skills/ai-provider-cohere-sdk
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Cohere SDK Patterns

Quick Guide: Use the cohere-ai npm package with CohereClientV2 for all new Cohere integrations. V2 API requires model on every call. Use chatStream for streaming with content-delta events. Embeddings require inputType matching your use case (search_document for indexing, search_query for querying). Rerank scores documents by relevance. RAG works by passing documents to chat() -- the model returns inline citations automatically. Tool use follows a 4-step loop: user message, model returns tool_calls, you execute and return results, model generates cited response.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST use CohereClientV2 (not CohereClient) for all new code -- V2 is the current API with required model parameter)

(You MUST specify inputType on every embed call -- search_document for indexing, search_query for querying -- mismatched types produce garbage similarity scores)

(You MUST handle the tool use loop correctly: append the full assistant message (with tool_calls) to messages, then append tool role results with matching tool_call_id)

(You MUST check finish_reason in responses -- MAX_TOKENS means the output was truncated)

(You MUST never hardcode API keys -- pass via token constructor parameter sourced from environment variables)

</critical_requirements>


Auto-detection: Cohere, cohere-ai, CohereClientV2, CohereClient, command-a, command-r, command-r-plus, embed-v4, rerank-v4, chatStream, content-delta, inputType, search_document, search_query, embeddingTypes, topN, CO_API_KEY, COHERE_API_KEY

When to use:

  • Building applications with Cohere Command models (chat, generation, summarization)
  • Creating semantic search pipelines with Cohere embeddings
  • Adding relevance scoring to search results with Cohere Rerank
  • Implementing RAG with inline document grounding and automatic citations
  • Building agentic workflows with Cohere tool use / function calling
  • Streaming chat responses for real-time user interfaces

Key patterns covered:

  • Client setup with CohereClientV2 (token, timeout, platform configs)
  • Chat and streaming (chat, chatStream, event types)
  • Embeddings with inputType for search/classification/clustering
  • Rerank for relevance scoring and search result ordering
  • RAG with documents and automatic citation handling
  • Tool use / function calling with multi-step loops
  • Model selection (Command-A, Command-R, Embed v4, Rerank v4)

When NOT to use:

  • Multi-provider applications needing OpenAI/Anthropic/Google switching -- use a unified provider SDK
  • React-specific chat UI hooks -- use a framework-integrated AI SDK
  • Simple text completion without Cohere-specific features (rerank, citations)

Examples Index





<decision_framework>

Decision Framework

Which Client Class to Use

New project?
+-- YES -> CohereClientV2 (always)
+-- Existing V1 code?
    +-- Working fine? -> Keep CohereClient but plan migration
    +-- Need V2 features? -> Migrate to CohereClientV2

Which Model to Choose

What is your task?
+-- General chat/generation -> command-a-03-2025 (most capable)
+-- Reasoning / multi-step -> command-a-reasoning-08-2025
+-- Image/document analysis -> command-a-vision-07-2025
+-- Translation -> command-a-translate-08-2025
+-- Lightweight / low latency -> command-r7b-12-2024
+-- Embeddings -> embed-v4.0 (or embed-english-v3.0 for English-only)
+-- Rerank quality -> rerank-v4.0-pro
+-- Rerank speed -> rerank-v4.0-fast

Embed inputType Selection

What are you embedding?
+-- Documents for a search index -> "search_document"
+-- Search queries against an index -> "search_query"
+-- Text for a classifier -> "classification"
+-- Text for clustering -> "clustering"
+-- Images -> "image" (embed-v4+ only)

When to Use Rerank

Do you have search results to re-order?
+-- YES -> Use rerank as a second-stage ranker
|   +-- Quality matters most? -> rerank-v4.0-pro
|   +-- Latency matters most? -> rerank-v4.0-fast
+-- NO -> Not applicable (rerank needs existing results to score)

RAG Approach

Do you need grounded answers with citations?
+-- YES -> Pass documents to chat()
|   +-- Have pre-retrieved documents? -> Pass directly via documents param
|   +-- Need retrieval first? -> Use embed + vector search + rerank pipeline, then pass top results to chat()
+-- NO -> Use plain chat without documents

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Using CohereClient instead of CohereClientV2 for new code (V1 is legacy)
  • Missing model parameter in V2 API calls (required on every call, unlike V1)
  • Using wrong inputType for embeddings (search_query for documents or vice versa -- silently degrades results)
  • Hardcoding API keys instead of using environment variables
  • Not appending the full assistant message (with tool_calls) before appending tool results in the tool use loop

Medium Priority Issues:

  • Not specifying embeddingTypes (defaults may not match your storage format)
  • Ignoring finish_reason: "MAX_TOKENS" (output was silently truncated)
  • Not handling CohereTimeoutError separately from CohereError
  • Processing all stream events without checking type (only content-delta has text)
  • Using V1 parameter names (preamble, connectors, conversation_id) with V2 client

Common Mistakes:

  • Accessing response.text instead of response.message.content[0].text (V2 response shape changed)
  • Forgetting that embeddingTypes is required in V2 Embed API
  • Not matching tool_call_id when submitting tool results (model cannot correlate results)
  • Using documents with string values instead of { data: { text: "..." } } objects in V2
  • Expecting response.message.citations to exist when no documents were provided (citations only appear with grounded responses)

Gotchas & Edge Cases:

  • The SDK is in beta -- pin your cohere-ai version in package.json to avoid breaking changes
  • V2 API is NOT yet supported for cloud deployments (Bedrock, SageMaker, Azure, OCI) -- use V1 client for cloud platforms
  • inputType is camelCase in TypeScript SDK (inputType) but snake_case in the REST API (input_type)
  • Embed API accepts max 96 texts per call -- batch larger sets yourself
  • embed-v4.0 supports outputDimension for flexible sizing (256, 512, 1024, 1536) but v3 models have fixed dimensions
  • Rerank relevanceScore is normalized 0-1 but not calibrated across queries -- compare scores within a single query only
  • Stream events include tool-plan-delta before tool-call-start -- the model's reasoning about which tool to call
  • V2 uses system role for instructions (V1 used preamble parameter)
  • Citation sources in tool use responses reference tool_call_id values, not document indices
  • The clientName constructor parameter is for logging/analytics, not authentication
  • responseFormat: { type: "json_object" } is NOT supported in RAG mode (with documents, tools, or toolResults)
  • toolChoice is only supported on command-r7b-12-2024 and newer models
  • First requests with strictTools: true and a new tool set take longer (schema compilation)
  • thinking (reasoning mode) is only available on reasoning-capable models like command-a-reasoning-08-2025

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST use CohereClientV2 (not CohereClient) for all new code -- V2 is the current API with required model parameter)

(You MUST specify inputType on every embed call -- search_document for indexing, search_query for querying -- mismatched types produce garbage similarity scores)

(You MUST handle the tool use loop correctly: append the full assistant message (with tool_calls) to messages, then append tool role results with matching tool_call_id)

(You MUST check finish_reason in responses -- MAX_TOKENS means the output was truncated)

(You MUST never hardcode API keys -- pass via token constructor parameter sourced from environment variables)

Failure to follow these rules will produce broken embeddings, missing citations, or insecure AI integrations.

</critical_reminders>

Files (skills)
  • examples
    • core.md 5.9 KB
      # Cohere SDK -- Setup, Chat & Streaming Examples
      
      > Client initialization, chat completions, streaming, multi-turn conversations, and error handling. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [embeddings-rerank.md](embeddings-rerank.md) -- Embeddings, rerank, semantic search pipeline
      - [tools-rag.md](tools-rag.md) -- Tool use, RAG with documents, citation handling
      
      ---
      
      ## Basic Client Setup
      
      ```typescript
      // lib/cohere.ts
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({
        token: process.env.CO_API_KEY,
      });
      
      export { client };
      ```
      
      ---
      
      ## Production Configuration
      
      ```typescript
      // lib/cohere.ts
      import { CohereClientV2 } from "cohere-ai";
      
      const TIMEOUT_MS = 30_000;
      
      const client = new CohereClientV2({
        token: process.env.CO_API_KEY,
        timeout: TIMEOUT_MS,
        clientName: "my-app",
      });
      
      export { client };
      ```
      
      ---
      
      ## Basic Chat Completion
      
      ```typescript
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({ token: process.env.CO_API_KEY });
      
      async function chat(userMessage: string): Promise<string> {
        const response = await client.chat({
          model: "command-a-03-2025",
          messages: [
            { role: "system", content: "You are a helpful assistant. Be concise." },
            { role: "user", content: userMessage },
          ],
        });
      
        const content = response.message.content[0].text;
        console.log(
          `Tokens: input=${response.usage?.tokens?.inputTokens}, output=${response.usage?.tokens?.outputTokens}`,
        );
      
        if (response.finishReason === "MAX_TOKENS") {
          console.warn("Response was truncated");
        }
      
        return content;
      }
      
      const answer = await chat("What is TypeScript in one sentence?");
      console.log(answer);
      ```
      
      ---
      
      ## Multi-Turn Conversation
      
      ```typescript
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({ token: process.env.CO_API_KEY });
      
      const messages: Array<{ role: string; content: string }> = [
        { role: "system", content: "You are a TypeScript expert." },
        { role: "user", content: "What is a union type?" },
      ];
      
      const response = await client.chat({
        model: "command-a-03-2025",
        messages,
      });
      
      // Append assistant response for next turn
      const assistantText = response.message.content[0].text;
      messages.push({ role: "assistant", content: assistantText });
      messages.push({ role: "user", content: "Give me a real-world example." });
      
      const followUp = await client.chat({
        model: "command-a-03-2025",
        messages,
      });
      
      console.log(followUp.message.content[0].text);
      ```
      
      ---
      
      ## Streaming with Content Deltas
      
      ```typescript
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({ token: process.env.CO_API_KEY });
      
      const stream = await client.chatStream({
        model: "command-a-03-2025",
        messages: [
          { role: "system", content: "You are a helpful assistant." },
          { role: "user", content: "Explain async/await in TypeScript." },
        ],
      });
      
      for await (const event of stream) {
        if (event.type === "content-delta") {
          process.stdout.write(event.delta?.message?.content?.text ?? "");
        }
        if (event.type === "message-end") {
          console.log(`\nFinish reason: ${event.delta?.finishReason}`);
        }
      }
      ```
      
      ---
      
      ## Streaming with All Event Types
      
      ```typescript
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({ token: process.env.CO_API_KEY });
      
      const stream = await client.chatStream({
        model: "command-a-03-2025",
        messages: [{ role: "user", content: "Tell me about TypeScript." }],
      });
      
      for await (const event of stream) {
        switch (event.type) {
          case "message-start":
            // Stream has begun
            break;
          case "content-delta":
            process.stdout.write(event.delta?.message?.content?.text ?? "");
            break;
          case "citation-start":
            // Citation generated (only with documents/tools)
            break;
          case "tool-plan-delta":
            // Model reasoning about which tool to call
            break;
          case "tool-call-start":
            // Tool call initiated
            break;
          case "message-end":
            console.log(`\nDone. Reason: ${event.delta?.finishReason}`);
            break;
        }
      }
      ```
      
      ---
      
      ## Production Error Handling
      
      ```typescript
      import { CohereClientV2, CohereError, CohereTimeoutError } from "cohere-ai";
      
      const TIMEOUT_MS = 30_000;
      
      const client = new CohereClientV2({
        token: process.env.CO_API_KEY,
        timeout: TIMEOUT_MS,
      });
      
      async function safeChat(prompt: string): Promise<string | null> {
        try {
          const response = await client.chat({
            model: "command-a-03-2025",
            messages: [
              { role: "system", content: "You are a helpful assistant." },
              { role: "user", content: prompt },
            ],
          });
      
          if (response.finishReason === "MAX_TOKENS") {
            console.warn("Response was truncated");
          }
      
          return response.message.content[0].text;
        } catch (error) {
          if (error instanceof CohereTimeoutError) {
            console.error("Request timed out");
            return null;
          }
      
          if (error instanceof CohereError) {
            console.error(`Cohere API Error [${error.statusCode}]: ${error.message}`);
      
            if (error.statusCode === 429) {
              console.error("Rate limited -- back off and retry");
            }
      
            if (error.statusCode === 401) {
              throw new Error(
                "Invalid API key. Check CO_API_KEY environment variable.",
              );
            }
      
            return null;
          }
      
          // Unknown errors should be re-thrown
          throw error;
        }
      }
      
      const result = await safeChat("Hello!");
      if (result) {
        console.log(result);
      }
      ```
      
      ---
      
      ## Temperature and Output Control
      
      ```typescript
      const MAX_TOKENS = 500;
      
      const response = await client.chat({
        model: "command-a-03-2025",
        messages: [{ role: "user", content: "Summarize this article." }],
        temperature: 0, // Deterministic output
        maxTokens: MAX_TOKENS,
      });
      
      if (response.finishReason === "MAX_TOKENS") {
        console.warn("Output was truncated -- increase maxTokens");
      }
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • embeddings-rerank.md 6.6 KB
      # Cohere SDK -- Embeddings & Rerank Examples
      
      > Embedding generation, input type pairing, cosine similarity, rerank scoring, and semantic search pipeline. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, chat, streaming, error handling
      - [tools-rag.md](tools-rag.md) -- Tool use, RAG with documents, citation handling
      
      ---
      
      ## Basic Embedding Generation
      
      ```typescript
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({ token: process.env.CO_API_KEY });
      const EMBEDDING_MODEL = "embed-v4.0";
      
      // Embed documents for storage in a vector database
      const response = await client.embed({
        model: EMBEDDING_MODEL,
        inputType: "search_document",
        texts: [
          "TypeScript is a typed superset of JavaScript.",
          "React is a library for building user interfaces.",
          "Node.js is a JavaScript runtime.",
        ],
        embeddingTypes: ["float"],
      });
      
      // response.embeddings.float is number[][] -- one vector per input text
      const vectors = response.embeddings.float;
      console.log(
        `Generated ${vectors.length} vectors of dimension ${vectors[0].length}`,
      );
      ```
      
      ---
      
      ## Correct Input Type Pairing
      
      The most critical gotcha with Cohere embeddings: `inputType` must match between indexing and querying.
      
      ```typescript
      const EMBEDDING_MODEL = "embed-v4.0";
      
      // INDEXING: Use "search_document" for documents going into your vector store
      const docEmbeddings = await client.embed({
        model: EMBEDDING_MODEL,
        inputType: "search_document", // Documents being indexed
        texts: documents,
        embeddingTypes: ["float"],
      });
      
      // QUERYING: Use "search_query" for the user's search query
      const queryEmbedding = await client.embed({
        model: EMBEDDING_MODEL,
        inputType: "search_query", // Query being searched
        texts: [userQuery],
        embeddingTypes: ["float"],
      });
      ```
      
      ```typescript
      // BAD: Using the same inputType for both
      const docs = await client.embed({
        model: EMBEDDING_MODEL,
        inputType: "search_query", // WRONG: documents should use "search_document"
        texts: documents,
        embeddingTypes: ["float"],
      });
      ```
      
      **Why bad:** Cohere trains separate embedding spaces for documents vs queries. Using the wrong type silently produces vectors that don't align well, resulting in degraded search quality. There is no error -- results just get worse.
      
      ---
      
      ## Cosine Similarity Search
      
      ```typescript
      function cosineSimilarity(a: number[], b: number[]): number {
        let dotProduct = 0;
        let normA = 0;
        let normB = 0;
      
        for (let i = 0; i < a.length; i++) {
          dotProduct += a[i] * b[i];
          normA += a[i] * a[i];
          normB += b[i] * b[i];
        }
      
        return dotProduct / (Math.sqrt(normA) * Math.sqrt(normB));
      }
      
      // Rank documents by similarity to query
      const queryVector = queryEmbedding.embeddings.float[0];
      const scored = docVectors.map((vec, index) => ({
        index,
        score: cosineSimilarity(queryVector, vec),
      }));
      
      scored.sort((a, b) => b.score - a.score);
      console.log("Most similar:", scored[0]);
      ```
      
      ---
      
      ## Reduced Dimensions with embed-v4
      
      `embed-v4.0` supports `outputDimension` for faster similarity search at minimal quality loss.
      
      ```typescript
      const REDUCED_DIMENSION = 512;
      
      const response = await client.embed({
        model: "embed-v4.0",
        inputType: "search_document",
        texts: documents,
        embeddingTypes: ["float"],
        outputDimension: REDUCED_DIMENSION, // 256 | 512 | 1024 | 1536
      });
      
      // Vectors are now 512-dimensional instead of default 1536
      console.log(`Dimension: ${response.embeddings.float[0].length}`);
      ```
      
      ---
      
      ## Compressed Embedding Types
      
      Use `int8` or `binary` for storage-efficient embeddings with minimal quality loss.
      
      ```typescript
      const response = await client.embed({
        model: "embed-v4.0",
        inputType: "search_document",
        texts: documents,
        embeddingTypes: ["float", "int8"],
      });
      
      // Access both formats
      const floatVectors = response.embeddings.float; // Full precision
      const int8Vectors = response.embeddings.int8; // 4x compressed
      ```
      
      ---
      
      ## Batch Embedding (96 texts per call)
      
      ```typescript
      const BATCH_SIZE = 96; // Cohere max per embed() call
      
      async function embedAllDocuments(
        texts: string[],
        model: string,
      ): Promise<number[][]> {
        const allVectors: number[][] = [];
      
        for (let i = 0; i < texts.length; i += BATCH_SIZE) {
          const batch = texts.slice(i, i + BATCH_SIZE);
          const response = await client.embed({
            model,
            inputType: "search_document",
            texts: batch,
            embeddingTypes: ["float"],
          });
          allVectors.push(...response.embeddings.float);
        }
      
        return allVectors;
      }
      ```
      
      ---
      
      ## Basic Rerank
      
      Score and reorder documents by relevance to a query.
      
      ```typescript
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({ token: process.env.CO_API_KEY });
      const RERANK_MODEL = "rerank-v4.0-pro";
      const TOP_N = 5;
      
      const result = await client.rerank({
        model: RERANK_MODEL,
        query: "What is TypeScript?",
        documents: [
          "TypeScript is a typed superset of JavaScript developed by Microsoft.",
          "Python is a general-purpose programming language.",
          "TypeScript adds static typing to JavaScript.",
          "Java is a class-based, object-oriented language.",
          "TypeScript compiles to plain JavaScript.",
        ],
        topN: TOP_N,
      });
      
      for (const item of result.results) {
        console.log(`Doc[${item.index}] score: ${item.relevanceScore.toFixed(4)}`);
      }
      ```
      
      ---
      
      ## Embed + Rerank Pipeline
      
      Two-stage retrieval: embed for initial recall, rerank for precision.
      
      ```typescript
      const EMBEDDING_MODEL = "embed-v4.0";
      const RERANK_MODEL = "rerank-v4.0-pro";
      const INITIAL_TOP_K = 20;
      const FINAL_TOP_N = 5;
      
      // Stage 1: Embed query and retrieve initial candidates via vector search
      const queryEmbedding = await client.embed({
        model: EMBEDDING_MODEL,
        inputType: "search_query",
        texts: [userQuery],
        embeddingTypes: ["float"],
      });
      
      // ... perform vector search to get INITIAL_TOP_K candidates ...
      const candidates: string[] = vectorSearchResults;
      
      // Stage 2: Rerank candidates for precision
      const reranked = await client.rerank({
        model: RERANK_MODEL,
        query: userQuery,
        documents: candidates,
        topN: FINAL_TOP_N,
      });
      
      // Top results by relevance
      const topDocuments = reranked.results.map((r) => ({
        text: candidates[r.index],
        score: r.relevanceScore,
      }));
      ```
      
      ---
      
      ## Classification Embeddings
      
      Use `inputType: "classification"` for text classifier inputs.
      
      ```typescript
      const response = await client.embed({
        model: "embed-v4.0",
        inputType: "classification",
        texts: ["This product is amazing!", "Terrible experience."],
        embeddingTypes: ["float"],
      });
      
      // Use these vectors as features for your classifier
      const featureVectors = response.embeddings.float;
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • tools-rag.md 11.2 KB
      # Cohere SDK -- Tool Use & RAG Examples
      
      > Function calling, document grounding, citation handling, and full RAG pipelines. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, chat, streaming, error handling
      - [embeddings-rerank.md](embeddings-rerank.md) -- Embeddings, rerank, semantic search pipeline
      
      ---
      
      ## RAG with Inline Documents
      
      Pass documents directly to `chat()` for grounded answers with automatic citations.
      
      ```typescript
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({ token: process.env.CO_API_KEY });
      
      const response = await client.chat({
        model: "command-a-03-2025",
        messages: [
          { role: "user", content: "What is TypeScript and who created it?" },
        ],
        documents: [
          {
            data: {
              text: "TypeScript is a typed superset of JavaScript that compiles to plain JavaScript.",
              title: "TypeScript Overview",
              url: "https://typescriptlang.org",
            },
          },
          {
            data: {
              text: "TypeScript was developed by Microsoft and first released in 2012.",
              title: "TypeScript History",
            },
          },
        ],
      });
      
      // Grounded response text
      console.log(response.message.content[0].text);
      
      // Citations show which documents support each claim
      if (response.message.citations) {
        for (const citation of response.message.citations) {
          console.log(`"${citation.text}" [${citation.start}-${citation.end}]`);
          for (const source of citation.sources) {
            console.log(`  Source: ${JSON.stringify(source)}`);
          }
        }
      }
      ```
      
      **Why good:** Documents include metadata (title, url) for richer citations, response is grounded in provided facts
      
      ---
      
      ## V2 Document Format
      
      V2 uses `{ data: { ... } }` format for documents. The `data` object can contain any fields -- the model uses them for grounding.
      
      ```typescript
      // Good: V2 document format with data wrapper
      const documents = [
        { data: { text: "Content here", title: "Title", id: "doc-1" } },
        { data: { text: "More content", source: "internal-wiki" } },
      ];
      
      // BAD: V1 format (string or flat object) -- does not work with V2
      const documents = [
        "Content here", // V1 string format
        { text: "Content", title: "Title" }, // V1 flat object
      ];
      ```
      
      **Why bad:** V2 requires the `data` wrapper object. V1 formats cause errors or produce no citations.
      
      ---
      
      ## Full RAG Pipeline: Embed + Rerank + Chat
      
      ```typescript
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({ token: process.env.CO_API_KEY });
      
      const EMBEDDING_MODEL = "embed-v4.0";
      const RERANK_MODEL = "rerank-v4.0-pro";
      const CHAT_MODEL = "command-a-03-2025";
      const TOP_N = 3;
      
      async function ragQuery(
        query: string,
        corpus: Array<{ text: string; title: string }>,
      ): Promise<{ answer: string; citations: unknown[] }> {
        // Step 1: Embed the query
        const queryEmbed = await client.embed({
          model: EMBEDDING_MODEL,
          inputType: "search_query",
          texts: [query],
          embeddingTypes: ["float"],
        });
      
        // Step 2: Retrieve candidates (simplified -- use vector DB in production)
        // In production, use the query embedding against your vector store
        const candidateTexts = corpus.map((doc) => doc.text);
      
        // Step 3: Rerank for precision
        const reranked = await client.rerank({
          model: RERANK_MODEL,
          query,
          documents: candidateTexts,
          topN: TOP_N,
        });
      
        // Step 4: Build documents for chat
        const topDocs = reranked.results.map((r) => ({
          data: {
            text: corpus[r.index].text,
            title: corpus[r.index].title,
          },
        }));
      
        // Step 5: Chat with grounded documents
        const response = await client.chat({
          model: CHAT_MODEL,
          messages: [{ role: "user", content: query }],
          documents: topDocs,
        });
      
        return {
          answer: response.message.content[0].text,
          citations: response.message.citations ?? [],
        };
      }
      ```
      
      ---
      
      ## Tool Use: Basic Function Calling
      
      4-step loop: user message -> model returns tool_calls -> execute tools -> return results.
      
      ```typescript
      import { CohereClientV2 } from "cohere-ai";
      
      const client = new CohereClientV2({ token: process.env.CO_API_KEY });
      
      // Step 1: Define tools with JSON Schema
      const tools = [
        {
          type: "function" as const,
          function: {
            name: "get_weather",
            description: "Get the current weather for a location",
            parameters: {
              type: "object",
              properties: {
                location: { type: "string", description: "City name" },
                unit: { type: "string", enum: ["celsius", "fahrenheit"] },
              },
              required: ["location"],
            },
          },
        },
      ];
      
      // Step 2: Send user message with tools
      const messages: Array<Record<string, unknown>> = [
        { role: "user", content: "What is the weather in Tokyo?" },
      ];
      
      const response = await client.chat({
        model: "command-a-03-2025",
        messages,
        tools,
      });
      
      // Step 3: Check if model wants to call tools
      if (response.message.toolCalls && response.message.toolCalls.length > 0) {
        // CRITICAL: Append the FULL assistant message (including tool_calls and tool_plan)
        messages.push({
          role: "assistant",
          toolPlan: response.message.toolPlan,
          toolCalls: response.message.toolCalls,
        });
      
        // Step 4: Execute each tool and append results
        for (const toolCall of response.message.toolCalls) {
          const args = JSON.parse(toolCall.function.arguments);
      
          // Execute the actual function (simplified)
          const result = await getWeather(args.location, args.unit);
      
          // Append tool result with matching tool_call_id
          messages.push({
            role: "tool",
            toolCallId: toolCall.id,
            content: [
              {
                type: "document",
                document: { data: JSON.stringify(result) },
              },
            ],
          });
        }
      
        // Step 5: Get final response with grounded citations
        const finalResponse = await client.chat({
          model: "command-a-03-2025",
          messages,
          tools,
        });
      
        console.log(finalResponse.message.content[0].text);
      
        // Citations reference the tool outputs
        if (finalResponse.message.citations) {
          for (const citation of finalResponse.message.citations) {
            console.log(`"${citation.text}": ${JSON.stringify(citation.sources)}`);
          }
        }
      }
      ```
      
      ---
      
      ## Tool Use: Multiple Tools
      
      ```typescript
      const tools = [
        {
          type: "function" as const,
          function: {
            name: "search_docs",
            description: "Search documentation and return relevant snippets",
            parameters: {
              type: "object",
              properties: {
                query: { type: "string", description: "Search query" },
                topK: { type: "integer", description: "Max results" },
              },
              required: ["query"],
            },
          },
        },
        {
          type: "function" as const,
          function: {
            name: "get_user_info",
            description: "Get information about a user by ID",
            parameters: {
              type: "object",
              properties: {
                userId: { type: "string", description: "User ID" },
              },
              required: ["userId"],
            },
          },
        },
      ];
      
      // The model may call multiple tools in a single response
      // Process ALL tool calls before sending results back
      const response = await client.chat({
        model: "command-a-03-2025",
        messages: [
          { role: "user", content: "Find docs about auth and look up user 123" },
        ],
        tools,
      });
      
      if (response.message.toolCalls) {
        messages.push({
          role: "assistant",
          toolPlan: response.message.toolPlan,
          toolCalls: response.message.toolCalls,
        });
      
        // Execute ALL tool calls
        for (const toolCall of response.message.toolCalls) {
          const args = JSON.parse(toolCall.function.arguments);
          const fn = functionMap[toolCall.function.name];
          const result = await fn(args);
      
          messages.push({
            role: "tool",
            toolCallId: toolCall.id,
            content: [
              { type: "document", document: { data: JSON.stringify(result) } },
            ],
          });
        }
      
        // Model generates final response using all tool results
        const finalResponse = await client.chat({
          model: "command-a-03-2025",
          messages,
          tools,
        });
      }
      ```
      
      ---
      
      ## Tool Result Document Format
      
      Tool results use the document format with `type: "document"` and `data`.
      
      ```typescript
      // Good: Document format with data
      messages.push({
        role: "tool",
        toolCallId: toolCall.id,
        content: [
          {
            type: "document",
            document: {
              data: JSON.stringify({
                temperature: 22,
                condition: "sunny",
                location: "Tokyo",
              }),
            },
          },
        ],
      });
      
      // BAD: Plain string content
      messages.push({
        role: "tool",
        toolCallId: toolCall.id,
        content: "Temperature: 22, Condition: sunny", // May work but no citation grounding
      });
      ```
      
      **Why bad:** Using plain strings instead of document format means the model cannot generate fine-grained citations referencing specific fields from the tool output.
      
      ---
      
      ## Streaming with Tool Use
      
      ```typescript
      const stream = await client.chatStream({
        model: "command-a-03-2025",
        messages: [{ role: "user", content: "What is the weather in Paris?" }],
        tools,
      });
      
      let toolPlan = "";
      const toolCalls: Array<{ id: string; name: string; arguments: string }> = [];
      let currentToolCall: { id: string; name: string; arguments: string } | null =
        null;
      
      for await (const event of stream) {
        switch (event.type) {
          case "tool-plan-delta":
            // Model's reasoning about which tools to call
            toolPlan += event.delta?.message?.toolPlan ?? "";
            break;
      
          case "tool-call-start":
            currentToolCall = {
              id: event.delta?.message?.toolCalls?.id ?? "",
              name: event.delta?.message?.toolCalls?.function?.name ?? "",
              arguments: "",
            };
            break;
      
          case "tool-call-delta":
            if (currentToolCall) {
              currentToolCall.arguments +=
                event.delta?.message?.toolCalls?.function?.arguments ?? "";
            }
            break;
      
          case "tool-call-end":
            if (currentToolCall) {
              toolCalls.push(currentToolCall);
              currentToolCall = null;
            }
            break;
      
          case "content-delta":
            process.stdout.write(event.delta?.message?.content?.text ?? "");
            break;
      
          case "message-end":
            console.log(`\nFinish reason: ${event.delta?.finishReason}`);
            break;
        }
      }
      
      // If tool calls were collected, execute them and continue the conversation
      if (toolCalls.length > 0) {
        // ... execute tools and submit results as shown in previous examples
      }
      ```
      
      ---
      
      ## Citation Processing Helper
      
      ```typescript
      interface ProcessedCitation {
        text: string;
        start: number;
        end: number;
        sourceIds: string[];
      }
      
      function processCitations(
        responseText: string,
        citations: Array<{
          start: number;
          end: number;
          text: string;
          sources: unknown[];
        }>,
      ): ProcessedCitation[] {
        return citations.map((c) => ({
          text: c.text,
          start: c.start,
          end: c.end,
          sourceIds: c.sources.map((s: { id?: string }) => s.id ?? "unknown"),
        }));
      }
      
      // Use with any response that has citations
      const response = await client.chat({
        model: "command-a-03-2025",
        messages: [{ role: "user", content: "Tell me about TypeScript." }],
        documents: [{ data: { text: "TypeScript was created by Microsoft." } }],
      });
      
      if (response.message.citations) {
        const processed = processCitations(
          response.message.content[0].text,
          response.message.citations,
        );
        console.log("Citations:", processed);
      }
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
  • reference.md 9.2 KB
    # Cohere SDK Quick Reference
    
    > Client configuration, model IDs, API methods, error types, and streaming events. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## Package Installation
    
    ```bash
    npm install cohere-ai
    ```
    
    ---
    
    ## Client Configuration
    
    ```typescript
    import { CohereClientV2 } from "cohere-ai";
    
    const client = new CohereClientV2({
      token: process.env.CO_API_KEY, // Required: API key
      timeout: 30_000, // Request timeout in ms (default: 120_000)
      clientName: "my-app", // Optional: for logging/analytics
    });
    ```
    
    ### Constructor Parameters
    
    | Parameter    | Type     | Default                      | Purpose                         |
    | ------------ | -------- | ---------------------------- | ------------------------------- |
    | `token`      | `string` | (required)                   | API key for authentication      |
    | `timeout`    | `number` | `120_000` (2 min)            | Request timeout in ms           |
    | `clientName` | `string` | `undefined`                  | Client identifier for logging   |
    | `baseUrl`    | `string` | `"https://api.cohere.ai/v1"` | Override for custom deployments |
    
    ---
    
    ## Model IDs
    
    ### Command Models (Chat / Text Generation)
    
    | Model                         | Context | Max Output | Use Case                                 |
    | ----------------------------- | ------- | ---------- | ---------------------------------------- |
    | `command-a-03-2025`           | 256K    | 8K         | General purpose, strongest Command model |
    | `command-a-reasoning-08-2025` | 256K    | 32K        | Multi-step reasoning tasks               |
    | `command-a-vision-07-2025`    | 128K    | 8K         | Image/chart/document analysis            |
    | `command-a-translate-08-2025` | 8K      | --         | Translation (23 languages)               |
    | `command-r7b-12-2024`         | 128K    | 4K         | Lightweight, fast, RAG-optimized         |
    | `command-r-08-2024`           | 128K    | 4K         | General purpose (legacy)                 |
    | `command-r-plus-08-2024`      | 128K    | 4K         | Enhanced quality (legacy)                |
    
    ### Embed Models
    
    | Model                           | Context | Dimensions      | Notes                         |
    | ------------------------------- | ------- | --------------- | ----------------------------- |
    | `embed-v4.0`                    | 128K    | 256-1536 (flex) | Multimodal (text/images/PDFs) |
    | `embed-english-v3.0`            | 512     | 1024            | English-focused               |
    | `embed-english-light-v3.0`      | 512     | 384             | Lightweight English           |
    | `embed-multilingual-v3.0`       | 512     | 1024            | 23-language support           |
    | `embed-multilingual-light-v3.0` | 512     | 384             | Lightweight multilingual      |
    
    ### Rerank Models
    
    | Model                      | Context | Notes                         |
    | -------------------------- | ------- | ----------------------------- |
    | `rerank-v4.0-pro`          | 32K     | Quality-focused, multilingual |
    | `rerank-v4.0-fast`         | 32K     | Latency-optimized             |
    | `rerank-v3.5`              | 4K      | English, JSON support         |
    | `rerank-english-v3.0`      | 4K      | English-only                  |
    | `rerank-multilingual-v3.0` | 4K      | Non-English documents         |
    
    ---
    
    ## API Methods Reference
    
    ### Chat (V2)
    
    ```typescript
    const response = await client.chat({
      model: "command-a-03-2025", // Required
      messages: [], // Required: { role, content }[]
      temperature: 0.3, // 0-1 (default: 0.3)
      maxTokens: 4096, // Max output tokens
      tools: [], // Tool definitions
      documents: [], // Documents for RAG grounding
      citationOptions: { mode: "fast" }, // Citation generation mode
      safetyMode: "CONTEXTUAL", // "CONTEXTUAL" | "STRICT" | "OFF"
      stopSequences: [], // Stop generation at these strings
      toolChoice: "REQUIRED", // "REQUIRED" | "NONE" (optional, command-r7b+ only)
      strictTools: true, // Force tool calls to match schema exactly
      responseFormat: { type: "json_object" }, // Force JSON output
      thinking: { type: "enabled", tokenBudget: 4096 }, // Reasoning mode
    });
    
    // Response access
    response.message.content[0].text; // Generated text
    response.message.citations; // Citation objects (if documents provided)
    response.message.toolCalls; // Tool calls (if tools provided)
    response.message.toolPlan; // Tool reasoning (if tools provided)
    response.finishReason; // "COMPLETE" | "MAX_TOKENS"
    response.usage; // { billedUnits, tokens }
    ```
    
    ### Chat Stream (V2)
    
    ```typescript
    const stream = await client.chatStream({
      model: "command-a-03-2025", // Required
      messages: [], // Required
      // Same optional params as chat()
    });
    
    for await (const event of stream) {
      // Handle events by type
    }
    ```
    
    ### Embed (V2)
    
    ```typescript
    const response = await client.embed({
      model: "embed-v4.0", // Required
      inputType: "search_document", // Required for v3+
      texts: [], // Up to 96 strings
      embeddingTypes: ["float"], // "float" | "int8" | "uint8" | "binary" | "ubinary"
      outputDimension: 1024, // embed-v4+ only: 256 | 512 | 1024 | 1536
      truncate: "END", // "NONE" | "START" | "END"
    });
    
    // Response access
    response.embeddings.float; // number[][] (one per input text)
    ```
    
    ### Rerank (V2)
    
    ```typescript
    const response = await client.rerank({
      model: "rerank-v4.0-pro", // Required
      query: "search query", // Required
      documents: ["doc1", "doc2"], // Required: string[] (max 1000)
      topN: 3, // Limit results
      maxTokensPerDoc: 4096, // Truncate long documents
    });
    
    // Response access
    response.results; // { index, relevanceScore }[]
    ```
    
    ---
    
    ## Message Roles (V2)
    
    | Role        | Purpose                                          |
    | ----------- | ------------------------------------------------ |
    | `system`    | Instructions the model prioritizes               |
    | `user`      | User input                                       |
    | `assistant` | Model response (including tool_calls, tool_plan) |
    | `tool`      | Tool execution results (with tool_call_id)       |
    
    ---
    
    ## Streaming Event Types
    
    | Event Type        | Key Fields                                    | Description              |
    | ----------------- | --------------------------------------------- | ------------------------ |
    | `message-start`   | `id`, `delta.message.role`                    | Stream begins            |
    | `content-start`   | `index`                                       | Content block begins     |
    | `content-delta`   | `delta.message.content.text`                  | Text token received      |
    | `content-end`     | `index`                                       | Content block ends       |
    | `tool-plan-delta` | `delta.message.tool_plan`                     | Tool reasoning token     |
    | `tool-call-start` | `delta.message.tool_calls`                    | Tool call begins         |
    | `tool-call-delta` | `delta.message.tool_calls.function.arguments` | Tool arguments streaming |
    | `tool-call-end`   | `index`                                       | Tool call complete       |
    | `citation-start`  | `delta.message.citations`                     | Citation generated       |
    | `citation-end`    | `index`                                       | Citation complete        |
    | `message-end`     | `delta.finish_reason`, `delta.usage`          | Stream complete          |
    
    ---
    
    ## Error Types
    
    | Error Class          | When Thrown                             |
    | -------------------- | --------------------------------------- |
    | `CohereError`        | API errors (4xx, 5xx) with `statusCode` |
    | `CohereTimeoutError` | Request exceeds configured timeout      |
    
    ```typescript
    import { CohereError, CohereTimeoutError } from "cohere-ai";
    
    // CohereError properties
    error.statusCode; // HTTP status code
    error.message; // Error message
    error.body; // Full error response body
    ```
    
    ---
    
    ## Input Type Reference (Embeddings)
    
    | Value             | Use Case                                        |
    | ----------------- | ----------------------------------------------- |
    | `search_document` | Embeddings for documents stored in vector DB    |
    | `search_query`    | Embeddings for search queries against vector DB |
    | `classification`  | Embeddings for text classification inputs       |
    | `clustering`      | Embeddings for clustering algorithms            |
    | `image`           | Embeddings for image inputs (embed-v4+ only)    |
    
    **Critical:** Always pair `search_document` (indexing) with `search_query` (querying). Using the same type for both silently degrades similarity scores.
    
    ---
    
    ## Embedding Types
    
    | Type      | Format     | Models | Storage Size   |
    | --------- | ---------- | ------ | -------------- |
    | `float`   | `number[]` | All    | Full precision |
    | `int8`    | `number[]` | v3.0+  | 4x compressed  |
    | `uint8`   | `number[]` | v3.0+  | 4x compressed  |
    | `binary`  | `number[]` | v3.0+  | 32x compressed |
    | `ubinary` | `number[]` | v3.0+  | 32x compressed |
    | `base64`  | `string`   | v3.0+  | Base64-encoded |
    
    ---
    
    ## Citation Structure
    
    ```typescript
    interface Citation {
      start: number; // Start position in response text
      end: number; // End position in response text
      text: string; // Cited text span
      sources: Source[]; // Source references
    }
    
    // In RAG mode, sources reference document indices
    // In tool use mode, sources reference tool_call_id values
    ```
    
  • SKILL.md 20.1 KB
    ---
    name: ai-provider-cohere-sdk
    description: Official Cohere TypeScript SDK patterns -- CohereClientV2, chat, embeddings, rerank, RAG with citations, tool use, streaming, and model selection
    ---
    
    # Cohere SDK Patterns
    
    > **Quick Guide:** Use the `cohere-ai` npm package with `CohereClientV2` for all new Cohere integrations. V2 API requires `model` on every call. Use `chatStream` for streaming with `content-delta` events. Embeddings require `inputType` matching your use case (`search_document` for indexing, `search_query` for querying). Rerank scores documents by relevance. RAG works by passing `documents` to `chat()` -- the model returns inline citations automatically. Tool use follows a 4-step loop: user message, model returns `tool_calls`, you execute and return results, model generates cited response.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST use `CohereClientV2` (not `CohereClient`) for all new code -- V2 is the current API with required `model` parameter)**
    
    **(You MUST specify `inputType` on every embed call -- `search_document` for indexing, `search_query` for querying -- mismatched types produce garbage similarity scores)**
    
    **(You MUST handle the tool use loop correctly: append the full assistant message (with `tool_calls`) to messages, then append `tool` role results with matching `tool_call_id`)**
    
    **(You MUST check `finish_reason` in responses -- `MAX_TOKENS` means the output was truncated)**
    
    **(You MUST never hardcode API keys -- pass via `token` constructor parameter sourced from environment variables)**
    
    </critical_requirements>
    
    ---
    
    **Auto-detection:** Cohere, cohere-ai, CohereClientV2, CohereClient, command-a, command-r, command-r-plus, embed-v4, rerank-v4, chatStream, content-delta, inputType, search_document, search_query, embeddingTypes, topN, CO_API_KEY, COHERE_API_KEY
    
    **When to use:**
    
    - Building applications with Cohere Command models (chat, generation, summarization)
    - Creating semantic search pipelines with Cohere embeddings
    - Adding relevance scoring to search results with Cohere Rerank
    - Implementing RAG with inline document grounding and automatic citations
    - Building agentic workflows with Cohere tool use / function calling
    - Streaming chat responses for real-time user interfaces
    
    **Key patterns covered:**
    
    - Client setup with `CohereClientV2` (token, timeout, platform configs)
    - Chat and streaming (`chat`, `chatStream`, event types)
    - Embeddings with `inputType` for search/classification/clustering
    - Rerank for relevance scoring and search result ordering
    - RAG with documents and automatic citation handling
    - Tool use / function calling with multi-step loops
    - Model selection (Command-A, Command-R, Embed v4, Rerank v4)
    
    **When NOT to use:**
    
    - Multi-provider applications needing OpenAI/Anthropic/Google switching -- use a unified provider SDK
    - React-specific chat UI hooks -- use a framework-integrated AI SDK
    - Simple text completion without Cohere-specific features (rerank, citations)
    
    ---
    
    ## Examples Index
    
    - [Core: Setup, Chat & Error Handling](examples/core.md) -- CohereClientV2 init, basic chat, streaming, error handling
    - [Embeddings & Rerank](examples/embeddings-rerank.md) -- Semantic search, input types, rerank scoring, RAG pipeline
    - [Tool Use & RAG](examples/tools-rag.md) -- Function calling, document grounding, citation handling
    - [Quick API Reference](reference.md) -- Model IDs, method signatures, event types, error classes
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    The Cohere TypeScript SDK (`cohere-ai`) provides **direct access to Cohere's API surface** -- chat, embeddings, rerank, and RAG with citations. The SDK is auto-generated from Cohere's API spec using Fern.
    
    **Core principles:**
    
    1. **V2 API is current** -- `CohereClientV2` provides the modern API. `model` is required on every call. V1 methods on `CohereClient` are legacy.
    2. **Embeddings are typed** -- The `inputType` parameter (`search_document`, `search_query`, `classification`, `clustering`) is mandatory for v3+ models. Mismatching input types between indexing and querying silently degrades results.
    3. **RAG is first-class** -- Pass `documents` directly to `chat()` and the model returns grounded answers with inline citations. No external retrieval framework required for the grounding step.
    4. **Rerank is a standalone primitive** -- Score and reorder search results without building a full RAG pipeline. Feed any list of documents and a query, get relevance scores back.
    5. **Citations are automatic** -- When documents are provided (via RAG or tool results), the model generates fine-grained citations with start/end positions and source references.
    
    **When to use the Cohere SDK directly:**
    
    - You want Cohere-specific features: rerank, citation grounding, multilingual embeddings
    - You need semantic search with embed + rerank pipeline
    - You want RAG with automatic inline citations
    - You are building on Cohere's platform (or Bedrock/Azure/OCI with Cohere models)
    
    **When NOT to use:**
    
    - You need to switch between multiple LLM providers -- use a unified provider SDK
    - You want React-specific chat UI hooks -- use a framework-integrated AI SDK
    - You only need basic chat completion without Cohere differentiators
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Client Setup
    
    Initialize `CohereClientV2`. The `token` parameter is required (pass from environment).
    
    ```typescript
    // lib/cohere.ts -- basic setup
    import { CohereClientV2 } from "cohere-ai";
    
    const client = new CohereClientV2({
      token: process.env.CO_API_KEY,
    });
    
    export { client };
    ```
    
    ```typescript
    // lib/cohere.ts -- production configuration
    const TIMEOUT_MS = 30_000;
    
    const client = new CohereClientV2({
      token: process.env.CO_API_KEY,
      timeout: TIMEOUT_MS,
    });
    ```
    
    **Why good:** Explicit token from env var, named timeout constant, named export
    
    ```typescript
    // BAD: Hardcoded key, default CohereClient (V1)
    import { CohereClient } from "cohere-ai";
    const client = new CohereClient({ token: "sk-abc123" });
    ```
    
    **Why bad:** Hardcoded API key is a security breach risk, `CohereClient` is the legacy V1 client
    
    **See:** [examples/core.md](examples/core.md) for error handling, platform configs (Bedrock, Azure)
    
    ---
    
    ### Pattern 2: Chat Completion
    
    V2 chat uses `messages` array with `system`, `user`, `assistant`, and `tool` roles.
    
    ```typescript
    const response = await client.chat({
      model: "command-a-03-2025",
      messages: [
        { role: "system", content: "You are a helpful coding assistant." },
        { role: "user", content: "Explain TypeScript generics." },
      ],
    });
    
    console.log(response.message.content[0].text);
    ```
    
    **Why good:** System message for instruction, `model` explicitly specified, correct V2 content access path
    
    ```typescript
    // BAD: Missing model (required in V2), wrong response access
    const response = await client.chat({
      messages: [{ role: "user", content: "Hello" }],
    });
    console.log(response.text); // WRONG: V2 uses response.message.content[0].text
    ```
    
    **Why bad:** V2 requires `model`, response shape is `response.message.content[0].text` not `response.text`
    
    **See:** [examples/core.md](examples/core.md) for multi-turn, token tracking, temperature control
    
    ---
    
    ### Pattern 3: Streaming
    
    Use `chatStream` with `for await` and check event `type` for `content-delta`.
    
    ```typescript
    const stream = await client.chatStream({
      model: "command-a-03-2025",
      messages: [{ role: "user", content: "Explain async/await." }],
    });
    
    for await (const event of stream) {
      if (event.type === "content-delta") {
        process.stdout.write(event.delta?.message?.content?.text ?? "");
      }
    }
    ```
    
    **Why good:** Checks event type before accessing delta, handles nullable content safely
    
    ```typescript
    // BAD: Not checking event type
    for await (const event of stream) {
      console.log(event.delta?.message); // Many events don't have message delta
    }
    ```
    
    **Why bad:** Only `content-delta` events have text content -- other events (`message-start`, `citation-start`, `tool-plan-delta`) have different shapes
    
    **See:** [examples/core.md](examples/core.md) for full streaming with all event types
    
    ---
    
    ### Pattern 4: Embeddings
    
    `inputType` is required for v3+ models. Mismatching types between indexing and querying silently degrades results.
    
    ```typescript
    const EMBEDDING_MODEL = "embed-v4.0";
    
    // Index documents with search_document
    const docEmbeddings = await client.embed({
      model: EMBEDDING_MODEL,
      inputType: "search_document",
      texts: ["TypeScript is a typed superset of JavaScript."],
      embeddingTypes: ["float"],
    });
    
    // Query with search_query
    const queryEmbedding = await client.embed({
      model: EMBEDDING_MODEL,
      inputType: "search_query",
      texts: ["What is TypeScript?"],
      embeddingTypes: ["float"],
    });
    ```
    
    **Why good:** Correct `inputType` pairing, `embeddingTypes` explicitly specified, named model constant
    
    ```typescript
    // BAD: Same inputType for both indexing and querying
    const docs = await client.embed({
      model: "embed-v4.0",
      inputType: "search_query", // WRONG for documents
      texts: documents,
      embeddingTypes: ["float"],
    });
    ```
    
    **Why bad:** Using `search_query` for document indexing silently produces worse similarity scores -- documents must use `search_document`
    
    **See:** [examples/embeddings-rerank.md](examples/embeddings-rerank.md) for cosine similarity, dimension control, batch embedding
    
    ---
    
    ### Pattern 5: Rerank
    
    Score documents by relevance to a query. Returns ordered results with relevance scores.
    
    ```typescript
    const RERANK_MODEL = "rerank-v4.0-pro";
    const TOP_N = 3;
    
    const result = await client.rerank({
      model: RERANK_MODEL,
      query: "What is TypeScript?",
      documents: [
        "TypeScript is a typed superset of JavaScript.",
        "Python is a general-purpose language.",
        "TypeScript compiles to JavaScript.",
      ],
      topN: TOP_N,
    });
    
    for (const item of result.results) {
      console.log(`Doc ${item.index}: score ${item.relevanceScore}`);
    }
    ```
    
    **Why good:** Named constants, `topN` limits results, accesses `index` and `relevanceScore`
    
    **See:** [examples/embeddings-rerank.md](examples/embeddings-rerank.md) for embed + rerank pipeline, rank fields
    
    ---
    
    ### Pattern 6: RAG with Documents
    
    Pass `documents` to `chat()` and the model returns grounded answers with inline citations.
    
    ```typescript
    const response = await client.chat({
      model: "command-a-03-2025",
      messages: [{ role: "user", content: "What is TypeScript?" }],
      documents: [
        {
          data: {
            text: "TypeScript is a typed superset of JavaScript.",
            title: "TS Docs",
          },
        },
        {
          data: {
            text: "TypeScript was developed by Microsoft.",
            title: "History",
          },
        },
      ],
    });
    
    console.log(response.message.content[0].text);
    
    // Citations reference which documents support each claim
    if (response.message.citations) {
      for (const citation of response.message.citations) {
        console.log(`"${citation.text}" from doc ${citation.sources}`);
      }
    }
    ```
    
    **Why good:** Documents passed inline with metadata, citations accessed from response, no external retrieval framework needed
    
    **See:** [examples/tools-rag.md](examples/tools-rag.md) for full RAG pipeline with embed + rerank + chat
    
    ---
    
    ### Pattern 7: Tool Use / Function Calling
    
    4-step loop: user message -> model returns `tool_calls` -> execute tools -> return results with `tool_call_id`.
    
    ```typescript
    const tools = [
      {
        type: "function" as const,
        function: {
          name: "get_weather",
          description: "Get weather for a city",
          parameters: {
            type: "object",
            properties: {
              location: { type: "string", description: "City name" },
            },
            required: ["location"],
          },
        },
      },
    ];
    
    const response = await client.chat({
      model: "command-a-03-2025",
      messages: [{ role: "user", content: "Weather in Paris?" }],
      tools,
    });
    
    // Check if model wants to call tools
    if (response.message.toolCalls) {
      // See examples/tools-rag.md for the complete tool execution loop
    }
    ```
    
    **Why good:** Standard JSON Schema tool definition, checks for toolCalls before executing
    
    **See:** [examples/tools-rag.md](examples/tools-rag.md) for complete multi-step tool loop with tool result submission
    
    ---
    
    ### Pattern 8: Error Handling
    
    Catch `CohereError` for API errors, `CohereTimeoutError` for timeouts.
    
    ```typescript
    import { CohereError, CohereTimeoutError } from "cohere-ai";
    
    try {
      const response = await client.chat({
        model: "command-a-03-2025",
        messages: [{ role: "user", content: "Hello" }],
      });
    } catch (error) {
      if (error instanceof CohereTimeoutError) {
        console.error("Request timed out");
      } else if (error instanceof CohereError) {
        console.error(`API Error [${error.statusCode}]: ${error.message}`);
        console.error("Body:", error.body);
      } else {
        throw error; // Re-throw unknown errors
      }
    }
    ```
    
    **Why good:** Specific error types with status codes, re-throws unexpected errors, timeout handled separately
    
    **See:** [examples/core.md](examples/core.md) for production error handling patterns
    
    </patterns>
    
    ---
    
    <performance>
    
    ## Performance Optimization
    
    ### Model Selection for Cost/Speed
    
    ```
    General purpose (best)      -> command-a-03-2025 (256K context, strongest)
    Reasoning tasks             -> command-a-reasoning-08-2025 (multi-step reasoning)
    Vision/document analysis    -> command-a-vision-07-2025 (images, charts, OCR)
    Translation                 -> command-a-translate-08-2025 (23 languages)
    Lightweight / edge          -> command-r7b-12-2024 (7B, fast, 128K context)
    Legacy (still supported)    -> command-r-08-2024, command-r-plus-08-2024
    Embeddings (best)           -> embed-v4.0 (multimodal, 128K context, flexible dims)
    Embeddings (English)        -> embed-english-v3.0 (1024 dims)
    Embeddings (multilingual)   -> embed-multilingual-v3.0 (23 languages)
    Rerank (quality)            -> rerank-v4.0-pro (32K context, multilingual)
    Rerank (speed)              -> rerank-v4.0-fast (32K context, latency-optimized)
    ```
    
    ### Key Optimization Patterns
    
    - **Batch embeddings** -- pass up to 96 texts per `embed()` call instead of calling per-document
    - **Use `topN` in rerank** -- limit results to reduce response size and cost
    - **Use `outputDimension` with embed-v4** -- reduce dimensions (256/512/1024) for faster similarity search at minimal quality loss
    - **Check `finish_reason === "MAX_TOKENS"`** -- detect truncated output
    - **Use `temperature: 0`** for deterministic output (enables caching)
    - **Use embed-v4 `int8`/`binary` types** for compressed storage with minimal quality loss
    - **Use `strictTools: true`** to force tool calls to follow the schema exactly (structured outputs)
    - **Use `thinking: { type: "enabled" }`** with reasoning models for complex multi-step tasks
    - **Use `toolChoice: "REQUIRED"`** when you always want the model to call a tool (command-r7b+ only)
    
    </performance>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Client Class to Use
    
    ```
    New project?
    +-- YES -> CohereClientV2 (always)
    +-- Existing V1 code?
        +-- Working fine? -> Keep CohereClient but plan migration
        +-- Need V2 features? -> Migrate to CohereClientV2
    ```
    
    ### Which Model to Choose
    
    ```
    What is your task?
    +-- General chat/generation -> command-a-03-2025 (most capable)
    +-- Reasoning / multi-step -> command-a-reasoning-08-2025
    +-- Image/document analysis -> command-a-vision-07-2025
    +-- Translation -> command-a-translate-08-2025
    +-- Lightweight / low latency -> command-r7b-12-2024
    +-- Embeddings -> embed-v4.0 (or embed-english-v3.0 for English-only)
    +-- Rerank quality -> rerank-v4.0-pro
    +-- Rerank speed -> rerank-v4.0-fast
    ```
    
    ### Embed `inputType` Selection
    
    ```
    What are you embedding?
    +-- Documents for a search index -> "search_document"
    +-- Search queries against an index -> "search_query"
    +-- Text for a classifier -> "classification"
    +-- Text for clustering -> "clustering"
    +-- Images -> "image" (embed-v4+ only)
    ```
    
    ### When to Use Rerank
    
    ```
    Do you have search results to re-order?
    +-- YES -> Use rerank as a second-stage ranker
    |   +-- Quality matters most? -> rerank-v4.0-pro
    |   +-- Latency matters most? -> rerank-v4.0-fast
    +-- NO -> Not applicable (rerank needs existing results to score)
    ```
    
    ### RAG Approach
    
    ```
    Do you need grounded answers with citations?
    +-- YES -> Pass documents to chat()
    |   +-- Have pre-retrieved documents? -> Pass directly via documents param
    |   +-- Need retrieval first? -> Use embed + vector search + rerank pipeline, then pass top results to chat()
    +-- NO -> Use plain chat without documents
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Using `CohereClient` instead of `CohereClientV2` for new code (V1 is legacy)
    - Missing `model` parameter in V2 API calls (required on every call, unlike V1)
    - Using wrong `inputType` for embeddings (`search_query` for documents or vice versa -- silently degrades results)
    - Hardcoding API keys instead of using environment variables
    - Not appending the full assistant message (with `tool_calls`) before appending tool results in the tool use loop
    
    **Medium Priority Issues:**
    
    - Not specifying `embeddingTypes` (defaults may not match your storage format)
    - Ignoring `finish_reason: "MAX_TOKENS"` (output was silently truncated)
    - Not handling `CohereTimeoutError` separately from `CohereError`
    - Processing all stream events without checking `type` (only `content-delta` has text)
    - Using V1 parameter names (`preamble`, `connectors`, `conversation_id`) with V2 client
    
    **Common Mistakes:**
    
    - Accessing `response.text` instead of `response.message.content[0].text` (V2 response shape changed)
    - Forgetting that `embeddingTypes` is required in V2 Embed API
    - Not matching `tool_call_id` when submitting tool results (model cannot correlate results)
    - Using `documents` with string values instead of `{ data: { text: "..." } }` objects in V2
    - Expecting `response.message.citations` to exist when no documents were provided (citations only appear with grounded responses)
    
    **Gotchas & Edge Cases:**
    
    - The SDK is in beta -- pin your `cohere-ai` version in package.json to avoid breaking changes
    - V2 API is NOT yet supported for cloud deployments (Bedrock, SageMaker, Azure, OCI) -- use V1 client for cloud platforms
    - `inputType` is camelCase in TypeScript SDK (`inputType`) but snake_case in the REST API (`input_type`)
    - Embed API accepts max 96 texts per call -- batch larger sets yourself
    - `embed-v4.0` supports `outputDimension` for flexible sizing (256, 512, 1024, 1536) but v3 models have fixed dimensions
    - Rerank `relevanceScore` is normalized 0-1 but not calibrated across queries -- compare scores within a single query only
    - Stream events include `tool-plan-delta` before `tool-call-start` -- the model's reasoning about which tool to call
    - V2 uses `system` role for instructions (V1 used `preamble` parameter)
    - Citation `sources` in tool use responses reference `tool_call_id` values, not document indices
    - The `clientName` constructor parameter is for logging/analytics, not authentication
    - `responseFormat: { type: "json_object" }` is NOT supported in RAG mode (with `documents`, `tools`, or `toolResults`)
    - `toolChoice` is only supported on `command-r7b-12-2024` and newer models
    - First requests with `strictTools: true` and a new tool set take longer (schema compilation)
    - `thinking` (reasoning mode) is only available on reasoning-capable models like `command-a-reasoning-08-2025`
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST use `CohereClientV2` (not `CohereClient`) for all new code -- V2 is the current API with required `model` parameter)**
    
    **(You MUST specify `inputType` on every embed call -- `search_document` for indexing, `search_query` for querying -- mismatched types produce garbage similarity scores)**
    
    **(You MUST handle the tool use loop correctly: append the full assistant message (with `tool_calls`) to messages, then append `tool` role results with matching `tool_call_id`)**
    
    **(You MUST check `finish_reason` in responses -- `MAX_TOKENS` means the output was truncated)**
    
    **(You MUST never hardcode API keys -- pass via `token` constructor parameter sourced from environment variables)**
    
    **Failure to follow these rules will produce broken embeddings, missing citations, or insecure AI integrations.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related